Technical field guide
The sections below preserve the service-specific depth behind eDiscovery Early Case Intelligence, edited for the current national practice and its documented engagement model. Methods are selected for the source, authorization, system state and assigned specialty. No single tool or artifact establishes a conclusion, and legal, regulatory or certification decisions remain with the responsible authority.
What This Solves
After collection, you are looking at several hundred thousand documents. Most of them are not relevant. A substantial portion are exact duplicates sitting in multiple custodians' inboxes. Hundreds of email threads are present in fragmented form, with the same conversation appearing across a dozen custodians at different stages of the reply chain. Review all of that document by document and you have weeks of attorney time spent on material the court will never see.
Early case intelligence (ECA) is the phase where GDF applies analytics and reduction techniques before a single document goes to an attorney for review. The goal is to move only the right data forward: unique, potentially responsive documents, in a form reviewers can navigate efficiently. The process is not about cutting corners. Material reduction decisions should be documented, methodology is logged, and the parameters used to exclude documents can be reproduced and explained. Modern ECA, done properly, reduces review volume by 60 to 80 percent without losing responsive material.
Deduplication: Global and Custodial
Hash-based deduplication is the foundation of volume reduction. Processing can compute cryptographic identifiers for collected items and should record the algorithms used. Two files with identical hash values are exact duplicates, byte for byte. One copy moves to review. The duplicate is suppressed but not destroyed: GDF maintains a cross-reference log mapping suppressed duplicates to their retained canonical copy, so that custodian-specific context is preserved if needed.
The choice between global and custodial deduplication matters, and it needs to be agreed upon with counsel before processing begins. Global deduplication removes a document from review if it appears anywhere in the collection, regardless of which custodian held it. A single copy proceeds to review. Custodial deduplication removes duplicates only within each custodian's own data set. If the same document appears in five custodians' email, all five copies proceed to review (one per custodian), because the fact that each person held the document may itself be relevant.
GDF documents which deduplication methodology was applied, the parameters used, the hash algorithm, and the resulting document counts at each stage. This documentation goes into the processing log and is available for opposing counsel review or court submission if challenged.
Email Threading and Inclusive Email Analysis
An email thread starts as a single message. By the time discovery is collected, that same conversation may exist as dozens of individual messages, each held by a different custodian at a different point in the reply chain. Every message in a thread is technically a separate document, but reviewing each separately is redundant: later messages in the thread quote all prior content.
Email threading identifies the most inclusive email in each conversation thread. The inclusive email is the latest message in the chain that contains all prior content within it. If a reviewer reads the inclusive email, they have seen everything said in that thread. Earlier messages (non-inclusive) are flagged as thread members but can be set aside unless a specific review decision requires looking at earlier branch points separately.
Threading also handles replies, forwards, and divergent branches. GDF's threading analysis identifies where a thread splits (when someone forwards a message and starts a parallel conversation) and ensures that each branch has its own inclusive email identified. The result is a structured view of the processed conversations, not a flat list of individual messages.
Near-Duplicate Detection
Exact hash deduplication removes byte-identical files. Near-duplicate detection handles the next layer: documents that are substantially similar but not identical. A contract with a single clause changed, a draft email with one paragraph removed, or a report with an updated figure, each is a different document by hash but may require only a fraction of the review time if identified as a near-duplicate of something already reviewed.
GDF's near-duplicate analysis computes similarity scores between documents and groups near-duplicates into clusters. Reviewers can see the pivot document for each cluster, review the differences within the cluster, and apply consistent review decisions across clustered documents without reading each one independently. The similarity threshold used and the resulting clusters are documented in the processing log.
Communications Mapping and Data Visualization
Communications mapping analyzes the email and message metadata across the entire collection to build a structured picture of who communicated with whom, how frequently, and during which time periods. This analysis serves two purposes. First, it helps counsel identify the most active communicators in the relevant time period, which often focuses document review on the custodians and relationships that matter most to the theory of the case. Second, it identifies unexpected communications: messages between parties who should not have been in contact, or communication patterns that are anomalous relative to the surrounding time period.
GDF delivers communications maps as visual outputs: network graphs showing communication volume and frequency between custodians, timeline charts showing communication spikes around key events, and frequency tables exportable to Excel or PDF. These visualizations support attorney strategy sessions and can be adapted for presentation to clients or co-counsel.