Search terms filter a collection by keyword before review. Predictive coding ranks the whole collection by likely relevance and lets review stop when recall is defensible. Terms are simpler and easier to negotiate; predictive coding measures its own performance, which is what makes it defensible at volume.
| Dimension | Search termsBoolean keyword filtering applied before review. | Predictive codingMachine learning ranks the population by likely relevance. |
|---|---|---|
| How scope is set | Negotiated keyword list, usually with hit counts exchanged. | Attorney coding decisions train a model applied to everything. |
| What gets missed | Anything phrased differently — code names, euphemism, misspellings, discussion by implication. | Material unlike anything coded so far; concepts absent from the training decisions. |
| Measurability | None inherent. Hit counts show volume, not recall. | Recall estimated from a sample of what was excluded, with a confidence interval. |
| Cost profile | Low setup, cost scales with the number of documents the terms return. | Higher setup, cost scales with how far down the ranking review must go. |
| Suits which data | Any format; works on short documents and chat where models struggle. | Text-rich documents; poor on spreadsheets, images and very short messages. |
| Negotiation burden | High and recurring — term lists are argued over, revised, argued again. | Front-loaded into one methodology and validation agreement. |
| Defensibility if challenged | Rests on the reasonableness of the list, which is a matter of assertion. | Rests on a measured recall estimate, which is a matter of evidence. |
Choose Search terms when
Choose search terms for smaller collections where the setup cost of a model is not recoverable, for populations that are mostly non-text, and for matters where the relevant vocabulary is genuinely well defined — a named project, specific contract numbers, identified counterparties. They are also the pragmatic choice where the opposing party will not engage on methodology and the volume does not justify the fight.
Choose Predictive coding when
Choose predictive coding when the population is large, text-rich, and the relevant material is a small fraction of it — which describes most modern email collections. Choose it also when defensibility matters more than simplicity, because it is the only approach that can produce a measured estimate of what the process missed rather than an argument that the process was reasonable.
Where this goes wrong
The failure that costs most is treating search terms as though they measure anything. A negotiated list returning 200,000 documents tells you how many documents matched, not what proportion of the responsive material you found — and the answer is frequently much lower than anyone assumes, because people discussing something sensitive rarely use the obvious word for it. Parties then produce from that set, certify completeness, and are exposed when the requesting party finds responsive material that no term would ever have hit. The reverse failure is reaching for predictive coding on a 5,000-document collection, where validation overhead exceeds the review it saves.
They are not actually mutually exclusive
The framing as a choice is somewhat artificial, and the most common real-world design uses both: culling by date range and custodian first, applying broad terms to remove obviously irrelevant categories, then ranking what remains.
What matters is understanding which mechanism is doing the load-bearing work, because that determines what has to be defended. A process that culls aggressively by keyword and then ranks the remainder has already made its most consequential decision at the keyword stage — and validating the ranking says nothing about what the culling discarded.
This is the subtlety that gets missed. Validation measures the step it was applied to, not the pipeline.
Why keyword recall is worse than intuition suggests
The foundational research on this is decades old and has been replicated repeatedly: experienced litigators asked to construct search terms substantially overestimate how much of the responsive material their terms will retrieve. The gap is large and consistent.
The reason is not incompetence. It is that people do not write about sensitive subjects using the terms an outsider would search for. They use first names, internal project code names, abbreviations, and — most defeating of all — reference the matter obliquely because everyone in the thread already knows what is being discussed.
What to agree before starting either
- For terms: whether hit counts are exchanged before the list is finalised, and what happens when a term returns an unworkable volume. A list agreed without hit counts is a list nobody has costed.
- For predictive coding: the validation protocol — target recall, sample size, who reviews the sample — and the stopping rule. Agreeing a stopping rule after review has begun is agreeing it at the moment one side wants to stop.
- For both: what is excluded from the process entirely and handled separately. Spreadsheets, images, foreign-language material and short chat messages are the usual candidates, and leaving them unaddressed is how they end up unreviewed.
The disclosure question
Most ESI protocols now require disclosing that technology-assisted review is being used. Fewer require exposing training decisions, and demanding them is usually a fight not worth having — the early case law that generated seed-set disputes largely predates continuous active learning, where there is no discrete seed set to argue about.
The productive demand is not for the training data. It is for the validation results.
From our work
Which one does your matter need?
Our examiners and testifying experts work these questions for a living. Tell us what you're facing.
Reviewed by Law & Forensics. See our editorial standards.
