Skip to content
Daily AI Intel

AI in Law & Legal Services · AI in E-Discovery & Document Review

How does predictive coding work in document review?

Predictive coding trains a machine learning model on human relevance decisions for a sample of documents, then applies that model to rank or classify the rest of a much larger document set.

Key takeaways

  • Predictive coding is a machine learning technique that learns patterns from a set of human-coded example documents to predict relevance in a larger set.
  • The process typically involves iterative training, where reviewers continue coding documents and the model's predictions are refined over multiple rounds.
  • Statistical sampling is generally used to validate how accurately the model's predictions match human judgment before relying on it broadly.
  • Predictive coding is one specific technique underlying the broader technology-assisted review process used in e-discovery.

The core idea behind predictive coding

Predictive coding refers to the machine learning method that powers much of technology-assisted review in e-discovery. Rather than requiring reviewers to individually read every document in a large collection, predictive coding works by learning from a smaller set of examples that have already been reviewed and tagged by humans, then extending that learned pattern across a much larger set of unreviewed documents to predict which ones are likely relevant.

The iterative training process

A typical predictive coding workflow begins with reviewers coding an initial batch of sample documents as relevant or not relevant to the matter. The model is trained on those decisions and then used to generate predictions across the remaining document population. Rather than stopping after one round, most implementations continue this process iteratively — reviewers examine additional documents, often ones the model is uncertain about, and their coding decisions further refine the model’s predictions. This cycle typically continues until the model’s performance is judged sufficiently reliable through statistical testing.

Why validation matters as much as the model itself

Because the entire point of predictive coding is to reduce how many documents need full manual review, confidence in the model’s accuracy is essential before relying on it to any significant degree. This is typically established through statistical sampling — reviewing a random sample of documents the model did not use in training and checking whether its predictions match what a human reviewer would conclude. This validation step is often as important to defending the process in litigation as the underlying technology itself, since opposing parties or courts may scrutinize whether the review process was conducted reliably.

Bottom line

Predictive coding uses machine learning trained on a sample of human-reviewed documents to predict relevance across a much larger discovery collection, refined through iterative rounds of human input and validated with statistical sampling before broader reliance.

Go deeper

Important caveats

  • This describes the general concept of predictive coding rather than any specific vendor's implementation.

Frequently asked questions

How much human review does predictive coding actually eliminate?

It doesn't eliminate human review entirely — it's designed to reduce the volume of documents needing full manual review by prioritizing and classifying based on learned patterns, with humans still training and validating the model.

Does predictive coding get more accurate the more documents are reviewed?

Generally, yes — most implementations improve with more training rounds, as the model continues to learn from additional human coding decisions.

Is predictive coding accepted by courts?

Courts in a number of jurisdictions have accepted predictive coding and technology-assisted review as appropriate discovery methods when properly validated.

Sources

  1. [1]E-discovery standards and practice resources — American Bar Association
  2. [2]Federal court rules and discovery guidance — United States Courts
ET

Written by Editorial Team

Last updated July 28, 2026

Get one well-sourced answer a week

No spam. Unsubscribe anytime.