| TL;DR: AI can help legal teams work through thousands of documents much faster. But before the system starts, someone has to decide what it should look for. We spoke with Law In Order about what happens when those early decisions start carrying much more weight. |
Give a lawyer 10,000 emails to review and the limitation is obvious: one person can only read so much.
Give the same collection to AI, and that limitation starts to disappear.
But the point of review is not simply to read everything. It is to work out what matters to the case.
Before the system can decide which documents are relevant, someone has to decide what relevant means.
AI can apply that decision across the entire collection. That is what makes it useful. But if the decision is wrong, the error can be repeated just as widely.
Yvonne Ling of Law In Order, an Australian company that has helped legal teams manage and review digital evidence for more than two decades, says the questions clients are asking about AI have changed.
“Clients are no longer asking whether they should use AI,” Ling told SaaSTake. “They are asking how to use it responsibly, effectively, and in ways that deliver measurable results.”
Once AI is part of the actual workflow, what it can do stops being the only question. Teams have to decide where to trust it, how to check it, and where a person still has to make the call.
We asked Law In Order about that, and Ling’s reply explained why:
“If the relevance criteria or issue definitions are loose, ambiguous, or built on an incomplete understanding of the matter, the system will apply that flawed reasoning consistently across the entire population.”
The AI still needs a brief
A legal review is not a keyword hunt.
An email can contain every obvious term in a case and still tell you very little. Another can look unremarkable until you understand who sent it, what was happening at the time, and how it connects to what the legal team is trying to establish.
AI needs some version of that context too.
Law In Order’s explanation of AI-assisted legal review describes a system being given an overview of the case, key people, issues, and coding instructions. In simple terms, someone has to brief the AI before asking it to work through the documents.
Which raises an obvious question: who writes the brief?
Law In Order told SaaSTake that defining those criteria is legal work. It requires an understanding of what each side is arguing, the facts and strategy behind the case, and what the review is actually trying to establish.
That is also why the company does not describe AI as removing judgment from the process.

Legal teams were already using technology-assisted review, email threading, clustering, and other tools long before generative AI arrived. Law In Order notes that these technologies have been part of document review for more than a decade. So this is not a story about lawyers suddenly handing a job they once did entirely by hand to a machine.
What is changing is how much the system can do once it has been told what to look for.
The scale can be striking. Law In Order uses legal technology platforms including Relativity in its review work. In one case study published by the company, 29,000 documents arrived weeks behind schedule. After Law In Order narrowed the collection, aiR for Case Strategy, Relativity’s generative AI tool for analyzing case documents, analyzed nearly 10,000 documents in about three hours and extracted 6,000 facts. According to Relativity, the wider review was completed in eight hours rather than the weeks it would ordinarily have taken if done manually.
That tool is a tad different from their product Relativity’s aiR for Review, which can classify documents for things such as relevance.
Still, the example shows what happens when analysis becomes much easier to scale. The person setting the direction may make fewer calls document by document. But the choices they do make can reach much further. And that creates another problem.
How do you know the AI followed those instructions without asking people to check everything themselves?
You can’t check the AI by doing its job again
There is one foolproof way to check 10,000 AI-reviewed documents: have people reread all 10,000.
There is also an obvious problem with that plan.
So legal teams test the review instead.
This is where terms such as control sets, sampling, precision, and recall start appearing. Strip away the terminology, and the questions are easier to understand: Did the system find what it was supposed to find? And how much of what it found was actually relevant?
Law In Order’s published hybrid-review workflows show one way teams can answer those questions. Human reviewers can code examples, compare AI predictions with their own decisions, look for incorrect classifications and missed documents, refine the instructions, and test again.
The company put the principle more simply in its response to SaaSTake:

Validation can go deeper than checking whether the final classification looks right. Law In Order told SaaSTake that current tools can also show the model’s reasoning behind a classification, which creates another task for the reviewer: deciding whether that reasoning itself makes sense.
“Current tools give you the model’s reasoning for each classification, and someone with legal training has to assess whether that reasoning is sound, not just whether the outcome looks right,” Ling said.
A classification can look right even when the reasoning behind it is not. And Law In Order is not alone in asking what should happen between an AI producing an answer and a legal team relying on it.
Add to it the fact that an American Bar Association examination of generative AI in discovery review points to a particularly awkward problem. An AI-generated result can look convincing even when material that should have been considered never appeared in the result at all.
That is harder to catch than an obviously bad answer. You cannot inspect a document the system never surfaced.
And that exposes the harder side of validation. Checking what the AI surfaced can tell you whether some of its calls were right. It does not, by itself, tell you what the system failed to surface in the first place. That is why testing recall, using control sets, and sampling across the wider collection matter. The review has to look for what the system may have missed, not only whether the documents it surfaced look convincing.
The Federal Court of Australia’s guidance on generative AI arrives at the issue from another direction. The Court recognizes the technology’s potential to improve efficiency, lower legal costs, and improve access to justice, while making clear that people using it remain responsible for understanding its limitations and complying with their existing professional obligations.
So the faster the review becomes, the less useful “the AI did it” becomes as an explanation.
Someone still has to be able to show why the process deserves to be trusted.
But testing and sampling only tell you how the system performs on average. They can’t hand it context that was never written down in the first place, and for some decisions, that missing context is the whole problem.
Then there are the calls that don’t fit neatly on the page
Privilege is a good example, and it’s worth explaining what that actually means. It’s the rule that keeps certain documents, mainly private conversations between a client and their lawyer, out of the other side’s hands during a case.
You see, just because a lawyer’s name shows up on an email doesn’t mean that email is protected. You’d need to know why the lawyer was involved, whether legal advice was actually being sought, who else was copied, and what happened afterward. Sometimes understanding a document means understanding things that aren’t written anywhere in it.
AI can help here too. Relativity has a separate aiR for Privilege tool built for exactly this kind of review, and an American Bar Association analysis of AI in e-discovery lays out both sides of it: AI can speed up privilege classification, but false positives, false negatives, and subtle context still mean a person needs to check the work.
Ling walked us through a short list when we asked where human judgment matters most. Getting the relevance criteria right came first. Testing the results came second. Then she got to the one she felt most strongly about:
“Privilege is the third, and it is the one we would not hand over.”
That’s not a claim that AI has no place in privilege review; the tools clearly exist and clearly help. It’s a claim about something narrower: being able to assist with a decision isn’t the same as being trusted to own it.
There’s also waiver to think about: whether the protection was given up somewhere along the way, which isn’t something you can tell from the document either. Getting any of these calls wrong can have consequences that are hard to undo once material has already reached the other side.
Law In Order says the same problem can show up with relevance too. One document may be peripheral in one dispute and central in another, depending on the allegation, the facts, and what the legal team is trying to establish.
So the boundary is messier than “AI handles easy work and people handle difficult work.”
Sometimes the technology can make the call. Sometimes it can help make the call. And sometimes the person overseeing the review needs to know why those two situations are not the same.
That changes what the reviewer is there to do.
The job is moving away from the documents
One line in Law In Order’s response captures that shift better than any list of AI capabilities could:
“The work has shifted from reading documents to designing, testing and defending the process that reads them.”
That does not mean people stop reading documents. Nor does it mean every review will suddenly be run by a handful of senior experts.
What changes is where human expertise matters most.
If a system can make thousands of classifications, the person overseeing it does not create value by trying to outrun the machine at reading. The harder work sits around the system: deciding what it should look for, checking what it did, recognizing when context changes the answer, and being able to defend the process afterward.
Experienced people can spend less time on repetitive review and more time on decisions that actually require their experience.
Law In Order has made that argument itself before. In an earlier piece on combining technology and people in document review, the company argued that technology can free junior resources from lower-level work that may otherwise stifle their development.
But removing lower-level work raises another question: what happens to the experience people used to gain while doing it?
If the important decisions are increasingly concentrated in “fewer, more experienced hands,” where do the next experienced hands come from?
Experience was never produced by a job title
The obvious answer would be to keep junior lawyers doing routine work because that is how things have always been done.
That is not a particularly convincing answer.
Repetitive work is not automatically good training. Making somebody spend hours on a task that software can perform more efficiently does not become useful simply because an earlier generation had to do it.
But people did learn while doing some of that work.
They saw patterns. They made mistakes. Someone corrected them. Gradually, they learned which detail was noise and which detail changed the interpretation of everything around it.
As AI removes or compresses parts of that work, legal firms have to think more deliberately about what replaces that learning.
An April 2026 Thomson Reuters Institute analysis of lawyer development in AI-enabled firms describes firms working on exactly this problem. One firm cited in the piece is trying to shorten the learning curve for legal judgment and plans more formal training around supervising and validating AI output.
That gives Law In Order’s observation a consequence beyond Law In Order. As more of the routine work is compressed, the expertise needed to supervise it becomes more important, not less.
The problem, then, is not preserving every task AI can take away. It is making sure the learning hidden inside some of those tasks does not disappear with them.
A machine can work through another thousand documents. But somebody still has to know when the instruction is wrong, recognize the context the document does not explain, and decide when a call that can be automated still should not be handed over.
AI is making legal review easier to scale. The harder question may be how legal teams scale the judgment behind it.




