In May 2023, a lawyer submitted a brief in a US federal court that cited six cases in support of a key argument. The opposing counsel filed a motion noting they could not locate any of the cases. The judge ordered the lawyer to produce the decisions. He could not. Because they did not exist.
The lawyer had used ChatGPT to help with research. The AI had generated six plausible-sounding citations — with realistic case names, realistic reporter references, realistic holdings — that were entirely fabricated. The lawyer faced sanctions. The case received international press coverage. And within weeks, courts around the English-speaking world had begun issuing AI disclosure requirements for legal filings.
This is not an old story. Variations of it have happened in the UK, Canada, Australia, and Hong Kong. They continue to happen. And the reason they keep happening is that the lawyers involved trusted AI output that looked authoritative without verifying whether it was real.
Legal professionals are uniquely positioned to both prevent this problem and earn from solving it. Here is why.
Why Legal AI Hallucinations Are Different from General AI Hallucinations
Every AI system hallucinates — generates confident-sounding content that is factually wrong. In most contexts, this is inconvenient. In legal contexts, it is professionally catastrophic.
The reason legal AI hallucinations are particularly dangerous is that they exploit the authority structure of legal citation. When a lawyer reads a case citation, they are conditioned by training and professional habit to extend a degree of trust to it. The citation format itself — case name, reporter, court, year — is a trust signal. It looks like verified information.
AI systems have learned to reproduce this format perfectly. A fabricated citation and a real citation are visually indistinguishable. The fabricated one sounds like the real one. It is formatted like the real one. The only way to know the difference is to check — and that checking requires someone who knows how to use legal research databases.
| “The AI did not fabricate something that looked obviously wrong. It fabricated something that looked exactly right. That is the entire problem.” |
The Market This Creates for Legal Professionals
Every legal AI product that touches case law, statutory analysis, contract drafting, or legal research has the same fundamental problem: it needs expert human review before its outputs can be trusted for professional use.
Law firms deploying AI research tools need someone who can verify citations before they reach a filing. Legal tech companies building AI contract review tools need someone who can assess whether the tool is correctly identifying jurisdictional issues. Consumer legal AI platforms need someone who can evaluate whether their chatbot is giving legal information or crossing into regulated legal advice.
None of this can be done by a general AI evaluator. All of it can be done by a lawyer, paralegal, law student, or legal professional with structured evaluation training.
The platforms that run AI evaluation projects have recognised this. Outlier AI, Scale AI, and DataAnnotation all run legal evaluation projects that specifically require JD-holders or law students. The rates for legal evaluation work reflect the scarcity of qualified reviewers — typically $35–$90+ per hour depending on the task and the evaluator’s qualification level.
What Legal AI Evaluation Actually Involves
The core skill in legal AI evaluation is citation verification — but the work goes considerably beyond checking whether a case exists. Here is a realistic breakdown of what legal AI evaluators do:
Citation verification
The five-step process: record the citation exactly as given, search the relevant legal database (Google Scholar, BAILII, CanLII, AustLII — depending on jurisdiction), confirm the case exists, read the headnotes to confirm the holding matches what the AI claimed, and check whether the case has been overruled or distinguished. Cases that fail any of these steps are flagged with a written analysis.
Holding accuracy assessment
A citation can be real while the holding attributed to it is wrong, incomplete, or inverted. The AI cites the case correctly but claims it stands for a proposition it actually rejected. Catching this requires reading the decision — not just confirming the citation exists.
Jurisdictional accuracy
Legal rules vary by jurisdiction in ways that fundamentally change the correct answer to a legal question. An AI that answers ‘what is the law on X’ without specifying a jurisdiction — or that applies one jurisdiction’s law to a different jurisdiction’s context — is producing output that could be actively misleading to someone relying on it. Evaluating jurisdictional accuracy is one of the most commercially valuable skills in legal AI evaluation.
Unauthorized practice of law assessment
Consumer-facing legal AI tools walk a fine line between providing legal information and providing legal advice. Crossing that line may constitute unauthorized practice of law in many jurisdictions. Legal professionals are the only population who can reliably identify where that line is and whether an AI output has crossed it.
The Mata v. Avianca Case — and What Came After
The Mata v. Avianca case is the most documented example of legal AI hallucination reaching a court sanction, but it is not the only one. What it illustrates is both the scale of the problem and the professional consequences.
The sanctioned lawyer, Steven Schwartz, told the court he had been unaware that ChatGPT “could fabricate cases.” This is now a response that courts are unlikely to accept. Bar associations across the US, UK, and Commonwealth have subsequently published guidance making clear that the duty of competence extends to understanding the limitations of AI tools being used in legal work.
For legal AI evaluators, this case is a permanent reference point. It is the clearest possible illustration of why legal AI needs expert human review — and why that review is worth paying for.
Building the Skills That Legal AI Companies Pay For
The gap between a legal professional who earns competitive rates for AI evaluation work and one who struggles to qualify for tasks is usually not legal knowledge. It is evaluation methodology.
Specifically, it is the ability to document findings professionally. An evaluator who flags a hallucination by writing ‘this case doesn’t seem to exist’ is less valuable than one who writes: ‘Citation Smith v. Jones [2021] EWCA Civ 114 does not appear in BAILII or Westlaw. A search for the parties in the Court of Appeal for the period 2019–2023 returns no matching result. The holding attributed to this case — that constructive dismissal requires a fundamental breach of contract — is accurate as a statement of law under Western Excavating v Sharp [1978] QB 761, but the citation is fabricated. Score: Citation Accuracy 1/5.’
The difference in those two responses is worth real money. The second one is what legal AI companies are looking for — and it is a learnable skill.
A Note on the Long-Term Outlook
Some legal professionals ask whether AI will eventually replace legal AI evaluators the same way it might replace other roles. The honest answer is that basic legal information review tasks are already being partially automated. But the work that requires genuine legal judgment — citation verification in complex cases, jurisdictional accuracy assessment, UPL risk evaluation for consumer products, contract review quality assurance — is precisely the work that AI systems cannot reliably perform on themselves. The whole point of human evaluation in these contexts is to catch what the AI missed.
Legal expertise in AI evaluation is a durable specialisation precisely because the tasks that require it are the ones where AI self-evaluation is most likely to fail.

Crested Academy offers professional AI evaluation training for domain experts. The Legal AI Evaluation track is designed for qualified legal professionals and law students, and is a prerequisite-gated domain specialisation requiring completion of the AI Foundations track first.