The GenAI Verification Burden in Legal Practice - Image shows a lawyer in a formal suit surrounded by stacks of legal textbooks and open papers, carefully reviewing Generative AI output displayed on a laptop screen. The setting includes a traditional law office with shelves of legal volumes and a scale of justice in the background, creating a professional and authoritative atmosphere.
|

The GenAI Verification Burden in Legal Practice

The GenAI verification burden in legal practice has been highlighted by Joshua Yuvaraj in a recent paper on ‘The Verification-Value Paradox: A Normative Critique of Gen AI in Legal Practice‘ [PDF].

The Abstract reads:

It is often claimed that machine learning-based generative AI products will drastically streamline and reduce the cost of legal practice. This enthusiasm assumes lawyers can effectively manage AI’s risks. Cases in Australia and elsewhere in which lawyers have been reprimanded for submitting inaccurate AI-generated content to courts suggest this paradigm must be revisited. This paper argues that a new paradigm is needed to evaluate AI use in practice, given (a) AI’s disconnection from reality and its lack of transparency, and (b) lawyers’ paramount duties like honesty, integrity, and not to mislead the court. It presents an alternative model of AI use in practice that more holistically reflects these features (the verification-value paradox). That paradox suggests increases in efficiency from AI use in legal practice will be met by a correspondingly greater imperative to manually verify any outputs of that use, rendering the net value of AI use often negligible to lawyers. The paper then sets out the paradox’s implications for legal practice and legal education, including for AI use but also the values that the paradox suggests should undergird legal practice: fidelity to the truth and civic responsibility.

This was met on LinkedIn with, on the whole, acknowledgement and support for this proposition.

However, Antti Innanen did not see it that way. He commented:

It would make a great paradox if it were true.

The supposed paradox – that increases in efficiency from AI use in legal practice will be met by a correspondingly greater need to manually verify outputs, making the net value of AI negligible – is not supported by any evidence.

The paper offers nothing beyond a few anecdotes and old, tired GPT-3.5 hallucination rates.

Meanwhile, massive evidence points in the opposite direction.

It might sound like a responsible thing to say, and it might even be fun to debate, but the data don’t support it.

The claim belongs in the same category as Brian Inkster’s original blog post: funny but not true.

I assume the reference to my original blog post is to Inkster’s Law:

The 2:1 ratio of time needed to check for accuracy of AI generated material

's Law - AI Content Review Principle - Image of the scales of justice showing the balance between GenAI content and the verification burden of checking it.

Although long before that I was advocating the need for The Legal Hallucinatory Detectorist over Legal Prompt Engineers.

I responded to Antti Innanen:

Please do provide to us citations for the massive evidence that points in the opposite direction.

He replied:

To counter a claim “increases in efficiency from AI use in legal practice will be met by a correspondingly greater need to manually verify outputs, making the net value of AI negligible”?

If someone argues that the “net value of AI is negligible in legal practice”, the burden of proof lies entirely on them.

Outrageous claims require evidence, not the other way around.

There’s already substantial proof that AI delivers meaningful efficiency gains when applied thoughtfully. Too tired to post the obvious here.

I countered:

So no “massive evidence” after all. As I thought! The “outrageous claims” concern AI providing huge efficiency gains for lawyers. As you say, this requires evidence. I will always be interested to see such evidence if anyone can actually produce it.

Antti Innanen responded:

Ok, I will bite.

PERFORMANCE:

Vals AI (October 2025) benchmark shows legal AI outperforming human lawyers by 9 percentage points on core legal tasks.

OpenAI’s GDPVal (September 2025) evaluated AI on authentic professional tasks across 44 occupations, including legal work, with Claude achieving near-human baseline performance while completing work much faster.

Then you have the Stanford LegalBench evaluations.

One of my favorite evals is Legalbenchmarks, by Anna Guo. The new study is a gem.

STUDIES:

“Better Call GPT” showed AI surpassing junior lawyers on contract analysis while operating at much lower cost.

“Building a Better Lawyer” demonstrated task completion time reduction with maintained quality.

“Generative AI and Legal Aid” field study showed that legal professionals reported increased productivity.

“AI-Powered Lawyering” showed that o1-preview increased productivity by a significant number.

Oldie but goldie (and non-legal specific): Harvard/BCG’s “Navigating the Jagged Technological Frontier” established 12.2% more tasks completed and 40% higher quality outputs with AI assistance.

ADOPTION:

LegalBenchmarks 2025: 97% of lawyers use AI tools for legal work. 83% actively use two or more platforms.

EFFICIENCY GAINS:

Thomson Reuters surveyed 1,700+ professionals (lawyers included) and found adoption is driven by time-savings.

Everlaw (vendor) reports legal teams saving up to ~32.5 working days per lawyer annually after deploying generative AI in review and drafting workflows.

Juros “State of In-house 2025”: 93% of CEOs & CFOs want their legal teams to increase AI adoption; 99% believe AI will change their legal roles within a year.

Clio 2025 Legal Trends Report says most of those using AI are seeing improvements in the quality of their work, responsiveness to clients, and work capacity.

All the studies that I mentioned point to efficiency gains.

Then there are countless papers from consulting firms:

BCG, McKinsey, Bain, Thomson Reuters’ 2025 Future of Professionals report, Deloitte’s survey of Chief Legal Officers, Forrester’s Economic Impact studies.

MARKET VALIDATION:

Harvey and Legora just both raised $150M with valuations at $8B and $1.8B.

These are just from the top of my head. I’m surely forgetting many more.

On the other side, Brian Inkster, we have your blog post (funny, but untrue) and this study that says:

“However, the actual paradox – based on the above analysis of the verification cost – is fundamentally theoretical, even though anecdotal evidence seems to bear it out. The paradox is therefore a hypothesis for further examination.”

Like your post, it’s an interesting way to point out that verification matters when using AI tools, and that verification costs can be high in legal work. As they should be, this is serious work!

But if you’re truly arguing that “the net value of AI use often negligible to lawyers”, you need to come up with evidence.

I just presented mine and it’s…MASSIVE.

Now it’s your turn to present yours.

But I suspect we’ll get 🦗 🏏 because there isn’t any.

Let me repeat it: there is no evidence that “increases in efficiency from AI use in legal practice will be met by a correspondingly greater imperative to manually verify any outputs of that use, rendering the net value of AI use often negligible to lawyers”.

It’s just a fun thing to say, not a substantiated claim.

Later in the thread Antti states:

Fermi style: where are all the lawyers manually verifying AI outputs all day. Where are they?

I replied to Antti’s “massive” evidence:

That’s not “massive”! It is predictable GenAI hype that does not actually address the point in hand. You won’t get 🦗 🏏 from me. You should know that by now 😉. But having not long got off a 15 hour flight with an 8 hour time difference you will have to wait until I get over the jet lag for a full and detailed response, which might warrant a complete blog post!

I am now over the jet lag and able to do what did, indeed, warrant the complete blog post that you are now reading.

Given how much faith Antti puts in GenAI, I asked Copilot:

Is there evidence that for legal related matters the time spent verifying GenAI output can outweigh the benefits of using it?

Copilot answered:

There is evidence suggesting that the time spent verifying Generative AI (GenAI) outputs can indeed outweigh the benefits, particularly in legal contexts where accuracy and authenticity are paramount.

Challenges of AI-Generated Evidence

The increasing use of AI-generated evidence in legal proceedings presents significant challenges. Judges and attorneys must determine the authenticity, validity, and reliability of digital content that may have been artificially created or manipulated. The rapid advancement of GenAI technology has outpaced the development of reliable detection tools, making it difficult to distinguish between genuine and AI-generated evidence.

Verification Time and Legal Risks

Increased Verification Time: Legal professionals often find themselves spending considerable time verifying the outputs generated by GenAI tools. This is particularly true when the outputs are used as evidence in court. The need for rigorous testing and validation of AI-generated content can lead to delays in legal processes, potentially outweighing the initial time savings offered by using GenAI.

Case Examples: There have been notable cases where attorneys faced sanctions for submitting legal documents that included fictitious citations generated by AI tools. For instance, a New York attorney was fined for failing to verify the authenticity of AI-generated citations, highlighting the professional risks associated with unverified reliance on GenAI outputs. Such incidents underscore the necessity for attorneys to exercise due diligence and maintain human oversight when incorporating AI into their practice.

Conclusion

In summary, while GenAI can enhance efficiency in legal research and documentation, the time and effort required to verify its outputs can be substantial. This verification process is crucial to uphold the integrity of legal proceedings and avoid potential sanctions or ethical violations. Therefore, legal professionals must weigh the benefits of using GenAI against the time spent ensuring the accuracy and reliability of its outputs.

Sources given by Copilot were:

It would be rather ironic if Antti now seeks to debunk what GenAI says on the matter!

Let me now also go into some detail by breaking down, and critically examining, Antti Innanen’s “massive” evidence.

PERFORMANCE:

  • Vals AI (October 2025) benchmark shows legal AI outperforming human lawyers by 9 percentage points on core legal tasks.

As reported by Bob Ambrogi, this particular benchmark also shows that:

Lawyers outperformed AI in roughly one-third of question categories, particularly those requiring deep interpretive analysis or nuanced reasoning, such as distinguishing similar precedents or reconciling conflicting authorities.

These areas underscore, as the report put it, “the enduring edge of human judgment in complex, multi-jurisdictional reasoning.”

It should also be noted that the benchmarking did not include the three largest AI legal research platforms: Thomson Reuters, LexisNexis and vLex.

  • OpenAI’s GDPVal (September 2025) evaluated AI on authentic professional tasks across 44 occupations, including legal work, with Claude achieving near-human baseline performance while completing work much faster.

However, with regard to legal work the report states:

The current version of the evaluation is also one-shot, so it doesn’t capture cases where a model would need to build context or improve through multiple drafts—for example, revising a legal brief after client feedback or iterating on a data analysis after spotting an anomaly. Additionally, in the real world, tasks aren’t always clearly defined with a prompt and reference files; for example, a lawyer might have to navigate ambiguity and talk to their client before deciding that creating a legal brief is the right approach to help them.

  • Then you have the Stanford LegalBench evaluations.

However, these evaluations have been criticised. It has been argued by Ryan McDonough that LegalBench focuses heavily on structured legal reasoning tasks (e.g., statutory interpretation, rule application) that may not reflect the messy, ambiguous, and pragmatic reasoning lawyers often perform in real-world settings.

This could lead to overestimating LLMs’ actual legal competence, since passing benchmark tasks doesn’t necessarily mean a model can handle real legal practice.

While LegalBench tasks are crafted by legal experts, Ryan McDonough has also noted that they are still synthetic and may not capture the complexity, nuance, or unpredictability of real legal documents or litigation scenarios.

For example, tasks often involve binary outputs (e.g., “Is this hearsay: Yes or No?”), which oversimplifies the interpretive nature of legal analysis.

  • One of my favorite evals is Legalbenchmarks, by Anna Guo. The new study is a gem.

Anna Guo has openly acknowledged that public benchmarks—even rigorous ones—can be manipulated. She cites examples like Meta allegedly gaming Hugging Face rankings to illustrate how leaderboard performance doesn’t always reflect real-world capability.

LegalBenchmarks are designed to test legal AI tools on structured tasks like contract drafting or statutory interpretation. However, Guo and others note that these tasks don’t always reflect how lawyers actually work—especially in messy, ambiguous, or collaborative environments.

This report was also aimed specifically at in-house Counsel and information extraction tasks.

STUDIES:

  • “Better Call GPT” showed AI surpassing junior lawyers on contract analysis while operating at much lower cost.

An article by Phuoc Nguyen suggested that Saul might disagree:

The comparison between the cost and time efficiency of LLMs and paralegals, while gleefully underscoring the supreme efficiency of LLMs over human lawyers (what a groundbreaking revelation, right?), conveniently skips over the crucial expense and necessity of having a lawyer double-check the LLM’s work. Unless, of course, you’re feeling particularly bold and ready to join the ranks of those pioneering souls who’ve flirted with professional jeopardy by putting too much faith in their digital counterparts, like that New York lawyer who got quite a stir for cozying up too closely with ChatGPT.

  • “Building a Better Lawyer” demonstrated task completion time reduction with maintained quality.

The study used law students to test AI-assisted legal task performance. Law students are, of course, not representative of practicing lawyers, especially in terms of experience, judgment, and workflow familiarity.

Lars Daniel, in Forbes article, pointed out:

Despite the positive results, the study identified several limitations. Neither AI tool consistently improved accuracy in legal research, with o1-preview showing a tendency to hallucinate sources in some cases.

Additionally, both tools were less effective for transactional work. The one assignment involving drafting a non-disclosure agreement showed no significant improvements in either quality or speed when using AI.

He went on to state:

While AI shows tremendous promise in improving legal work, human verification remains essential for final outputs to ensure no hallucinations or errors slip through. This balance between leveraging AI’s capabilities and maintaining human oversight will be key to harnessing the full potential of AI in law.

  • “Generative AI and Legal Aid” field study showed that legal professionals reported increased productivity.

The study involved 91 participants using paid generative AI tools over two months, with a subset receiving “concierge” support.

Colleen Chien and Miriam Kim acknowledge limitations in their own study. They note that:

  • The sample size was small and self-selected.
  • The study didn’t measure client outcomes or long-term impacts.
  • There were disparities in AI uptake across gender and roles

There is a worry that promoting AI tools in legal aid could lead to overreliance, especially among under-resourced practitioners who may lack time to verify outputs.

There is also, of course, concern that hallucinations or subtle errors in AI-generated legal advice could harm vulnerable clients if not caught and corrected.

  • “AI-Powered Lawyering” showed that o1-preview increased productivity by a significant number.

This was another study that used law students to test AI. As indicated previously, law students are, of course, not representative of practicing lawyers, especially in terms of experience, judgment, and workflow familiarity.

It was again a study that found the AI tools hallucinated.

The six legal tasks used in the study were controlled and standardised, which helps with measurement but may not reflect the messy, ambiguous, and context-rich nature of actual legal practice.

Real cases often involve incomplete information, conflicting precedents, and strategic judgment—factors that are hard to simulate.

  • Oldie but goldie (and non-legal specific): Harvard/BCG’s “Navigating the Jagged Technological Frontier” established 12.2% more tasks completed and 40% higher quality outputs with AI assistance.

As non-legal specific I will discount this entry from Antti as a no go rather than an oldie or a goldie. However, in any event, you may wish to read Conor E. Doherty’s take on why this study is no goldie but “deeply flawed and potentially dangerous“.

ADOPTION:

  • LegalBenchmarks 2025: 97% of lawyers use AI tools for legal work. 83% actively use two or more platforms.

If that AI tool is spell check in Word then this high percentage figure has no doubt been true for some time.

I asked Copilot if this figure was correct. It responded:

While the numbers are impressive, some experts have raised concerns about overgeneralization and the methodology behind the 97% figure. A critical review in the National Law Review suggests that claims of AI “universality” may be premature and that real-world usage varies by firm size, region, and practice area.

This review pointed out that the survey involved only 72 respondents and that quite possibly solicitors contacted who did not use AI simply opted out of taking part “for any number of reasons, including not wanting to appear out of date or due to embarrassment”.

I also asked Copilot for it’s thoughts on a more accurate percentage. It responded:

46% of lawyers now use AI tools, up from just 11% in 2023 — a 318% increase in under two years.

84% of legal professionals are either using or planning to adopt AI by the end of 2025.

55% of lawyers already use generative AI for tasks like research, drafting, and client communication.

61% of UK lawyers report using generative AI, though only 11% use it heavily for day-to-day work.

In-house legal teams lead adoption at 72%, compared to just 31% in the public sector.

Personal injury firms are the most aggressive adopters, with 56% automating tasks like medical record summarization.

Copilot gave its sources for this as aiqlabs and allaboutai.

Whilst not highlighted by Copilot the second of those two sources states that 21% of law firms use GenAI showing a distinction between law firm and individual use. This would then pose the question, is that individual use approved by the law firms that the individuals work for?

Some further (better perhaps) research via Google gave me from Legalit insider that ‘LexisNexis report: Over 60% of UK lawyers now ‘use GenAI’, but law firm culture slows progress‘. In particular this report shows that whilst 60% may say they are using GenAI only “11% of lawyers said they are using GenAI heavily for day-to-day work and that number drops to 7% at large and medium-sized firms”. Furthermore “the most common response in describing their organisation’s AI culture was ‘we’re experimenting but progress is slow.’ 19% reported interest but little investment.”

However, usage figures do not counter the GenAI verification burden. Many solicitors will be playing with GenAI rather than using it for real legal work. Many will, as they have done for many years, be using AI for spell check in Word or other such tasks that are taken for granted but are not GenAI use cases. Some will be using it when it is completely daft to do so. For example, I spotted just the other day this post on LinkedIn by Aaron Krigelski:

Someone told me they used AI to highlight every instance of a word in their document.

“It took some prompt engineering, but I got it to work!”

Ctrl+F. That’s it. That’s the tool.

Been around since the 1970s. Works every time. Takes zero seconds. No prompt engineering required.

But here we are, using artificial intelligence to do what a keyboard shortcut solved decades ago.

I see this constantly. Firms buying AI solutions for tasks their existing software already handles. Like buying a Ferrari to drive to your mailbox.

Last month, a legal assistant spent hours getting ChatGPT to count pages in PDFs. There’s literally software called PDF Page Counter. Does one thing. Does it perfectly.

The week before that? Someone used AI to rename files in sequence. Bulk Rename Utility. Free. Since 2000.

Here’s the thing: AI is incredible for certain tasks. I use it daily. But when you’re using a sledgehammer to push in a thumbtack, you’re not innovative. You’re inefficient.

The best tool isn’t always the newest tool. Sometimes it’s the boring one that’s worked perfectly for 20 years.

Know what problem you’re solving. Then pick the right tool. Not the trendy one.

Your Ctrl key is feeling neglected.

EFFICIENCY GAINS:

  • Thomson Reuters surveyed 1,700+ professionals (lawyers included) and found adoption is driven by time-savings.

The report also disclosed that:

However, even with usage skyrocketing, many of the firms, departments, and agencies in which these professionals work still have a way to go to fully extract value from GenAI. The report shows that few organizations are capturing return-on-investment (ROI) metrics, regularly training staff on GenAI updates, or integrating GenAI use into their policies. In addition, few professionals say their firms and their clients are having conversations around GenAI use. And even more worrisome, questions about GenAI’s impact on billing rates and costs remain unanswered.

  • Everlaw (vendor) reports legal teams saving up to ~32.5 working days per lawyer annually after deploying generative AI in review and drafting workflows.

This relates to ediscovery software which was, of course, doing this for lawyers (using AI) long before GenAI came along.

  • Juros “State of In-house 2025”: 93% of CEOs & CFOs want their legal teams to increase AI adoption; 99% believe AI will change their legal roles within a year.

This is rather than hire more staff. The CEOs & CFOs hope that AI can replace staff. The CEO of Klarna found out that was not the case:

Months after touting AI’s potential to replace human work, Klarna CEO Sebastian Siemiatkowski is backtracking and reversing an AI-induced hiring freeze to bring on more human staff.

Siemiatkowski, 43, told Bloomberg on Thursday that Klarna is hiring human workers again to ensure that customers always have a human presence to talk to, if needed.

“From a brand perspective, a company perspective, I just think it’s so critical that you are clear to your customer that there will always be a human if you want,” Siemiatkowski told the outlet.

Siemiatkowski tells Bloomberg that the AI-focused strategy Klarna employed for the past few years wasn’t the right path. He says that while AI customer service chatbots were cheaper to employ than human staff, they resulted in a “lower quality” output.

  • Clio 2025 Legal Trends Report says most of those using AI are seeing improvements in the quality of their work, responsiveness to clients, and work capacity.

The report also references the risks of AI hallucinations, particularly in the context of legal accuracy and professional responsibility.

Acknowledged Risk: The report notes that while AI tools offer significant productivity gains, they also pose risks – especially when they generate inaccurate or fabricated legal information (hallucinations).

Professional Implications: Lawyers are cautioned to maintain oversight and verify AI-generated outputs to avoid ethical breaches or reputational harm.

Training and Guardrails: The report emphasises the need for proper training and the implementation of safeguards to mitigate hallucination risks in legal workflows.

Contextual Framing: Hallucinations are discussed as part of a broader concern about cognitive load and decision-making quality – highlighting that AI should support, not replace, human judgment.

Thus the GenAI verification burden is acknowledged.

It was also noted recently that Kim Kardashian (who has been studying law) blamed ChatGPT for failing her law exams. So she didn’t see an improvement in the quality of her work by using it! This also shows that the verification burden is real if she was relying mainly on ChatGPT to get her through her exams without such verification.

Kim Kardashian blames CHatGPT for failing law exams

However, not all students are like Kim Kardashian. William Scates Frances, pointed out, just the other day, on LinkedIn that:

A significant number of university students hold a strong stance against generative AI and refuse its use outright or in part.

One of the five top reasons given for this is:

They don’t trust AI outputs. They view chatbots as unreliable and not worth the effort required to babysit them. It’s more trouble than just doing things the “old-fashioned” way.

Thus, the verification burden is being recognised in higher education.

  • Then there are countless papers from consulting firms: BCG, McKinsey, Bain, Thomson Reuters’ 2025 Future of Professionals report, Deloitte’s survey of Chief Legal Officers, Forrester’s Economic Impact studies.

Then there is the fact that Deloitte produced a $439,000 (AUS) report for the Australian Government riddled with GenAI hallucinated errors!

MARKET VALIDATION:

  • Harvey and Legora just both raised $150M with valuations at $8B and $1.8B.

Which does not validate in any way any argument that there is not a GenAI validation burden.

Furthermore, many (including Bill Gates) believe that such valuations are all pointing to a GenAI bubble that may soon burst.

Indeed, it has been suggested that OpenAI are looking for Government bailouts when the bubble bursts.

Just yesterday: Wall Street drops to one of its worst days since April as AI bubble fears return

LAWYERS MANUALLY VERIFYING?

  • Fermi style: where are all the lawyers manually verifying AI outputs all day. Where are they?

Well they should be at their desks doing so all day if they are in fact using GenAI. No doubt they aren’t broadcasting the fact to the world as they record their time privately on their time sheets.

If they are not we have a problem. That problem is borne out by the number of reported court cases where hallucinated citations have been put forward without verification (more on that below – under ‘Funny but Untrue?’)

However, for Antti’s information there is a GenAI legal technology company that is using lawyers to manually verify AI outputs all day. That is LawY:

Ask your legal questions and unlock matter-specific instant AI-answers, with optional verification by qualified lawyers. LawY blends cutting-edge legal AI with human expertise, to enhance your firm’s efficiency with confidence.

That shows recognition of the verification burden but is at the same time, in my opinion, quite ridiculous. It is a tool being sold to law firms who then use LawY lawyers (outside their own firms) to verify the GenAI output and then they should still, of course, do so themselves (as no PI insurance from LawY) thus doubling the verification burden!

Before I knew about LawY, the concept of it was an April Fool’s joke on this blog!: Elwood launches claiming to be Real AI for Law

Funny but Untrue?

Antti says that ‘Inkster’s Law’ is funny but untrue. However, he provides no real evidence to counter it. The foregoing “massive” evidence from him is merely pro GenAI sound bites (many from vendors incorporating GenAI into their products) that can easily be taken down. Indeed, many of them meanwhile actually acknowledge the problem of hallucinations and hence arguably back the principle of the GenAI verification burden.

Antti also considers Joshua Yuvaraj’s academic paper to be funny but untrue. However, again, he provides no real evidence to counter it. If anything he provides evidence in support of it, when he says:

verification matters when using AI tools, and that verification costs can be high in legal work. As they should be, this is serious work!

Yes, verification matters as can be seen by the numerous cases (545 identified so far by Damien Charlotin  who tracks AI Hallucination Cases) where GenAI hallucinated citations have been used in court without such verification. It should be noted that was 532 this time last week. So Damien has clocked 13 new ones just in the space of one week. These are in the public domain. It is frightening to think, on that basis, of what is being hallucinated behind closed doors in law offices around the world daily.

It should be clear to anyone who practices in court (I do and, I assume, Antti does not) that the time that would be spent verifying such erroneous output would outweigh (as Joshua Yuvaraj says) or far outweigh (as I say) any benefits of using GenAI for legal research in the first place.

I talk about this with Kayode Alabi in a recent interview. Although, at that time I had read of 122 reported cases (not the 545 reported, as of today’s date, by Damien Charlotin):

Guidance issued by The Bar Council (England & Wales) stresses the need for verification:

  • Due to possible hallucinations and biases, it is important for barristers to verify the output of LLM software and maintain proper procedures for checking generative outputs.
  • ‘Black box syndrome’ – LLMs should not be a substitute for the exercise of professional judgment, quality legal analysis and the expertise that clients, courts and society expect from barristers.

The verification burden is borne out by comments on Joshua Yuvaraj’s LinkedIn post, such as:

Haward Soper (Honorary Professor Of Law at University of Leicester):

Another must read! It reflects my data too. I have asked around ten AI tools to draft an exclusion of consequential loss for me. Redlining their work would have taken quite some time. This should be published later this year in my upcoming book on excluding consequential loss (which includes a note on the issue in Australia ).

Dr Eliza Mik (IT Lawyer, Academic):

100% true. LLMs accelerate text production- not reading / rewriting time.

Julian Webb (Professor of Law at The University of Melbourne):

Thanks for the HT Joshua Yuvaraj, I’m very much looking forward to reading this. From the summary, I think we will be in substantial agreement!

Jonathan Crass (Privacy, Cyber and AI Lawyer | CIPP/E):

On the reading list and thank you Joshua Yuvaraj for putting in the time to dig deeper.

Sounds like it confirms what I have found in practice and what other colleagues have told me they have found.

Marco Rizzi (Associate Professor, Graduate Research Coordinator and Co-Director of the Centre for Health Law & Policy at UWA Law School):

This aligns with my personal and very anecdotal experience. I have tried to use LLMs to generate scenarios, mostly for teaching purposes. The reality is that the amount of work required to tailor the output to the actual teaching needs eats up pretty much all the time ‘saved’ by outsourcing the initial thinking process. In fact it slows down the editing and re-writing process, because it is not text that I have generated in the first place, so I have less familiarity with it.

Alice Hewitt (Legal Information & Knowledge Management Champion | Discoverability & Accessibility Advocate | Information Literacy Tour Guide | Research Ninja | Interrogator of Mysteries (& databases & AI & IA & the why…):

A interesting read – and supports a lot of what I know many law librarians know (or at least strongly suspect).

Jane Cornwell (Senior Lecturer in Intellectual Property Law at University of Edinburgh Law School):

What a great a paper on generative AI, legal practice and legal education…

As someone who spent many years in legal practice – not only producing my own outputs in terms of advice, correspondence, pleadings and submissions, but also reviewing the work of others and signing off on that work in the name of the firm – I worry that misjudging the proper response to AI in legal education risks setting up students to fall short in the discharge of their professional obligations as they move forward in their careers. This paper’s conclusions really chimed with me – particularly the response to the argument that incorporating AI into legal education is essential because the use of AI will be all-pervasive.

I have previously written about the limitations/dangers of GenAI in legal practice, including the verification burden (although I hadn’t previously referred to it as that):

The GenAI verification burden is also borne out by others writing on the topic:

The Siren’s Song of GenAI: Why legal practitioners still fall for fabricated content‘ – Armin Almardani refers to the phenomenon of ‘verification drift’. He says that “this occurs when users initially approach AI-generated content cautiously, aware of the risks of inaccuracy and the potential for hallucination. However, as they engage with the content, they gradually become overconfident in its reliability and find verification less necessary. This misplaced trust may stem from GenAI’s authoritative tone and ability to present incorrect details alongside accurate, well-articulated data. This suggests that the challenge is not just a lack of awareness but a cognitive bias that lulls users into a false sense of security.” He also refers to the ‘verification burden’ and states that “for certain tasks, relying on generative AI followed by exhaustive verification may be less time-efficient than conducting traditional legal research.”

AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries‘ – A study reveals that “bespoke legal AI tools still hallucinate an alarming amount of the time: the Lexis+ AI and Ask Practical Law AI systems produced incorrect information more than 17% of the time, while Westlaw’s AI-Assisted Research hallucinated more than 34% of the time”.

Hallucinations in Legal Practice: A Comparative Case Law Analysis‘ – Bakht Munir concludes that “it shifts the onus to the end user to counter-verify the generated content otherwise face the music.

Law, Lies, and Language Models: Responding to AI Hallucinations in UK Jurisprudence‘ – Tahir Khan of The Barrister Group discusses the epistemological and regulatory threats that hallucinations pose in UK legal practice, including the need for manual review and correction.

Beyond Guesswork: Reducing Hallucinations in Legal GenAI Tools‘ – Ryan McDonough highlights how hallucinations force lawyers to spend time validating AI outputs and suggests strategies to reduce this overhead.

Law Firms Leveraging AI: Maximizing Benefits and Addressing Challenges‘ – Logan Lathrop explores the tension between automation and accuracy, showing how hallucinations increase legal risk and workload.

AI-hallucinated cases end up in more court filings, and Butler Snow issues apology for ‘inexcusable’ lapse‘ – Debra Cassens Weiss reports on Butler Snow acknowledging that “there is no excuse for using ChatGPT to obtain legal authority and failing to verify the sources it provided”.

When AI hallucinations hit the courtroom: Why content quality determines AI reliability in legal practice‘ – Thomson Reuters highlight how poor-quality AI content leads to courtroom errors, requiring more human oversight and correction.

The risks of using GenAI for legal research‘ – Alex Heshmaty explores whether GenAI should be used for legal research (NB I am quoted within this article).

Jurists of the Gaps: Large Language Models and the Quiet Erosion of Legal Authority – Philippe Tritto and Ilsse Ortega consider that Large Language Models (LLMs) are not merely tools to assist legal professionals—they represent a deeper epistemic and normative challenge to the foundations of legal authority. While LLMs allow humans to produce outputs that convincingly simulate legal reasoning, they lack the embodied judgment, ethical intentionality, and contextual awareness that define legitimate legal decision-making. Their paper argues that the social legitimacy of the legal profession relies on capacities that are not reproducible through computational systems.

These are just a small selection of articles on the topic available online. There are countless more. Just Google it or, dare I say, ask ChatGPT. I tried the latter (requesting a list of 10) and it produced five articles with sources given and five without sources. When I asked for the sources for the latter five it admitted that it could not find articles with the titles it had previously cited! Thus, these were hallucinations creating a verification burden for me before an accurate list of 10 could be reliably cited. Indeed, I found better articles through a Google search than that produced by ChatGPT and only two of its five (non-hallucinated ones) made the final cut. Those two also appeared via a Google search. I would have saved time had I just used Google in the first place.

It is also blatantly obvious. If you use GenAI to research a legal issue and it produces, say, 12 fake citations and misses out goodness knows how many correct ones, then the verification burden is going to be huge. Furthermore, not all case law is digitalised. Even the best software will not necessarily find what you are looking for. A search in an old fashioned law library may well be necessary.

OpenAI’s unified Usage Policies (effective October 29, 2025) explicitly prohibit using their services for:

provision of tailored advice that requires a license, such as legal or medical advice, without appropriate involvement by a licensed professional

automation of high-stakes decisions in sensitive areas without human review … [including] legal [and] medical

Specific GenAI legal technology vendors followed suit, as Antti Innanen acknowledged:

Antti Innanen on Harvey and Legora not providing legal advice

These are all effectively acknowledgements of the verification burden that clearly exists.

It is also an acknowledgement of what GenAI really is. Joshua Yuvaraj, in his paper, puts it like this:

AI models are fundamentally probabilistic – they learn from input data, including its biases and omissions, and return outputs that are statistically likeliest to reflect what is requested by users. They are not structurally linked to reality: namely, factual accuracy, and valid links between ‘factual propositions…[and] relevant legal documents.’ However accurate the training data, a machine learning model does not learn the facts underlying that training data, but reduces that data to patterns which it then ingests and seeks to reproduce with variations depending on the input of a user (instructions/prompts).

The reality flaw means hallucinations – outputs that are ‘false, incorrect, or outright nonsensical’, no matter how plausible-sounding – and other errors occur frequently enough to warrant significant concern. One study of ‘public-facing’ models like GPT-4/3.5 (OpenAI), PaLM2 (Google) and Llama 2 found that, in response to ‘a direct, verifiable question about a randomly selected [US] federal court case’, 58%-88% of responses were hallucinations. Similarly, a study of GPT-4o and Llama-3-8B legal analysis documented that 80% of responses had hallucinations.

One response to this problem is to enhance the quality of the training data. In one sense this can help address concerns about omissions, biases, mistakes and other failings in the training data that can impact the output. Yet even with more high-quality training datasets, or bespoke datasets built for particular contexts, the hallucination problem is not immediately resolved. Studies have repeatedly found ‘hallucinations’ even in machine learning tools built for the legal context. Even where leading legal research companies like Westlaw and Lexis have built AI into their search functions, they remain unreliable.

Or as Jim Amos put it recently on LinkedIn:

Understand this one thing about Large Language Models if you want to stay grounded and sane:

LLMs do not think. They do not reason. They do not understand a single word they output. They don’t understand cause and effect or right from wrong or true from false.

They’re language calculators: they use Natural Language Processing algorithms to intepret prompts then generate statistically plausible answers rendered from algebraic vectors — afterimages of training data compressed into non-euclidian space.

LLMs are a feat of science and a breakthrough in computation, but they are not an emergent intelligence.

It’s a card trick, not real magic. Don’t let anyone convince you otherwise.

I like the simple stochastic parrot metaphor which can be summed up in this image:

ChatGPT v Parrot

The dangers of not verifying GenAI output is summed up nicely in this image, that Mitch Kowalski shared on LinkedIn (thinking I would like it – I do):

GenAI and edible mushrooms - Image of GenAI telling someone a mushroom was edible and once they die confirming it was poisonous and would they like to learn more about poisonous mushrooms

Mapping ‘Inkster’s Law’ against the ‘Verification-Value Paradox’

When I asked Copilot to map ‘Inkster’s Law’ against Joshua Yuvaraj’s ‘verification‑value paradox’ I got this output:

Inkster's Law v Verification-Value Paradox

⚖️ How They Align

  • Both highlight the same tension: GenAI promises speed, but lawyers must verify everything, which can outweigh the benefits.
  • Inkster’s Law gives a catchy, numerical shorthand (2:1 ratio).
  • Verification‑Value Paradox provides the theoretical and ethical framework explaining why that ratio exists.

🧩 Why This Matters

  • Together, they show that the verification burden is not just anecdotal but systemic.
  • Inkster’s Law makes the issue memorable in practice; the Verification‑Value Paradox grounds it in academic critique.
  • Both reinforce that lawyers cannot abdicate responsibility — GenAI outputs must be checked as if the lawyer wrote them.

✅ So, Inkster’s Law is the practitioner’s shorthand, while the Verification‑Value Paradox is the academic framing. They converge on the same point: GenAI may not save time if verification duties outweigh drafting efficiencies.

A little verification burden on that would be to point out that the citations for the ‘Verification-Value Paradox’ may not be correct and the ‘Verification-Value Paradox’ does not necessarily agree with the 2:1 ratio of Inkster’s Law. Otherwise, Copilot has grasped the similarities between the two.

Conclusions

In my opinion the evidence is very clearly stacked against Antti Innanen’s viewpoint and what ‘evidence’ he has produced to counter Inkster’s Law and the findings by Joshua Yuvaraj in his paper. The verification burden is real and has to be taken very seriously. It cannot simply be dismissed as “funny but untrue”.

What that verification burden may be, will vary from case to case. Sometimes it will be a cancelling out of any benefit as Joshua Yuvaraj suggests. Other times it may be a 2:1 ratio as I suggest in Inkster’s Law. However, the verification burden could be higher than that in certain instances. Inkster’s Law is, I consider, a good average and warning point.

That is not to say there might be other uses of GenAI in legal practice that would not carry the same risks or verification burden. This may be the case in legal marketing, for example. Although, I would not personally use it for that due to the bland nature of the output and lack of any clear unique voice to associate with your brand.

Summaries

Reactions to The GenAI Verification Burden in Legal Practice

On LinkedIn the following comments have been made:-

Antti Innanen (⚫ Ⓜ️ 💯):

Good stuff!

Remember what we are debating: “Paradox suggests increases in efficiency from AI use in legal practice will be met by a correspondingly greater imperative to manually verify any outputs of that use, rendering the net value of AI use often negligible to lawyers.”

We are not debating whether AI outputs should be verified. They should. We are not debating whether verification takes time. It does.

The question is whether the verification burden “renders net value of AI use often negligible to lawyers.”

It doesn’t.

The blog post is mostly a collection of criticism of the studies I brought up. And the criticisms might be valid in some situations.

But I didn’t see anything that would prove this “paradox.”

Actually, it shouldn’t be that difficult to test.

The question is simple: “do you have to use so much time to verify AI outputs that the net value of AI is often negligible?”

If a large group of lawyers (something like half) said yes, the paradox would make sense.

But the more realistic answer is: Yes, verification takes time. No, it doesn’t make the net value of AI negligible.

Me:

You say “It doesn’t”. I say “It does”. In my own experience the time spent does negate the use of it and usually far outweighs that use. Others agree as cited in my blog post. The evidence is fairly clear especially where legal research is involved. As I say the ratio will vary from case to case but it will be rare in legal practice currently that the time involved does not cancel any perceived benefits of using GenAI in the first place.

Antti Innanen:

I have to admit that I actually thought “Inkster’s Law” was a joke.

Like an insightful quip that reminded us that verification is a burden and that we should always verify the outputs.

And as such it was fun and insightful. A bit absurd too.

😅 But now we are arguing that it is actually true.

Me:

That’s what I thought about the metaverse 😉

It was never a joke. It was based on real life experience and my thoughts on that in 2023. Gav Ward coined the term following a debate on the topic on LinkedIn. Two years later and I don’t think a lot has really changed.

Still can’t see why it would be “a bit absurd too”. My blog post sets out in detail why it is not. “Funny but untrue” simply does not cut it.

Yes, we are arguing that it is actually true because I believe it is and you do not.

ChatGPT will tell you that the “2:1 ratio is flexible; complex or poorly prompted AI outputs may require even more verification.” Likewise, those of the variety suggested by Jennifer Marsh, in this thread, will likely require less.

I have personally used GenAI for a timeline summary that was quick and fairly accurate and drew all its information from blog posts I had written over several years. I was able (from my expert knowledge on the subject and as author of the source material) to quickly verify it. It saved me time and the 2:1 ratio did not apply. But for legal research (and especially where you do not have subject expertise) the story is very likely to be a different one.

-+-+-+-+-+-

Mark Bennett (Associate Professor of Law at Te Herenga Waka | Victoria University of Wellington):

This is a weird debate.

The question has to be answered empirically.

I see lawyers who use AI answer that it does help their efficiency and effectiveness. I find the same benefits with my own academic research and administration work. The use cases – legal and otherwise – are described and demonstrated all over the place, and can be tested by trying to use AI yourself.

Why do we need the debate? If using AI wastes your time, don’t do it. If it saves you time and improves your work, do it.

Me:

Many don’t realise the pitfalls involved. Debates like this hopefully raise awareness amongst the uninitiated. Then, as you say it is their choice, but armed with information to assist that choice.

-+-+-+-+-+-

Jennifer Marsh (Product Leader | Former Practicing IP Attorney):

It’s true in general BUT the best use cases don’t require you to validate the output. For example, rephrase this sentence to x, explain this to me in plain English, conversational intake workflow, tell me why I might be wrong, etc.

Me:

But those are not the ‘use cases’ that are going to replace lawyers 😉

-+-+-+-+-+-

Daniel Yim (Founder @ Sideline (track changes in Outlook) | Helping Law Firms Develop Their Lawyers | Legal Tech & Innovation):

Yeah but more verification time = more billable hours = good. 👍

Anyway, everyone’s mileage will vary depending on the task and context. I think use the AI if it’s useful for your work, don’t use it if it’s not, and definitely don’t use it just because big tech (or hangers-on to big tech) tells you that you have to use it.

Me:

Indeed. And they told us AI would kill the billable hour!

-+-+-+-+-+-

John McCarthy (Helping Law Firm & Legal Tech Owners unlock £100K extra profit in 6 months, build a highly profitable, valuable business & plan for exit with the PROFIT System):

It comes down to risk- benefit ratio too Brian Inkster

Using AI may save time (good) but increase risk for lawyers if the appropriate checks and balances are not in place (bad).

My thoughts are that we will move towards AIs use but it should be done with caution ⚠️

Me:

Approach with caution indeed. The problem is many lawyers are throwing caution to the wind as per the 545 cases identified so far by Damien Charlotin.

-+-+-+-+-+-

Ryan McDonough (Head of Software Engineering – AI in Legal: Building, Governing, and Owning the Tech):

Thanks for including my work Brian, fantastic article! The bit I keep seeing in real workflows is that the drag isn’t mistakes, it’s the clean paragraphs with no real idea of how they were produced, which just pushes lawyers back into doing the reasoning themselves.

Accuracy always grabs the headlines, but predictable behaviour is what would actually reduce verification time. That feels like the real gap the next wave of tools need to look to close

Me:

And clearly a real danger in those clean paragraphs if you don’t know where they came from. Many will, unfortunately, accept them without question.

-+-+-+-+-+-

Chantal McNaught (That LegalTech Redhead | 🎙️People in Legal | Answering the question “how lawyers can navigate the conflicts between law as a profession and law as a business?”):

The marketing “lie” about generative AI in legal practice is that it is merely a “productivity tool”. It isn’t. It fundamentally changes what the practice of law looks like, and simply validating the outputs will absolutely add to the overall time it takes to complete legal work product. Lawyers should consider the high-value vs low-value legal work products in the delivery of legal services, and design their processes accordingly.

-+-+-+-+-+-

Robert Dean, J.D. (Director, Litigation | UnitedLex | eDiscovery for in-house and law firms | Legal data annotation for AI):

It takes your associate 10 hours to draft a memo. It takes you an hour to check the cites / sources before filing. Or, it takes your AI Copilot 10 minutes to spin up a brief. It takes you an hour to check the cites / sources before filing. I don’t see the burden – what am I missing?

Antti Innanen (⚫ Ⓜ️ 💯):

According to Inksters law you spend only 20 minutes verifying so it takes only half an hour 🤣

Robert Dean, J.D.:

That’s right. The act of “verifying” outputs is what lawyers in firms have been doing since the invention of associates. That the output is AI-generated doesn’t change the profession. It’s part of doing good work, not an argument against adopting more efficient tools.

Me:

The Associate probably won’t make stuff up (hallucinate) and their cites/sources will be easier/quicker to check. Your AI Copilot might well spit out fabricated nonsense and miss much that you would want included along the way. As a result you just might spend a lot longer checking the AI Copilot output than you would have to with the Associate output.

Robert Dean, J.D.:

Hallcuination itself has been a fundamental part of machine learning since Rosenblatt in the 1960s. A model is always limited by the training data, and the hyperplane will be drawn anywhere within the bias range. It is within that range where the machine “hallucinates” as it has to guess on which side of the plan the input data might fall. Nothing wrong with that; you can fill in the bias gap with fine tuning to add more data. It’s just math – and, the size of the models (ie the availability of more data to calibrate the hyperplane and reduce gaps) is only getting more robust. Kimi K2 just released a trillion (!) parameter model.

But your premise that “cites/sources” are easier to check with an associate generates the output versus the an AI Copilot is just not true. Ever generate a brief using GPT-5.1 with Deep Research? It spins up a brief in less than 20 minutes. You get Bluebook citations and can spend the same amount of time checking the cites as you would any other work product.

Bottom line, verifying is what lawyers have always done. Whatever extra time is spent verifying AI work product is easily captured by the efficiency of using AI in the first place.

Me:

It will be true whenever the Associate’s output is better than the Copilot’s output. You cannot say that will never be the case. Whilst hallucination may be a fundamental part of machine learning, I don’t think it has ever been a fundamental part of Associate learning!

Antti Innanen (⚫ Ⓜ️ 💯):

Robert Dean, J.D. – That is my point: legal work is largely about verifying sources and reviewing the final product. It always has been.

The need to verify AI generated or AI assisted work is nothing new. There might be some added burden, but it does not erase the efficiency/other gains.

Many people may not realise how strong the latest frontier models and legal specific tools have become. In 2023 this was a much bigger problem than it is now.

Me:

Antti Innanen – The problem remains big in 2025 as daily reports like this one prove: https://www.linkedin.com/posts/ketanjoshi1_two-federal-judges-say-use-of-ai-led-to-errors-activity-7396272782596014080-If_Y?utm_source=share&utm_medium=member_desktop&rcm=ACoAAAL3d5kBQ5O9lwNOHjqRNQpnUNzdQSl6_48

-+-+-+-+-+-

Anna Guo (📕 Lawyer | Legal AI Researcher):

I honestly don’t see how this can be properly debated, because the tradeoff is different for everyone and depends on what they’re using AI for.

I agree with what Jennifer Marsh said. Some tasks, like using AI to brainstorm ideas or for polishing, require very little verification, whereas using AI to summarize open-ended queries can require much more in a high-risk scenario.

And let’s not forget the study published by METR earlier this year, which found that engineers did spend more time than expected verifying and correcting AI-generated code, but also that the cognitive burden of reviewing AI output was lower than writing the code themselves. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

Both halves of that result matter.

At the end of the day, it’s important to understand both the failure modes and the benefits of using AI assistance.

That, I think, is the consensus we can all come to here.

Me:

It has been fully acknowledged that the trade-off is different for everyone and depends on what they’re using AI for.

The debate arose from a suggestion that the GenAI verification burden in legal practice was “funny but untrue”.

Allowing lawyers to think so creates the danger of perpetuating the non-verification seen in the 545 cases identified so far by Damien Charlotin which are rising in number by the day.

Debates like this are necessary for educating the uninitiated. I trust that my blog post, and the comments arising on here from that, will do just that.

If that results in a greater exercise within legal practice of the caution that John McCarthy refers to, then it will have been a worthwhile debate.

And this assists, as you say, “to understand both the failure modes and the benefits of using AI assistance.”

Then, yes, I too think there is the consensus we can all come to here.

Antti Innanen (⚫ Ⓜ️ 💯):

No, we are not debating verification burden in general.

We are debating if ”increases in efficiency from AI use in legal practice will be met by a correspondingly greater imperative to manually verify any outputs of that use, rendering the net value of AI use often negligible to lawyers.”

Lets not move the goalposts.

Me:

I’m not moving any goal posts. That debate is about the verification burden and the weight of it. You dismiss that weight, I do not.

Antti Innanen:

You are.

What I am saying that verification is important and it surely takes time.

But it doesn’t ”render the net value of AI use often negligible to lawyers”.

Me:

Not always (I accept that), but more often than not (especially when actual legal tasks are involved).

Sam Davidoff (CEO at Align | Former Trial Lawyer | Electronic Binders for Litigators | Litigate with your iPad | Using my 20+ years of litigation experience to give lawyers practical tools and advice that will make them more effective):

Anna Guo – something else about the METR article that seems relevant to this debate is that it emphasizes that people are bad at assessing their own efficiency. It’s not just that the devs WERE 19% LESS efficient, it’s that they also THOUGHT they were 20% MORE efficient.

It always makes me skeptical of claims that run along the lines of “but I use it and I’m more efficient” (whether made individually or in aggregate via surveys). I’m not even saying those claims are wrong. I’m just saying I have no idea how to evaluate how reliable that evidence is coming from lawyers. (Though I will note I have some innate skepticism from my time as a lawyer and seeing how bad lawyers are at estimating how long something SHOULD take to do.)

Josh McBride (Barrister at Richmond Chambers):

Brian Inkster – honestly this is like saying Microsoft Word is complicated and takes ages to learn before you will realise any efficiencies, so you should just keep writing with a fountain pen!

Me:

Not a good analogy. What we are debating here is the inherent hallucinatory problems of GenAI (not to mention its other drawbacks) and the verification burden placed on detecting and correcting those hallucinations (and filling in the gaps it will simply not output). This is not an issue when using Microsoft Word.

Josh McBride:

it’s an excellent analogy. “Hallucinations” are only an issue with generative AI if you don’t know how to use it. Just like Word won’t correct typos if you don’t know how to use spellcheck.

Me:

Spellcheck runs for me in Word every time I use it. I didn’t realise there was a way of switching it on and off. Just as well it has always been on! Do please tell us how you turn hallucinations off.

Abigail Hope S. (Empowering GCs across APAC to lead strategy, digital transformation & innovation | Senior Manager, Legal Transformation at SCG | MIT Sloan MBA | JD/LLM):

Anna Guo – I love your point about context-dependent tradeoffs, Anna and it deserves more attention than it gets. Most corporate AI policies today are blanket AI policies. For our company it’s “use it everywhere” and “verify everything equally” policies, where both miserably fail. High-stakes summarization for a tax opinion from legal needs different guardrails than UI copy generation from corporate branding, but most AI governance frameworks treat verification as binary. The failure modes aren’t just different in degree, they’re different in kind.

Brandon Barney (Mission: Scalable C-Suite IP Intelligence. Partner to PE/VCs on SaMD, SiMD, FemTech, Robots, DeepTech IP. Also, Stents and Catheters!):

Anna Guo, Abigail Hope S., Brian Inkster – I think that there needs to be a new benchmark for the time it takes humans paralegals and attorneys to verify the output of an AI and for that time to be analyzed systematically across all tasks. That is the work that we are doing currently and we do not think its wise to make a lot of claims yet. What I can say is Artificial Intelligence is true to its name.

Abigail Hope S.:

That it is artificial/non-human or that it is intelligent? Or both?

Brandon Barney:

I focus on the denotation, as the connotations of ‘intelligence’ are often speculative.

I see AI in a superposition of three states:

Non-human

Asset-Intelligence (when it’s right)

Liability-‘Intelligence’ (when it’s wrong)

The wisdom is in knowing the difference—which is precisely what we are trying to benchmark.

Me:

I look forward to seeing your benchmark report in due course.

-+-+-+-+-+-

Robyna May (CIO – Specialising in Innovation, Strategy, Change & Leadership – Writer & Speaker – AI enthusiast | McInnes Wilson Lawyers.):

AI hallucinates. It’s not a problem that’s been solved. Therefore, where using it generate content, it has to be checked carefully. The analogy to an eager intern rings true. Where we use AI differently, say summarising a document, it can still hallucinate but the burden is lower. Where I see the burden being the absolute highest is where the other side is using AI and it has not been checked whether that’s a self represented litigant, or a lawyer who doesn’t quite understand AI.

Me:

Many think it is unlikely ever to be solved. An inbuilt ‘feature’ apparently! There is, I believe, a difference between GenAI and an eager intern though. An eager intern may well make mistakes but is unlikely to make up case law that never existed. Furthermore, they will hopefully learn from those mistakes as they are corrected by a supervisor. However, GenAI will not and you are always back to square one when using it.

Robyna May:

I think the analogy is more helpful just in terms of it HAS to be checked. Rather than the nature of the mistakes made. That said, there is much more reward in guiding a new lawyer to think about their reasoning rather than picking out completely fabricated cases. And that perhaps is a whole other discussion entirely.

Me:

Agreed! Although, I have seen some using the analogy as though there is no difference between the two when, of course, there is.

Jennifer Marsh (Product Leader | Former Practicing IP Attorney):

Robyna May – It will not be solved because that’s how GenAI works. It creates the best response, not the correct response. It needs to be used in such a way where that is what you need.

Robyna May:

Agreed. It’s trying to replicate what a great response looks like, and it fills in gaps. It’s interesting that RAG hasn’t closed the gap.

Jim Amos (CTO, Human-first technologist, full-stack software engineer, coach, writer. Follow for original critical thinking and leadership perspectives in the AI era.):

Robyna May – not just hallucinations, also bias, and totally misconstrued concepts based on misinterpretation of certain language idioms and contexts. Relying on an alien to discern human legal narratives seems like a very bad idea to me.

-+-+-+-+-+-

Gerard Stegmaier (Practical Problem Solver, Trusted Consigliere + Lawyer):

Begs and important question about whether AI-native lawyers will have the other skills required for this human-in-the-loop work and where it will come from. Filing of pleadings and issuance of decisions with hallucinations happens for a reason and that reason is merely a new symptom of an older problem—quality assurance.

Me:

Indeed. Might it be the case that those skills will diminish as there is more reliance placed on AI?

-+-+-+-+-+-

Teresa Villa (A trauma-informed technology ecosystem restoring transparency, due process, and civil rights protections through ethical AI. Includes The Justice Engine™, Notary of Justice™, BitOfJustice™, Solutions4Value™, LegalFlow):

What do I think ?
This is exactly why I Built Justice Engine and why it matters.
People verifying legal AI can’t afford:
• hallucinations
• fabricated citations
• made-up rules
• “almost true” answers
The stakes are TOO high.
lived experience fighting wrongful garnishments, misapplied payments, and fake registry numbers

Do you see hallucinations in GenAI output? Here’s my real-world answer transparent, no fluff:
Yes. GenAI does hallucinate.
All models do. Even the best ones still occasionally produce:
• incorrect facts
• outdated info
• misattributed citations
• overconfident answers
• “logic leaps” that sound right but aren’t
Not because they’re malicious but because they’re probability machines, not truth machines.

Do I find myself having to verify GenAI
yes people do have to verify GenAI output.
Especially when:
• the topic is legal
• the conversation involves benefits
• the data affects someone’s rights
• the system must handle trauma safely
• the stakes are high

How much time is that taking me?
Verification can take anywhere from 30 seconds to 30 minutes, depending on:
• whether the user knows the topic
• how complex the question is
• how sensitive the consequences are

Me:

Thanks. Good lists. Also important is what it misses out. The verification burden of ascertaining that can be high.

Teresa Villa:

You’re right the verification burden is enormous.
And that’s actually the part I built JusticeTree around.

Most tools treat verification as an afterthought.
In JusticeTree, verification is the architecture:
• Every output is tied to a source, statute, or record trace
• If the model can’t show its grounding, it isn’t used
• Human review isn’t optional it’s the governing layer
• Overrides create a documented trail, not a productivity penalty
• Survivors and claimants stay in control of their own narrative

So instead of burying people under “prove the AI wrong,” the system reduces that load by design.

In legal contexts, verification shouldn’t be a burden it should be a right.

-+-+-+-+-+-

Michael Lawrence (CTO | Private Intelligence Operating Systems Architect):

This entire debate assumes a world where AI is used as a loose tool giving unstructured answers.
In a governed system — where outputs are jurisdiction-bound, rule-constrained, logged by design, and routed through human oversight at strategic checkpoints — hallucination is engineered out and the verification burden drops to near-zero.
The problem isn’t AI.
The problem is architecture. Juris by CalyxOS

Me:

Indeed. On the whole that, unfortunately, appears to be the world that most lawyers are living in at the moment. I have seen #LegalTech vendors add a GenAI layer to their product without testing, requesting permission, providing warnings, not giving any training on its use etc. What was previously a governed system suddenly became an ungoverned one. Some vendors are creating rather than solving the architecture problem.

-+-+-+-+-+-

Brian Belt (Founding Shareholder | Experienced Corporate, RE & Hospitality Attorney | Legal Tech Enthusiast):

The article by Joshua Yuvaraj was purely theoretical- no real world testing of the theory. I see no recent real world testing of the supposition. This also does not match our experience. It’s too much easier to verify output from a quality, legal AI vendor than from an associate, since documents and clauses are so readily accessible (hovering over or clicking on links). There is no comparison with the time necessary to verify another human’s output. Everything that I have read, is that when legal AI is applied to the correct use cases, a considerable net time savings results.

Me:

How do you measure “quality”. There are currently many #LegalTech AI vendors and new ones appearing on the block all the time. Many lawyers are using more readily available GenAI tools often not realising the issues associated with those. Specific LegalTech AI tools are not without problems as research cited by me demonstrates. Before GenAI, LegalTech tools were available to research case law. These were possibly more reliable than they are now with a GenAI layer added. Just because you can doesn’t mean you need to. You would hope a good Associate would cite their sources (even hyperlink them in this day and age). If not, they have not been well trained. And what are the “correct use cases”?

-+-+-+-+-+-

Kevin Keller (General Counsel and Tech Builder | Inventor | Investor | Advisor | co-Founder | Board Member – Helping develop hardware, software and SAAS solutions for 25+ years; developing applications to enhance learning using AI):

This is what validation/verification of AI output looks like in the system we’ve built – every reasoning step and bit of output has provenance and a confidence score – human expert validated facts/reasoning approaches get scored higher, web search and synthetic data lower. You can look at the overall scores and what goes into each step to determine where to dive in more yourself to understand whether and how much to trust the output.

Kevin Kellar - validation - verification of AI output in a system

Me:

Thanks. We need more of this built within GenAI systems.

-+-+-+-+-+-

Paul Correa (Experienced legal leader):

Reductionist arguments on both sides. Neither capture the nuance or value. 1. You are already supposed to “verify” by reading cases, reading documents, knowing your case deeply, and using your expertise. (I require all my legal writers to use pinpoint citations with a parenthetical quote from the case. )So, the “verify” work already existed, if you were doing your job right. 2) Now, AI uncovers more data, it researches more than we were able to before, it brings new insights. That’s additive, adds time, improves quality. 3) Ultimately, if it takes longer to use AI, but the output is better, that’s a gain. So, the added time cost does not cancel out the benefits! You have more work to do and you have better output. The robots aren’t here to do your job for you.

Me:

It does not necessarily follow that “AI uncovers more data, it researches more than we were able to before, it brings new insights”. It depends on what the AI has been trained on (does it cover your jurisdiction, legal subject area and has all the law in that area been digitalised and given to the LLM in question?). It also depends on the prompts given to it. “New insights” might be the hallucinated ones 😉 If the output is better that will be a gain. If it is worse it will not be.

-+-+-+-+-+-

Josh McBride (Barrister at Richmond Chambers):

I’m genuinely puzzled by all this! Have you tried Notebook LM? Or other RAG tools? When deployed alongside tools like Claude, Notion, and Wispr, the efficiencies are well beyond extraordinary. To suggest it takes “longer” because I then need to check the work – produced in seconds/minutes – is risible and demonstrably untrue.

Me:

Notebook LM is for uploading specific documents to and then using it to analyse them. As stated elsewhere in this thread RAG has not eliminated hallucinations. The arguments by me and Joshua Yuvaraj, on the whole, revolve around using GenAI for legal research. When you are preparing for a court case do you rely solely on GenAI for the authorities you cite? You may have found a set of AI tools that work for you. Many lawyers have not as per the 545 cases identified so far by Damien Charlotin.

Nikolaos Papavasiliou (Legal Counsel at Ray White | Independent Legal AI Consultant | LLB(Hons); BSc(BiomedSc); GDipLegPrac; DipSusLiv):

Brian Inkster – agreed that folk unable to establish, integrate, and operate AI workflow / suites properly should not attempt each of same; however, they absolutely should learn how to do so.

advantageous to spec into new mastery where the option is (checks notes) available by asking the product to train you. time commit now is an infitismal contribution (here, a genuine 0-approach if ai progress is accepted-exponential) as against efficiency gains.

qualifying, we could wait for the products to improve before training. there’d be a shorter inflection period from novice to mastery – but that is at the cost of time gained now at a ‘promise’ to self to pick it up later. may as well hold out for entire automative duties in such case.

Josh McBride:

Brian Inkster that’s not what your post says though. You say “increases in efficiency from AI use in legal practice will be met by a correspondingly greater imperative to manually verify any outputs of that use, rendering the net value of AI use often negligible to lawyers”.

That is quite a claim! I and many others struggle with that proposition, given what we are witnessing IRL with these tools.

If you are talking only about legal research (for most lawyers only a small fraction of their working day) then I suggest you try getting Chat GPT to draft bespoke Boolean searches of legal databases for research topics.

It is a game changer. And yes, I rely on this, and yes it saves lots and lots of time.

Me:

Nikolaos Papavasiliou – Thanks for those thoughts. Important too is training on when GenAI could be used and when it probably should not be used in legal practice.

Josh McBride:

Nikolaos Papavasiliou – 💯

Me:

Josh McBride – That is what Joshua Yuvaraj says in his academic paper on the matter which I agree with when it comes to legal research. I trust you have read the paper?

Giving what we are witnessing in real life with the lack of GenAI verification in legal practice and the consequences of that, it is clear that the verification burden is high.

You have not answered my question: When you are preparing for a court case do you rely solely on GenAI for the authorities you cite?

Josh McBride:

Yes I read the paper when it came out. Unfortunately, it seems to be an academic hypothesis as opposed to a study of how AI can actually be used in legal practice. And yes, I use generative AI to carry out research. I use it to generate Boolean search terms which I then run over legal databases. The results are terrific and extremely fast. And with zero hallucinations.

Antti Innanen (⚫ Ⓜ️ 💯):

Josh McBride – a scientific paper about verification makes an unverified and sweeping claim.

(“rendering the net value of AI use often negligible to lawyers”).

Yes, there is a paradox hiding there somewhere…

Me:

Josh McBride – Sounds like you are using GenAI in a very specific way and not how many lawyers are using it causing the problems and verification needs raised in the paper. Those lawyers could no doubt benefit from your insights. Have you written in detail anywhere about your methodology?

-+-+-+-+-+-

Viraj Deshwal (AI | Physics):

The verification burden is real, but the root problem is deeper. GenAI citations are fundamentally probabilistic.

Current systems return similarity scores for source citation like “95% similar” which are mathematically sound but legally meaningless. No lawyer can present “95% similar” as evidence in court.

Regulated industries can’t work with “the AI said so,” they need deterministic, verifiable proof. Until this changes, the 2:1 verification burden will persist.

-+-+-+-+-+-

Refat Ametov (Driving Business Automation & AI Integration | Co-founder of Devstark and SpreadSimple | Stoic Mindset):

What often gets overlooked is that verification isn’t just about catching hallucinations – it’s about catching misframings. Even when the facts are correct, the model’s implicit assumptions can quietly distort the legal position.

Me:

Indeed. Also, we can’t forget about what it misses out completely. That is a huge part of the verification burden.

-+-+-+-+-+-

Alex Heshmaty (Legal Content Specialist):

At present AI output definitely needs to be carefully checked, although I suppose the same could be said of work done by, say, a junior paralegal. However, if law firms decide to make a junior paralegal redundant and pay for AI, they are basically putting someone out of a job and lining the pockets of Big Tech – so the ethics of replacing humans with machines is highly questionable.

Even if AI gets more reliable and/or the firm makes more profit due to lower wage bills, ultimately taxes will need to go up as a result of more unemployment and the consequent welfare spending increase. The only party that significantly gains financially is Big Tech which generally manages to avoid a lot of tax through base erosion and profit shifting (BEPS).

Me:

You can’t compare AI to a junior paralegal. The latter learns from mistakes and, hopefully, does not make them again. The former does not. Furthermore, there is so much more that the latter can do that the former will never be able to do. I saw this earlier today: https://www.linkedin.com/posts/sangeetpaul_reshuffle-activity-7396452829944541185-JGAm?utm_source=share&utm_medium=member_desktop&rcm=ACoAAAL3d5kBQ5O9lwNOHjqRNQpnUNzdQSl6_48

Alex Heshmaty:

Arguably AI can also learn from its mistakes, albeit with the lack of understanding about why something was a mistake. But yes, humans are obviously more versatile and that article you linked to is trying to be hopeful about the ability of humans to adapt to a jobs market where certain tasks have been automated. However, if “AI” ends up being able to carry out the majority of economically viable tasks, white collar workers in a services economy like the UK are in a spot of bother. At the moment all the momentum seems to be with the concept of replacing humans with AI wherever possible (ultimately as a result of fiduciary duty to shareholders to maximise profits), but I guess we could reach a tipping point where AI is considered to be a signficant enough liability (eg due to repeated mistakes) that the savings made from the difference between software licensing costs and the wage of a human worker is no longer worth it.

Me:

Not sure GenAI can unless you tell it that it got it wrong. But then it will be a correction possibly in that thread only. It will usually start from scratch if you ask it the same thing again. Often you will get different answers at different times to the same question. It will also train on its own mistakes if they are published somewhere for it to do so! This it has been said could lead to model collapse as it feeds off its own ‘AI slop’.

…. on the much more that a junior paralegal can do point see: thetimeblawg.com/2024/06/09/moving-house-automation-lawyers-and-genai/

Alex Heshmaty:

Yes – the “enshittification of the internet” could definitely cause AI to eat itself. I was interested to see Google’s CEO yesterday warn about AI errors, but I wonder if that’s an attempt to shore up their primary business model which appears to be threatened by AI (even though it’s trying to cover all bases with its own AI!). On the point of correcting mistakes, my understanding is that you can essentially retain a “thread” indefinitely so you effectively teach the AI over time and it remembers all your responses etc.

Me:

But is that just the particular thread or does it actually ‘learn’ from that for the purposes of others asking the same question? And, if so, that is all very well if the person ‘teaching’ it has got it right!

Alex Heshmaty:

My understanding is that they now all have “persistent memory” features, eg https://www.independent.co.uk/tech/google-gemini-update-ai-memory-b2807453.html

Robots are certainly still useless at doing the dishes! https://www.tiktok.com/@wallstreetjournal/video/7568535282358734093

Me:

We have dishwashers for that!

It’s like getting GenAI to do tasks that can be done by document automation.

The “persistent memory” appears to be linked to a specific user rather than it ‘learning’ from one user to put that to use with other users?

Alex Heshmaty:

Correct – but I guess you could have shared corporate versions which learn from multiple users? And presumably the big LLMs leverage user data to attempt to improve the next models, as well as pulling stuff from the internet.

Me:

But in reality they are reading and stringing together words?

Alex Heshmaty:

Indeed. On a philosophical level, GenAI is moronic. But if a company can make more profit by licensing a stochastic parrot and reducing its wage bill, it’s generally going to put profit over any philosophical concerns. Having said that, I’ve heard that law firms are being asked by some clients specifically not to use any AI – whilst others are demanding their lawyers use it (and reduce their fees accordingly!).

Me:

Just saw this: https://www.linkedin.com/feed/update/urn:li:activity:7397306787466694656?utm_source=share&utm_medium=member_android&rcm=ACoAAAL3d5kBQ5O9lwNOHjqRNQpnUNzdQSl6_48

-+-+-+-+-+-

Jim Amos (CTO, Human-first technologist, full-stack software engineer, coach, writer. Follow for original critical thinking and leadership perspectives in the AI era.):

I’m convinced most of the productivity gains from AI in most industries are overstated and misconstrued. People want so badly to believe they are now part of a tribe of efficiency superhumans because that’s how these products have been marketed.

-+-+-+-+-+-

Ralf Haller (Founder & Business Development Leader | Driving Growth in High-Tech, AI & Data | Global Product & GTM Expertise | MIT & KIT MSc):

This paper – you keep mentioning – looked only at pure LLMs — which, according to OpenAI’s own usage policy, are not suitable for legal applications.

I’d recommend evaluating solutions that are actually designed for legal work.
Otherwise it’s like judging all lawyers based on a group of people who never studied law and concluding that all lawyers must be incompetent.

The comparison simply doesn’t hold.

Me:

Many ‘designed’ for legal work are those same LLMs in a wrapper. The paper also points that: “They [structural flaws inherent to AI technology] are also structural features of both general-purpose AI services (that members of the public and/or organisations can use for any purpose, and purchase escalating tiers of service for, like ChatGPT, Microsoft Copilot, Gemini), and bespoke legal profession-targeted services (e.g. CoCounsel, Harvey AI). This distinction matters because claimed points of difference between the latter and former categories is their reliability and reliance on high-quality data for lawyers. Such claims do not, however, address the fundamental structural flaws inherent to machine learning models outlined below.”

The lawyers who are using these technologies (general-purpose or bespoke legal) without proper verification are indeed incompetent as the 545 cases identified so far by Damien Charlotin demonstrate.

-+-+-+-+-+-

Mitch Kowalski (General Counsel/Senior Real Estate Counsel | Commercial • Development • Governance (ICD.D) • Risk • Airport | Dual-Qualified (Ontario & Massachusetts) – US Green Card – EU/Canadian Citizen):

I’ll prepare a response later this week. I read Joshua’s paper twice and will read it again this week. My view is that it is pretty “light” on data to support his thesis, and most of the paper is just a rehashing of old news on Gen AI (like poor use by lawyers) and is general knowledge filler which does not advance knowledge in this area.

I highly suggest that litigators (which is the very narrow area that Joshua reviews) watch this video from ClioCon 2024 which lays out the litigation use cases -and shreds Joshua’s thesis.

I’ll write more later.

Start at the 20 minute mark.
https://www.youtube.com/watch?v=Q8T24F-5-AE

Me:

I look forward to seeing it. In the meantime, I don’t think the video you refer to “shreds Joshua’s thesis” at all.

It could be argued that the video can be discredited immediately due to its reference to drills and holes 😉 It is also effectively a #LegalTech vendor’s sales pitch.

However, putting that to one side, in brief:

➡️ It references audio to text technology, timeline creation and brainstorming. All of which are relatively safe uses of GenAI that will potentially safe time and not engage a high verification burden in the way that legal research would.

➡️ It looks at a predictive model which Joshua’s paper specifically excludes.

➡️ It looks at a specific legal legal tech tool where you upload documents for analysis.

➡️ The dangers of general LLMs is made which is the same as Joshua does in his paper and I do in my blog post.

➡️ However, the suggestion that we don’t need to know what is going on in the black box as we don’t know what is going on in a human brain and if the output looks like reasoning it can be accepted as such is bonkers and very dangerous!

Mitch Kowalski:

As promised:

There has been much chatter on LinkedIn about a new academic article about the Verification-Value Paradox (of GenAI use by lawyers).

The article claims that it is doubtful that GenAI delivers value to lawyers because any efficiency gains are erased by the time spent verifying its output; a framing that the author calls the “verification-value paradox.”

The “paradox” is:

More AI = more verification = less value.

The author admits that this paper is not based on fresh, robust empirical evidence, as he waffles back and forth in much of his discussion of the paradox; essentially stating that GenAI is sometimes helpful and sometimes not.

From the standpoint of a legal practitioner, this is the equivalent of saying walking across the road “can” be dangerous.

The decision to implement GenAI in a legal practice is to attain gains over elements of different files – not to say, “Well, if I can’t use it for legal research, then it is of absolutely no use to me whatsoever.”

The author’s ultimate, ordinary and unhelpful conclusion is, “for lawyers to be cautious, critical and/or reticent as to AI use in legal practice.”

This is precisely what law societies, bar associations and courts have been saying for the last 2 years. So, in other news, the author is telling lawyers that the sky is blue.

Ironically, the paper does little to support the proposition that GenAI is not useful for lawyers; the evidence is selectively presented, and the conclusion is not supported by practitioner studies or time-motion analysis.

Be that as it may, there are 5 important points that I wish the author had dealt with as they are far more interesting and relevant to practicing lawyers.

1. Verification isn’t a paradox; it’s the lawyer’s job

Lawyers verify everything: junior memos, discovery summaries, expert reports, and even partner-drafted factums. This isn’t burdensome; it’s what competent practice looks like.

AI doesn’t create a new obligation to verify. It shifts where verification happens and how much time we spend on low-value drafting versus high-value analysis. That’s workflow design, not a paradox.

The assumption that verification cancels out gains ignores basic realities observed in firms using these tools (RSG Study). A decent first draft (say 60% quality) saves meaningful time. Instead of starting from a blank page, the lawyer starts from a structured, reasonably coherent draft and moves immediately into:

(i) refinement,

(ii) strategy,

(iii) judgment, and

(iv) nuance.

GenAI is not designed to replace any of the above as these items are the core values that lawyers bring to the table.

Editing is faster, deeper, and more analytical than authoring every word from scratch – even with verification.

This framing – author to editor – is the much more interesting and relevant shift in the profession, but the article ignores it entirely.

In fact, the author spends a mere paragraph in his 28-page paper on “transactional and in-house lawyers who never engage with the court” and simply says the verification costs of GenAI use by these lawyers is also very high. Empirical practitioner studies and time-motion analysis are needed to support his assertion, but he has none.

2. The uncomfortable truth: many litigators don’t read the cases upon which they rely

The author suggests that because the need for accuracy is so high in litigation, the verification costs are therefore too high to use GenAI for legal research. However, he does not provide any data to support his view that verifying cases pulled by GenAI is slower than cases pulled and verified by other means.

He fails to mention that GenAI hasn’t created a verification crisis for legal research, it has revealed one.

For decades, litigation practice has tolerated:

(i) template-based arguments with inherited citations,

(ii) reliance on headnotes instead of underlying reasons, and

(ii) recycled factums passed from junior to junior.

GenAI didn’t cause this prior behaviour, but it is now shining a very bright light upon it.

Instead of blaming the technology and suggesting that it is causing problems, we should thank it for surfacing real problems in the profession.

3. It is absurd to use general-purpose LLMs for legal research, then blame Gen AI when it goes wrong

This may be the weakest assumption in the article: that using a general purpose LLM, such as ChatGPT or Gemini for legal research is a meaningful test of GenAI’s value in legal practice.

It’s not.

It’s negligence.

General-purpose LLMs are not legal research engines.

That’s like using a butter knife as a scalpel and concluding “surgery is unreliable.”

Wrong tool, wrong workflow, wrong conclusion.

The author cites only one South African Legal Research GenAI start-up (Legal Genius) as being his proof that legal-specific GenAI is unreliable; the bigger players in legal research including the author’s part-time employer, Thomson Reuters, are probably rolling their eyes at his example. Moreover, the use of only one example seems to be a leap unworthy of an academic paper, especially since as of September 2025, Legal Genius claims to have eliminated hallucinations using RAG.

4. The author assumes humans are more accurate than technology. They aren’t.

A silent premise in the article is that humans are inherently more reliable than machines.

Humans hallucinate too, it’s just that when humans, misstate judgments, rely on summaries instead of full judgments, forget to update cases, and simply get sloppy under fatigue or time pressure, we call it “human error,” “an oversight,” or “carelessness.”

Suggesting that GenAI’s fallibility somehow makes it unfit for legal work ignores the very real fallibility of the humans using it.

5. The article treats GenAI as if the technology is frozen

The technology in this area is moving rapidly. But the article leans on now-dated hallucination studies and ignores current tool design such as:

(i) retrieval-augmented generation,

(ii) curated legal information,

(iii) citation validation,

(iv) hallucination suppression, and

(v) firm-level governance.

It’s like assessing today’s satellite GPS accuracy using the performance of the very first civilian receivers from the 1990s. You might as well warn drivers that “GPS can be off by 300 metres.” True then, but absurd now.

This article will likely only be formally published in 2026, making it even more out-dated.

If anything, this article highlights the danger of academic journals publishing in a rapidly changing area; especially if the article does not have solid empirical data and it does not provide a solid value proposition to the profession.

Conclusion

There is no “verification-value paradox.”

The article leans on an inflated belief in human infallibility, the rehashing of well-known critiques and long-standing sloppy legal habits, then wraps it all in a catchy phrase built on outdated assumptions and the misuse of the wrong tools.

Me:

I must disagree, using your numbering and headings:

1. Verification isn’t a paradox; it’s the lawyer’s job

But it becomes a paradox when the verification is of hallucinated output.

The author’s concentration is on research and not, as you say, transactional work. That is where the main problems arise and where the verification burden is a real problem.

With regard to starting with a blank page (which lawyers seldom do) see: https://thetimeblawg.com/2024/09/01/blank-page-law/

2. The uncomfortable truth: many litigators don’t read the cases upon which they rely

Presumably then it is those litigators that are relying on GenAI!

You say “However, he does not provide any data to support his view that verifying cases pulled by GenAI is slower than cases pulled and verified by other means.” He doesn’t have to. It is obvious. If you have a made up case, with made up citations, that does not exist, that you have to find, you are never going to find it. It could take you quite a while to ascertain that. Case citations in legal textbooks, case citators and existing case law will not present that problem.

The litigation practice you refer to is thankfully not one I am familiar with. I’ve been practising for 3+ decades and have never seen this. Maybe things are different in North America from the UK?

If GenAI has surfaced real problems in the profession that does not obviate the GenAI verification burden. It also clearly shines a light on that.

3. It is absurd to use general-purpose LLMs for legal research, then blame Gen AI when it goes wrong

It may be absurd but many lawyers are, unfortunately, doing it. New cases of this are arising by the day.

Accepting that it is negligence to use them accepts that the verification burden is too high to use them. The author was, in my view, generous with his view of cancelling out the benefits. In my opinion, the dangers of GenAI in legal research far outweigh any perceived benefits.

However, the author is looking at what is actually going on in practice. You cannot ignore that.

The author states “Studies have repeatedly found ‘hallucinations’ even in machine learning tools built for the legal context. Even where leading legal research companies like Westlaw and Lexis have built AI into their search functions, they remain unreliable.”

4. The author assumes humans are more accurate than technology. They aren’t.

I can’t agree that humans hallucinate like GenAI. A lawyer (pre GenAI) would find case law in the way trained to do so using established legal research methods. That resulted in them locating and reading real case law. If they use GenAI, without proper verification, they are relying on nonsense that they would never otherwise have found or have been bothered by.

You state “Suggesting that GenAI’s fallibility somehow makes it unfit for legal work ignores the very real fallibility of the humans using it.” However, you contradict yourself in that you have previously stated that it is unfit for legal research. On who is to blame see: https://thetimeblawg.com/2024/01/14/lawyers-or-machines-who-do-you-blame-for-genai-hallucinations/

5. The article treats GenAI as if the technology is frozen

Hallucinations are still a big problem. It has been suggested that they are getting worse, not better, with new models: https://www.newscientist.com/article/2479545-ai-hallucinations-are-getting-worse-and-theyre-here-to-stay/

RAG does not eliminate hallucinations and the verification burden remains. And obviously you need to ensure the knowledge base you are using contains the information that you need. The verification burden is as much about what will be omitted by GenAI than included in its output.

Conclusion

There is a “verification-value paradox” in certain situations. Not all. But many. It is dangerous to suggest otherwise. There has to be a realisation of what lawyers, on the whole, are actually doing. That is a real assumption and based on the tools that they are actually using.

[N.B. Mitch’s post on this and my response are reproduced above from: https://www.slaw.ca/2025/11/20/genai-the-verification-value-paradox-a-critique/]

Mitch Kowalski:

We need to continue this at a pub 😋😋

Me:

Happy to do so! 🍻

-+-+-+-+-+-

Antti Innanen (⚫ Ⓜ️ 💯):

New study, yes from vendor (Harvey & RSGI) but anyways maybe worth a look. Low number of participants, but from major players.

GenAI in Legal Practice - Harvey - Estimated Hours saved per month

-+-+-+-+-+-

Wessel Wijtvliet, PhD (AI and Legal Tech at Loyens & Loeff; Research Fellow at KU Leuven):

Of course it’s an important question, and it should receive an empirical answer, not a normative critique.

Antti Innanen (⚫ Ⓜ️ 💯):

Wessel Wijtvliet, PhD I want to see a study that proves ”the net value of AI use is often negligible to lawyers” 😅

Real empirical study of verification burden would of course be very welcome.

Me:

Looks like Brandon Barney is doing one.

Wessel Wijtvliet:

Keep sharing that info.

Dave CRAIG (Project Specialist for Stemovators: Prosper’s STEM Programme in Scotland):

Wessel Wijtvliet – Do you mean try it out and see if it works?

If so, I am not convinced.

You can cross a railway line blindfold many times successfully. This will never prove it is ‘safe’, because what we mean by ‘safe’ is less than say 1 chance a 1 billion of dying. So the experiment would need to run, in a repeatable and controlled way, 10 billion times to come to a valid conclusion.

If you are sending someone to jail for 10 years, a process that usually seems to get the right answer is not good enough.

Wessel Wijtvliet:

Not really, just that empirical statement should be based on empirical evidence, not normative assumptions. If you then conduct such research, the design itself should be such that it produces that are meaningful, including external validity. Surely coming up with the appropriate design is hard and requires creativity, but does not negate the fact that empirical statements should require empirical support.

Dave CRAIG:

OK, so what you are saying is that you want empirical evidence that:

[(research aided by AI) + validation]
costs more than
[research not aided by AI + validation]

Good luck with that. It obviously depends on what the case is.

But I support Brian’s point: you cannot assume that that AI will help, without considering validation.
And where validation is more important, AI is less likely to be a benefit. In some cases, AI may make things worse.

-+-+-+-+-+-

Anthony P. (Risk Mgmt, Taipei):

Thank you.
This is one of the most meaningful posts on this subject, followed by constructive discussion and some debate.

-+-+-+-+-+-

Dave CRAIG (Project Specialist for Stemovators: Prosper’s STEM Programme in Scotland):

Just get AI to do the verification, and take out insurance to cover liabilities arising from errors.
Compare the insurance premium with the cost of human paralegals.

Me:

If AI is causing the verification burden how is it then going to do the verification to correct its own hallucinations! I don’t think law firm’s PI insurers will provide such cover!

Anthony P. (Risk Mgmt, Taipei):

Dave CRAIG, ha ha, until the relevant exclusions are added to professional liability policies.

-+-+-+-+-+-

Alice Hewitt (Legal Information & Knowledge Management Champion | Discoverability & Accessibility Advocate | Information Literacy Tour Guide | Research Ninja | Interrogator of Mysteries (& databases & AI & IA & the why…):

I feel a bit honoured to be quoted in this!! Thanks Brian Inkster

-+-+-+-+-+-

Antti Innanen (⚫ Ⓜ️ 💯):

I feel like my position is pretty solid.

Yes, verification is important.
Yes, verification takes time.

No, it does not render the net value of AI use negligible to lawyers like the ”paradox” states.

I am done arguing this.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.