Gen AI v Verification - A robot sits using a laptop whilst a lawyer sits next to it verifying the output with a pile of legal text books
|

GenAI v Verification Debate – The Salient Points

The GenAI v Verification debate took place in November, after I blogged about the GenAI Verification Burden in Legal Practice. That blog post, following a heated debate on LinkedIn with those comments added to the post, ended up being 15,612 words long.

Most people won’t now read it. Our attention spans are too short in the instant dopamine hit world we currently live in.

I had GenAI summarise the post. Whilst this assists in giving a brief overview of the issues it is, not surprisingly, very bland and uninspiring. That is, even if, it did produce some graphs. Thus, it needs me to draw out some of the salient points in the GenAI v Verification debate. Which I will now do.

Proof of GenAI v Verification

Those arguing against the premise that the verification burden may outweigh the benefits of using GenAI, said there was no proof. They claimed that GenAI helps lawyers’ efficiency and effectiveness without providing any proof of this!

Empirical Answer v Normative Critique

It was suggested that “it’s an important question, and it should receive an empirical answer, not a normative critique”.

A counter point to this was that an empirical answer would be difficult to produce. This is because it will always depend on what the case is.

However, in the absence of an empirical study there is plenty of information to suggest that the verification burden (especially when applied to GenAI legal research) is high.

Reported cases of hallucinated citations

The main proof that GenAI is being used without verification causing real life problems for lawyers, their clients and the courts is in the number of reported cases of this. When I wrote my blog post on 15 November 2025 there were 545 such cases identified by Damien Charlotin  who tracks AI Hallucination Cases. Today (six weeks later) that figure is 712 cases. That is 28 new cases per week! There will, of course, be many more, that have not been identified/reported.

In those 712 cases no verification has been undertaken. How much time would have been required to verify and correct the hallucinations in each is a moot point. However, they demonstrate the real need for verification and the fact that the lawyers involved in those 712 cases would have been better off, in not verifying, to have not used GenAI in the first place as though it was magical. It ain’t. I will be doing a presentation on that at a Legal Tech event in 2026.

In a recent reported case in the United Arab Emirates a law firm, citing GenAI hallucinated cases, was ordered to pay AED 282,508 (GBP £56,990) to the opposing party in wasted costs. This demonstrates how costly the verification burden in legal practice can be.

We need more Lawyers to clean up after AI

In Australia an increase in lawyers being hired was reported in order to vet GenAI output:

The more artificial intelligence is used within a law firm, the more lawyers are needed to vet the technology’s outputs.

… The increased demand for legal professionals amid the increasing using of AI has been driven by strict rules around the use of AI-generated outputs by the courts and the tendency for the technology to fabricate answers.

Or as David Gerard, more bluntly, puts it:

The firms need proper lawyers to check all this chatbot spew and make sure it’s not lying nonsense.

It is no different from having an Associate

The old chestnut that GenAI is the equivalent of having an Associate was brought up in the GenAI v Verification debate. The argument being that verifying GenAI output is no different from verifying an Associate’s output. As I said:

The Associate probably won’t make stuff up (hallucinate) and their cites/sources will be easier/quicker to check. Your AI Copilot might well spit out fabricated nonsense and miss much that you would want included along the way. As a result you just might spend a lot longer checking the AI Copilot output than you would have to with the Associate output.

Furthermore, I pointed out then when an Associate makes mistakes:

they will hopefully learn from those mistakes as they are corrected by a supervisor. However, GenAI will not and you are always back to square one when using it.

We apparently have to accept that hallucinations are a fundamental part of machine learning when considering this comparison. I said:

Whilst hallucination may be a fundamental part of machine learning, I don’t think it has ever been a fundamental part of Associate learning!

As someone else put this part of the GenAI v Verification debate, perhaps, more succinctly:

Relying on an alien to discern human legal narratives seems like a very bad idea to me.

Concern was also expressed about an over reliance on GenAI, limiting the training of junior lawyers. And also questioning where the other skills required for the human-in-the-loop work will come from.

There is, of course, so much more that the junior lawyer can do that GenAI is unlikely to ever be able to do.

You should be using a Quality Legal AI vendor

It was argued that it’s much easier to verify output from a quality, legal AI vendor than from an Associate. This is apparently because documents and clauses are so readily accessible (hovering over or clicking on links). And when legal AI is applied to the correct use cases, a considerable net time savings results. I countered this with:

How do you measure “quality”. There are currently many #LegalTech AI vendors and new ones appearing on the block all the time. Many lawyers are using more readily available GenAI tools often not realising the issues associated with those. Specific LegalTech AI tools are not without problems as research cited by me demonstrates. Before GenAI, LegalTech tools were available to research case law. These were possibly more reliable than they are now with a GenAI layer added. Just because you can doesn’t mean you need to. You would hope a good Associate would cite their sources (even hyperlink them in this day and age). If not, they have not been well trained. And what are the “correct use cases”?

There was no response.

Furthermore many GenAI products ‘designed’ for legal work are general LLMs in a wrapper. The paper by Joshua Yuvaraj, that started the GenAI v Verification debate, points out that:

They [structural flaws inherent to AI technology] are also structural features of both general-purpose AI services (that members of the public and/or organisations can use for any purpose, and purchase escalating tiers of service for, like ChatGPT, Microsoft Copilot, Gemini), and bespoke legal profession-targeted services (e.g. CoCounsel, Harvey AI). This distinction matters because claimed points of difference between the latter and former categories is their reliability and reliance on high-quality data for lawyers. Such claims do not, however, address the fundamental structural flaws inherent to machine learning models outlined below.

It depends what you are using GenAI for

It was stated that the trade-off is different for everyone and depends on what they’re using AI for. Of course it does. However, the GenAI v Verification debate was principally about using GenAI for legal research. That is a task that it is clearly not designed for. It is also one that it is no use at (like many tasks it is wrongly being used for).

I saw this on LinkedIn:

I recently heard about a teacher who instead of trying to circumvent students using AI, which is impossible, she made assignments like “ask ChatGPT to write a report on this subject, and then research how and why it’s wrong.” Not only did the students discover that ChatGPT is extremely wrong much of the time, it also led them to realize that they should not use it as a primary source.

Lawyers would do well to carry out the same exercise.

Hallucinations are only an issue with GenAI if you don’t know how to use it

The argument that hallucinations are only an issue with GenAI if you don’t know how to use it comes from those lawyers that are apparently superior in their knowledge than all other mere mortals. They may have found a way of working with it to their benefit. However, I doubt that is the case when it comes to legal research. They also appear shy to share what this may be.

When pressed, it might be using it on specific known sets of documents as a search tool. Well, yes, but that is not what the GenAI v Verification debate was about.

Also, your own experiences if honed through much use and experimentation cannot be compared with most lawyers dabbling in it with little knowledge of the risks involved. Furthermore, most of those lawyers will not have the time nor inclination to train in expert prompt engineering. Even though, if they did, that would not eliminate hallucinations. They would still be better off becoming legal hallucinatory detectorists.

When the ‘experts’ were asked whether they had written in detail anywhere about their methodology there was a deafening silence.

It was also pointed out that:

What often gets overlooked is that verification isn’t just about catching hallucinations – it’s about catching misframings. Even when the facts are correct, the model’s implicit assumptions can quietly distort the legal position.

I added:

Also, we can’t forget about what it misses out completely. That is a huge part of the verification burden.

It was touched upon (but this is a big one) that lawyers are having to spend longer, than would otherwise be the case, verifying GenAI generated hallucinatory content produced by other lawyers and by their clients.

GenAI technology is not frozen

It was suggested that the technology in this area is moving rapidly. However, it is not rapidly eliminating hallucinations. Hallucinations are still a big problem. It has been suggested that they are getting worse, not better, with new models. Furthermore, the number of people using AI at work is suddenly falling. This would suggest that the technology is not moving rapidly enough to be actually useful in business. Indeed, it was recently reported that Microsoft has scaled back its AI goals because almost nobody is using Copilot.

Conclusion in the GenAI v Verification Debate

The debate, in my opinion (but I am of course biased), did not defeat my argument, or Joshua Yuvaraj’s argument, that the verification burden will often outweigh the benefits of using GenAI for legal tasks.

As someone, in the GenAI v Verification debate, nicely put it:

You cannot assume that AI will help, without considering validation. And where validation is more important, AI is less likely to be a benefit. In some cases, AI may make things worse.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.