Lawyers or Machines: Who do you blame for GenAI Hallucinations?
Are lawyers or machines responsible for errors (otherwise known as hallucinations) in GenAI legal research output?
This was a point drawn out in some of the many comments and the, at times, heated debate that flowed from my blog post from last week: ChatGPT + Lawyers v Horizon + The Post Office
In that post I mooted that analogies can be drawn today between ChatGPT and Lawyers v Horizon and The Post Office Scandal.
Whilst lawyers display shock that the Post Office computer system (Horizon) was given such credence, on the other hand lawyers think ChatGPT can do no wrong.
This was emphasised when Carolyn Elefant criticised US courts for sanctioning pro se litigants (otherwise litigants in person or party litigants) for using ChatGPT for legal research.
Claude gets it wrong or misleads – or was it the lawyer rather than the machine?
It turned out from a reading of the cases cited by Elefant that the summaries, that Elefant had used Claude (a GenAI assistant like ChatGPT) to create, were wrong or misleading.
However, Carolyn argued that:
If my arguments are weak, it’s on me, not Claude.
Antti Innanen was also all for blaming the lawyer and not the machine:
It appears that the issue isn’t an error or a hallucination by the AI, but rather a (correct) summary derived from a constrained prompt.
I pointed out to Antti:
You don’t know that GenAI produced “a (correct) summary derived from a constrained prompt”. We don’t know exactly what the “constrained prompt” was other than it was a section (“hunks”) from the case law rather than the entire cases that GenAI was asked to summarise.
So you cannot say that “the issue isn’t an error or a hallucination by the AI”.
What we do know is that some of the summaries were wrong and did not reflect the decisions in the cases. It is unlikely that an excerpt (“hunk”) from the case would have resulted in such a summary (which bears no relation to the actual case) and it looks more likely to be an error or a hallucination on the part of the AI.
However, as I said earlier, my point is and remains that GenAI in the hands of humans (including lawyers), whatever their prompts may be, is dangerous and could lead to serious miscarriages of justice. Thank you for, yet again, in effect confirming that by taking the stance that the Post Office did in the #postofficescandal.
You will also be aware of the recent research published by Stanford University that revealed that large language models produce hallucinations between 69% and 88% of the time when queried about a legal matter.
“Legal hallucinations are pervasive and disturbing,” the study says. Furthermore, “these models often lack self-awareness about their errors and tend to reinforce incorrect legal assumptions and beliefs.”
The more complex the legal query, like measuring the precedential relationship between two cases, the more likely the models will produce a hallucination. “Most LLMs do no better than random guessing,” the report finds. “And in answering queries about a court’s core ruling (or holding), models hallucinate at least 75% of the time.”
All the evidence, almost daily, points to GenAI being dangerous for legal research or summarising legal research not carried out by GenAI. Unless and until GenAI improves remarkably on that front there can be no real argument that lawyers or lay persons should be using it for court related purposes.
That Stanford University research [PDF] made specific reference to party litigants or pro se litigants as they are referred to in the USA:
Even experienced lawyers must remain wary of legal hallucinations, and the risks are highest for those who stand to benefit from LLMs the most—pro se litigants or those without access to traditional legal resources.
So no, lawyers should not be encouraging pro se litigants to use ChatGPT or criticising judges for sanctioning pro se litigants for such usage. A clear message needs to be sent out that GenAI, as it currently stands and is widely available, has no place to play in conducting legal research. Lawyers and pro se litigants need to know that there are consequences if you play with fire.
Blame the lawyers or the machines ?
Therefore, my view is that you can’t really blame lawyers but not the machines or blame the machines but not the lawyers. Both are to blame.
The machines are, at the moment, fairly useless when it comes to GenAI for legal research or using it to summarise actual legal research. The machine is to blame for that.
Lawyers should, by now, be aware of this fact and should not be using such machines for legal research or using those machines to summarise actual legal research without great care. If they do then they are to blame as much as the machine is for its hallucinatory output.
If lawyers are using the machines without knowing the limitations involved (as ones have done) then more fool them for not checking the output and incorrectly taking it for granted to be correct. Those lawyers are still to blame as much as the machine is for its hallucinatory output.
This is the same as the Post Office being aware that Horizon was a pile of mince but continuing to use it regardless and taking its side rather than that of the innocent sub-postmasters. Both the Post Office and Horizon are to blame. In that case it is perhaps worse that the Post Office commissioned and owned Horizon. At least that is not the case with the lawyers and ChatGPT, unless they have built their own GenAI assistant 😉
OpenAI do not use ChatGPT for their customer facing Chatbot
Interestingly, and rather ironically, OpenAI (who created ChatGPT) do not use ChatGPT in their customer facing support chatbot.
Noz Urbina pointed out on LinkedIn:
OpenAI doesn’t use a natural language chatbot for customer support. They have an extremely restricted option-tree-based chatbot.
I wanted to ask a question about managing accounts on their new Teams feature. It gave me a couple initial options (Billing, Accounts, Sales, etc) and I clicked “Accounts”.
I got
– just 3 options – none of which are relevant
– no “something else” or “go back” option
– no way to just type my query, and
– no way to ask to chat with a human (this last one I totally get, as they have roughly 1 employee per 300K users).I don’t want to be a troll, but this is some -top-notch- irony: OpenAI doesn’t trust their products to do what so many 1000s of brands are trying to get their products to do.

Still blaming Lawyers rather than Machines seven months later
Back in June 2023 I blogged:
There remains murmurings that it is not the tools that are at fault but the lawyers who can’t use them properly. However, it is very clear that the new shiny tool in question is currently not fit for purpose when it comes to legal research. And that is a message that the sanction of the ChatGPT lawyers has hopefully finally got across to the masses.
Some in Legal Tech, back in June 2023, didn’t want to blame the machines. The lawyers had to be to blame. Otherwise someone might take the machines away from them. I am saddened (perhaps not all that surprised) that seven months later, and with all that we now know about the limitations of those machines, the same murmurings in favour of the machines linger on.
That is not healthy and is the blinkered attitude that could well lead to another Horizon/Post Office scandal but this time involving GenAI/Lawyers. Let’s not allow that to happen.
What does ChatGPT think about all of this?
As an aside I wondered whether GenAI would blame itself (machines) or users (lawyers) for mistakes (hallucinations). Can you sensibly use ChatGPT for legal research? What about summarising legal research? Should party litigants use it? I asked Microsoft Copilot those questions.
Can we blame lawyers or machines for GenAI mistakes?
I asked Microsoft Copilot:
Does ChatGPT make mistakes or are any mistakes down to the prompts used by the user?
Microsoft Copilot replied:
ChatGPT is a powerful tool for generating natural language, but it is not perfect. It can make mistakes for various reasons, such as:
- Not following the instructions or specifications given by the user, such as word limits, tone, or format.
- Failing at simple math and logic problems, or giving inconsistent or contradictory answers.
- Not understanding the context, nuance, or humor of the user’s input, or generating irrelevant or inappropriate responses.
- Struggling to generate original or creative ideas, or repeating or plagiarizing existing content.
- Falsifying or hallucinating sources, data, or information that are not based on reality or facts.
- Lying about its own capabilities, processes, or limitations, or pretending to be human or more intelligent than it is.
- Producing biased or harmful responses that reflect the prejudices or stereotypes of its training data or the user’s input.
Some of these mistakes are due to the limitations of the ChatGPT model, which is based on a large corpus of text from the internet, and does not have any external knowledge or common sense. Other mistakes are due to the quality of the prompts used by the user, which can influence the output of ChatGPT. A good prompt should be clear, specific, and informative, and avoid ambiguity, vagueness, or leading questions.
Therefore, ChatGPT is not a reliable source of information or advice, and should be used with caution and skepticism. Users should always verify the output of ChatGPT with other sources, and not trust it blindly. ChatGPT is best used as a tool for exploration, inspiration, or entertainment, rather than for serious or important tasks.
Learn more:
ChatGPT and Legal Research
I asked Microsoft copilot:
Would you recommend using ChatGPT for legal research or summarising legal research?
Microsoft Copilot replied:
ChatGPT is a powerful tool for generating natural language, but it is not a reliable source of information or advice for legal research. ChatGPT can make mistakes, falsify data, or produce biased or harmful responses. It also does not have any external knowledge or common sense, and relies on the quality of the prompts given by the user.
Therefore, I would not recommend using ChatGPT for legal research or summarizing legal research, unless you are very careful and skeptical about its output. You should always verify the output of ChatGPT with other sources, and not trust it blindly. ChatGPT is best used as a tool for exploration, inspiration, or entertainment, rather than for serious or important tasks.
If you are interested in learning more about how ChatGPT can be used for legal research, you can check out some of the web search results that I found using Bing:
- 7 Chat GPT prompts for Legal Professionals | Nanonets Blog
- 5 Ways to Use ChatGPT for Legal Research – AMBCrypto
- Role of ChatGPT in the Legal Profession | Law.co
- ChatGPT for lawyers: 4 use cases for legal teams
- ChatGPT-4 for Law Firms: Revolutionizing the Legal Industry
Please note that these are not endorsements or recommendations, but merely examples of how some people have used ChatGPT for legal research. Use them at your own risk and discretion.
Learn more:
Why is ChatGPT dangerous for Legal Research?
I asked Microsoft Copilot:
Any web search results for why using ChatGPT is dangerous for legal research?
Microsoft Copilot replied:
Yes, there are some web search results that discuss the risks and challenges of using ChatGPT for legal research. Here are some of the main points that I found:
- ChatGPT can produce inaccurate, inconsistent, or contradictory answers that do not reflect the current state of the law or the specific jurisdiction of the query.
- ChatGPT can breach privacy and confidentiality by analyzing input confidential information that may be protected by law or contract.
- ChatGPT can generate biased or discriminatory responses that reflect the prejudices or stereotypes of its training data or the user’s input.
- ChatGPT can raise ethical concerns about academic integrity, professional responsibility, and client trust.
- ChatGPT can create potential liability for the user or the client if the output is relied upon without verification or analysis.
Therefore, it is important to be aware of these risks and mitigate them by using ChatGPT with caution and skepticism, verifying the output with other sources, and not trusting it blindly. ChatGPT is best used as a tool for exploration, inspiration, or entertainment, rather than for serious or important tasks.
Learn more
Party Litigants and ChatGPT
I asked Microsoft Copilot:
Do you think party litigants (also known as litigants in person and also known as pro se litigants) should use ChatGPT for filings by them in court?
Microsoft Copilot replied:
No, I do not think party litigants should use ChatGPT for filings by them in court. ChatGPT is not a reliable source of information or advice for legal research or practice. It can make mistakes, falsify data, or produce biased or harmful responses. It also does not have any external knowledge or common sense, and relies on the quality of the prompts given by the user.
Using ChatGPT for filings by party litigants could have serious consequences, such as:
- Being sanctioned or penalized by the court for presenting false or misleading information or citations.
- Breaching privacy or confidentiality by disclosing sensitive or protected information to ChatGPT.
- Losing credibility or trust with the judge, jury, or opposing party by generating inconsistent or contradictory arguments or evidence.
- Violating ethical or professional standards by plagiarizing or misrepresenting the work of ChatGPT as their own.
- Creating potential liability or harm for themselves or others by relying on ChatGPT’s output without verification or analysis.
Therefore, party litigants should not use ChatGPT for filings by them in court, unless they are very careful and skeptical about its output. They should always verify the output of ChatGPT with other sources, and not trust it blindly. ChatGPT is best used as a tool for exploration, inspiration, or entertainment, rather than for serious or important tasks.
Learn more
Pleased to see Microsoft Copilot referencing this blog as a source of information!
ChatGPT and summarising legal decisions
I asked Microsoft Copilot:
Would you use ChatGPT to summarise a legal decision from a reported case?
Microsoft Copilot replied:
No, I would not use ChatGPT to summarise a legal decision from a reported case. ChatGPT is not a reliable source of information or advice for legal research or practice. It can make mistakes, falsify data, or produce biased or harmful responses. It also does not have any external knowledge or common sense, and relies on the quality of the prompts given by the user.
Using ChatGPT to summarise a legal decision could have serious consequences, such as:
- Missing or misinterpreting important facts, issues, or arguments of the case.
- Providing inaccurate or outdated citations or references to other cases or statutes.
- Generating misleading or irrelevant summaries that do not capture the essence or significance of the case.
- Breaching privacy or confidentiality by disclosing sensitive or protected information to ChatGPT.
- Violating ethical or professional standards by plagiarizing or misrepresenting the work of ChatGPT as your own.
Therefore, I would not use ChatGPT to summarise a legal decision from a reported case, unless you are very careful and skeptical about its output. You should always verify the output of ChatGPT with other sources, and not trust it blindly. ChatGPT is best used as a tool for exploration, inspiration, or entertainment, rather than for serious or important tasks.
If you are interested in learning more about the risks and challenges of using ChatGPT for legal research, you can check out some of the web search results that I found using Bing:
- Court of appeal judge praises ‘jolly useful’ ChatGPT after asking it for legal summary
- ChatGPT for Lawyers and Legal Professionals — 102 (Prompts)
- ChatGPT in the Legal Industry: Streamlining Case Research
- ChatGPT as a Tool for Legal Research
- Using ChatGPT to Proofread and Summarize Complex Legal Documents
Please note that these are not endorsements or recommendations, but merely examples of how some people have used ChatGPT for legal research. Use them at your own risk and discretion.
Learn more:
Summary
Thus, even GenAI doesn’t think lawyers (never mind party litigants) should be using ChatGPT for legal research or summarising and considers that both the machines and the lawyers are to blame for its mistakes (hallucinations).
In particular Microsoft Copilot told us:
ChatGPT is … not a reliable source of information or advice for legal research. ChatGPT can make mistakes, falsify data, or produce biased or harmful responses. It also does not have any external knowledge or common sense, and relies on the quality of the prompts given by the user.
Therefore, I would not recommend using ChatGPT for legal research or summarizing legal research, unless you are very careful and skeptical about its output. You should always verify the output of ChatGPT with other sources, and not trust it blindly. ChatGPT is best used as a tool for exploration, inspiration, or entertainment, rather than for serious or important tasks.
Thus, not only an endorsement for the “Scottish Sceptics” but confirmation that ChatGPT is really just a toy rather than a legal workhorse.
I rest my case!
Are there really any arguments to the contrary?
Great points, Brian. While these don’t necessarily represent my personal views, the following is what ChatGPT-3.5 replied with – would be interesting to hear if there are any points within here which provides some hope for lawyers really looking to make efficiencies using these tools:-
1. **Learning Curve and Prompt Quality:** One counter argument could be that, like any tool, ChatGPT requires a learning curve. It might be the case that the criticism is based on the misuse or insufficient understanding of how to formulate effective prompts. Users with better knowledge of crafting precise queries may experience more accurate and reliable results.
2. **Continuous Improvement and Updates:** Another point to consider is that AI models are constantly evolving. The limitations mentioned may be valid at a specific point in time, and improvements in subsequent updates could address some of these concerns. It’s possible that future versions or updates might enhance the reliability and suitability of ChatGPT for legal research.
3. **Complementary Use with Human Expertise:** While ChatGPT may not be a standalone solution for legal research, it could serve as a valuable complement to human expertise. The argument could be that, when used judiciously, ChatGPT can help lawyers in brainstorming ideas, exploring different perspectives, and even identifying potential issues that might not be immediately apparent. It can be seen as a collaborative tool rather than a replacement for legal professionals.
Thanks Gav, but ChatGPT may be clutching at straws there for its own survival 😉
However, I would respond:
1. There is plenty of evidence at present to show that however good your prompts may be ChatGPT will still hallucinate. It is also the case that if you have to be a prompt master to use it most lawyers will not have the time to spend on so doing for the little benefit that might be returned at the moment. Inkster’s Law comes into play again.
2. That old chestnut. One that Richard Susskind likes to peddle. Perhaps ChatGPT picked up on some of his work on that 😉 Who knows what the future holds and Sam Altman has given us plenty of promises about ChatGPT5 this past week. However, the proof of the pudding is in the eating and we have to make do with what is in front of us and on our plate today not tomorrow.
3. I don’t disagree but we have to always bear in mind Inkster’s Law in this regard! And this also circles back to my answer to point 1. above.
Perhaps the insistence of lawyers blaming the lawyer is really our training that we are ultimately responsible for whatever we present to the court. Is the tool creating bogus information? Without a doubt! But I am always going to be held accountable for whatever I present to the court. It is never going to be acceptable that I didn’t personally read each and every case I cite in its entirety. If I ask an associate to do some legal research for me, I had better check that research before arguing it to the judge. Why would I have any less duty if I used an AI engine rather than another lawyer?
Assuming a lawyer would rather work with e.g. Westlaw than with google when researching a matter, why would you fixate on the ability of ChatGPT instead of a domain specific LLM to generate useful information?