Top large language models struggle to make accurate legal arguments
- Reference: 1704889810
- News link: https://www.theregister.co.uk/2024/01/10/top_large_language_models_struggle/
- Source link:
Last year, when OpenAI showed [1]GPT-4 was capable of passing the Bar Exam, it was heralded as a breakthrough in AI and led some people to question whether the technology could soon [2]replace lawyers. Some hoped these types of models could empower people who can't afford expensive attorneys to pursue legal justice, making access to legal help more equitable. The reality, however, is that LLMs can't even assist professional lawyers effectively, according to a recent study.
The biggest concern is that AI often fabricates false information, posing a huge problem especially in an industry that relies on factual evidence. A team of researchers at Yale and Stanford University analyzing the rates of hallucination in popular large language models found that they often do not accurately retrieve or generate relevant legal information, or understand and reason about various laws.
[3]
In fact, OpenAI's GPT-3.5, which currently powers the free version of ChatGPT, hallucinates about 69 percent of the time when tested across different tasks. The results were worse for PaLM-2, the system that was previously behind Google's Bard chatbot, and Llama 2, the large language model released by Meta, which generated falsehoods at rates of 72 and 88 percent, respectively.
[4]
[5]
Unsurprisingly, the models struggle to complete more complex tasks as opposed to than easier ones. Asking AI to compare different cases and see whether they agree upon an issue, for example, is challenging, and it will more likely generate inaccurate information than when faced with an easier task, such as checking which court a case was filed in.
Although LLMs excel at processing large amounts of text, and can be trained on huge amounts of legal documents – more than any human lawyer could read in their lifetime – they don't understand law and can't form sound arguments.
[6]
"While we've seen these kinds of models make really great strides in forms of deductive reasoning in coding or math problems, that is not the kind of skill set that characterizes top notch lawyering," Daniel Ho, co-author of [7]the Yale-Stanford paper and a law professor at Stanford, tells The Register .
"What lawyers are really good at, and where they excel is often described as a form of analogical reasoning in a common law system, to reason based on precedents."
Machines often fail in simple tasks too. When asked to inspect a name or citation to check whether a case is real, GPT-3.5, PaLM-2, and Llama 2 can make up fake information in responses.
[8]
"The model doesn't need to know anything about the law honestly to answer that question correctly. It just needs to know whether or not a case exists or not, and can see that anywhere in the training corpus," Matthew Dahl, a PhD law student at Yale University, says.
[9]Supreme Court supremo ponders AI-powered judges, concludes he's not out of a job yet
[10]Law secretly drafted by ChatGPT makes it onto the books
[11]Lawyers who cited fake cases hallucinated by ChatGPT must pay
[12]'Robot lawyer' DoNotPay not fit for purpose, alleges complaint
It shows that AI cannot even retrieve information accurately, and that there's a fundamental limit to the technology's capabilities. These models are often primed to be agreeable and helpful. They usually won't bother correcting users' assumptions, and will side with them instead. If chatbots are asked to generate a list of cases in support of some legal argument, for example, they are more predisposed to make up lawsuits than to respond with nothing. A pair of attorneys learned this the hard way when they were [13]sanctioned for citing cases that were completely invented by OpenAI's ChatGPT in their court filing.
The researchers also found the three models they tested were more likely to be knowledgeable in federal litigation related to the US Supreme Court compared to localized legal proceedings concerning smaller and less powerful courts.
Since GPT-3.5, PaLM-2, and Llama 2 were trained on text scraped from the internet, it makes sense that they would be more familiar with the US Supreme Court's legal opinions, which are published publicly compared to legal documents filed in other types of courts that are not as easily accessible.
They also were more likely to struggle in tasks that involved recalling information from old and new cases.
"Hallucinations are most common among the Supreme Court's oldest and newest cases, and least common among its post-war Warren Court cases (1953-1969)," according to the paper. "This result suggests another important limitation on LLMs' legal knowledge that users should be aware of: LLMs' peak performance may lag several years behind the current state of the doctrine, and LLMs may fail to internalize case law that is very old but still applicable and relevant law."
Too much AI could create a 'monoculture'
The researchers were also concerned that overreliance on these systems could create a legal "monoculture." Since AI is trained on a limited amount of data, it will refer to more prominent, well-known cases leading lawyers to ignore other legal interpretations or relevant precedents. They may overlook other cases that could help them see different perspectives or arguments, which could prove crucial in litigation.
"The law itself is not monolithic," Dahl says. "A monoculture is particularly dangerous in a legal setting. In the United States, we have a federal common law system where the law develops differently in different states in different jurisdictions. There's sort of different lines or trends of jurisprudence that develop over time."
"It could lead to erroneous outcomes and unwarranted reliance in a way that could actually harm litigants" Ho adds. He explained that a model could generate inaccurate responses to lawyers or people looking to understand something like eviction laws.
"When you seek the help of a large language model, you might be getting the exact wrong answer as to when is your filing due or what is the kind of rule of eviction in this state," he says, citing an example. "Because what it's telling you is the law in New York or the law of California, as opposed to the law that actually matters to your particular circumstances in your jurisdiction."
The researchers conclude that the risks of using these types of popular models for legal tasks is highest for those submitting paperwork in lower courts across smaller states, particularly if they have less expertise and are querying the models based on false assumptions. These people are more likely to be lawyers, who are less powerful from smaller law firms with fewer resources, or people looking to represent themselves.
"In short, we find that the risks are highest for those who would benefit from LLMs most," the paper states. ®
Get our [14]Tech Resources
[1] https://www.theregister.com/2023/03/14/openai_gpt4_ai/
[2] https://www.theregister.com/2023/01/16/in_brief_ai/
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZZ7NOhEIf6kVi0iAxoP7RgAAABE&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZZ7NOhEIf6kVi0iAxoP7RgAAABE&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZZ7NOhEIf6kVi0iAxoP7RgAAABE&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZZ7NOhEIf6kVi0iAxoP7RgAAABE&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[7] https://arxiv.org/abs/2401.01301
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZZ7NOhEIf6kVi0iAxoP7RgAAABE&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[9] https://www.theregister.com/2024/01/02/supreme_court_ai_roberts/
[10] https://www.theregister.com/2023/12/02/chatgpt_law_brazil/
[11] https://www.theregister.com/2023/06/22/lawyers_fake_cases/
[12] https://www.theregister.com/2023/03/13/robot_lawyer_donotpay_back_in/
[13] https://www.theregister.com/2023/06/22/lawyers_fake_cases/
[14] https://whitepapers.theregister.com/
Re: Reason
It's not even as simple as the AI being fed garbage data and not filtering it out, but sometimes the AI being fed applicable data and not being able to determine when it is applicable and when it is not. Admittedly, I've seen humans fail that test as well, but they're usually a bit better at it. For example, a person does a search for a legal issue and gets results that describe, accurately, the process for dealing with that issue in a place they're not in. The location where it applies will be written in that article, and most people will find that out and try to find another article. Language models will probably fail to correlate that mention of the location with all the words further along in the article and, if what it says is common enough, give it to anyone who asks about the issue, even if that person specifically mentioned a different location. It got correct data and nonetheless generates garbage. That is what an LLM does, and the sooner people realize that, the fewer idiots they will make of themselves.
Re: Reason
Most importantly, even if AI is fed correct and applicable data it might spew hallucinations - as stated in the article. This may even be an intrinsic protection against copyright theft accusations (not that it's working, as we've seen in cases where AI delivered textually identical content to what it was fed, like in the [1]NY Times case )
[1] https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html
"they don't understand law and can't form sound arguments"
Substitute 'anything' for 'law' and that will sum up the reality.
The anthropomorphic terminology that its proponents use in relation to AI in general (not just LLMs) -- 'understand', 'hallucinate' etc. -- belies that objective fact that these machines do not think at all. They have no correlates of human cognitive processes, they're just context-driven Markov chain generators drawing from frequency-weighted repositories of tokens. So any concept of 'meaning' is irrelevant, and apparently meaningful (let alone accurate) output is fortuitous.
Re: "they don't understand law and can't form sound arguments"
The one thing that is more specific to law than to other areas though is that it can be changed.
You could have millions of cases and other reputable legal texts that genuinely support one particular argument, but if your parliament or equivalent passed a new law last week that contradicts it, then that takes precedent.
Re: "they don't understand law and can't form sound arguments"
Absolutely! There is no AI or ML, just statistics and probability.
> Top large language models struggle to make accurate ...
Hold the Front Page! Statistical bullshit generator fails to generate accurate bullshit ...
On other pages: Grizzly bear fails to use public conveniences.. Pope Francis declines to attend wiccan nude mud-wrestling festival at solstice..
The issue I'd have with this research is that it's based on GPT-3.5, when it was GPT4 that demonstrated the capability to pass the bar exam.
Of course the version of GPT that failed the bar exam is bad at legal shiz, GPT3.5 has been shown multiple times to be incapable of regularly performing well enough to pass the exam. scoring in the bottom 10% of participants: https://openai.com/research/gpt-4
Would be far more interested in the results from GPT4, which OpenAI claim scores in the top 10% of test takers.
Great, the bullshit from GPT-4 is even more plausible.....
Doesn't change any of the arguments above, this is still just a markov thingy with slightly better probabilities. It's no closer to "understanding" than a brick is.
Who are you trying to kid? Us or yourself?
I put it to you
Interesting, this. I watched a youtube video recently where some amateur Go player managed to defeat an AI bot that had defeated the world's top Go player simply by exploiting the fact that the bot didn't know or care that it was playing Go. So if it's Matlock vs. LA Law AI Law, all Matlock needs to do is to figure out what the exploitable flaws are in the bot that the lazy legal eagles from LA Law are using and it should be possible to out-law them in every case.
Snake Oil doesn't work
I thought we all knew that already.
It is mightily impressive they have managed to get a computer to emulate the bullshit artist from every pub but I don't see how anyone thought that would be useful, or is surprised that it isn't.
Quite human
It's an interesting "feature" of LMLMs that instead of answering with "I don't know", they simply reply with made-up bullshit. Similar to (sadly) many people.
Re: Quite human
Similar to (sadly) many people search engines.
So a LLM made by a search engine corporation can be expected to be particularly bad.
I'd expect this to be an application at which ML would be particularly bad.
From my experience of listening to legal arguments in court, usually about whether something is admissible as evidence, they seem to hinge on decisions made in a given set of circumstances and how the current, novel set of circumstances, can be considered as equivalent or near enough so for the same decision to apply vs whether they're sufficiently different that it doesn't. Apart from a need for logic, judgement and ability to put things persuasively the ML is at an obvious disadvantage in relation to its material. It will have encountered the previous circumstances in the training material but the key adjective above was "novel"; it won't have encountered them before and without having the understanding of both, won't have any means of relating the two. No wonder it serves up some random response.
Reason
The AI simply recycles biases found in the training material.
Whereas intelligent person can most of the time tell they are being served BS. AI will just repeat it, won't challenge it.
AI may find holes in the legislation if you nudge it towards it. But it's a bit like leading a horse to the water.