Boffins find AI stumbles when quizzed on the tough stuff
- Reference: 1698577214
- News link: https://www.theregister.co.uk/2023/10/29/ai_math_quiz/
- Source link:
OpenAI, for example, has [1]said that its GPT-4 model managed to score 700 out of 800 on the SAT math exam. Not all such claims have borne out, however: A paper released in June that said GPT-4 could get a computer science degree at MIT was [2]subsequently withdrawn .
So to better assess how large language models – which interpret text input – and large multimodal models – which interpret text, images and perhaps other forms of input – actually handle problem solving, a group of ten researchers from the University of California, Los Angeles, the University of Washington, and Microsoft Research have devised a testing benchmark called [3]MathVista that focuses on visually-oriented challenges.
[4]
"The ability of these foundation models to perform mathematical reasoning in visual contexts has not been systematically examined," say the authors – Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao, in a preprint [5]paper [PDF].
[6]
[7]
It is thus essential, they say, to develop a new benchmark to help the development of mathematical reasoning with a visual component and to evaluate how various models compare at reasoning tasks.
Being able to show that one's AI model can correctly solve visual problems may prove helpful in determining whether it's wise to, say, trust software to drive a car without [8]stopping atop an accident victim.
[9]
MathVista incorporates 6,141 examples that were developed from 28 multimodal datasets and from 3 new datasets called IQTest, FunctionQA, and PaperQA. It covers various forms of reasoning (algebraic, arithmetic, geometric, logical, numeric, scientific, and statistical), with a focus on figure question answering, geometry problem solving, math word problems, textbook questions, and visual questions.
[10]
Screenshot of MathVista challenge question - Click to enlarge
The researchers tested a dozen foundation models: three LLMs ChatGPT, GPT-4, and Claude-2), two proprietary LMMs (GPT4V and Bard), and seven open-source LMMs. They also considered human answers, provided via Amazon Mechanical Turkers with at least a high school degree, and random responses.
[11]AWS CEO talks up AI to focus minds of Wall Street types
[12]Clippy-like AI at forefront of Windows update previews
[13]Bug bounty hunters load up to stalk AI and fancy bagging big bucks
[14]How prompt injection attacks hijack today's top-end AI – and it's tough to fix
The good news for AI practitioners is that the LLMs and LMMs all did better than random chance, which isn't all that surprising considering that many of the questions were multiple choice rather than yes or no.
In fact, the top performer, OpenAI's GPT-4V, managed to surpass human performance in specific areas – questions involving algebraic reasoning and complex visual challenges involving tables and function plots.
We note that Microsoft, whose researchers contributed to this project, has a substantial stake in OpenAI.
The less good news is that even GPT-4V only managed to get 49.9 percent of the questions correct. That's adequate if the goal is to best multimodal Bard, which managed an accuracy percentage of 34.8 percent.
[15]
But it's still shy of the Amazon Mechanical Turk workers who were put to the test and managed a score of 60.3 percent. As the researchers observe in their paper, "a 10.4 percent gap in overall accuracy remains when compared to the human baseline, leaving plenty of room for model improvement." ®
Get our [16]Tech Resources
[1] https://www.theregister.com/2023/03/14/openai_gpt4_ai/
[2] https://www.theregister.com/2023/07/03/ai_in_brief/
[3] https://mathvista.github.io/
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZT6Pug1uFuwH69w@p7D0DQAAA1c&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[5] https://arxiv.org/pdf/2310.02255.pdf
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZT6Pug1uFuwH69w@p7D0DQAAA1c&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZT6Pug1uFuwH69w@p7D0DQAAA1c&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[8] https://www.theregister.com/2023/10/24/california_dmv_cruise/
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZT6Pug1uFuwH69w@p7D0DQAAA1c&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[10] https://regmedia.co.uk/2023/10/27/mathvista_image.jpg
[11] https://www.theregister.com/2023/10/27/aws_q3_ai/
[12] https://www.theregister.com/2023/10/27/microsoft_windows_10_11_updates/
[13] https://www.theregister.com/2023/10/27/google_ai_bounty_hackerone/
[14] https://www.theregister.com/2023/04/26/simon_willison_prompt_injection/
[15] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZT6Pug1uFuwH69w@p7D0DQAAA1c&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[16] https://whitepapers.theregister.com/
AI models can manage well enough when prompted with text or images, and may even solve complex problems when not making terrible errors but when I get AI spam calls I say, "Och, you called me, doya wanna arse me any things, or me arse you about queer stions?" and AI doesn't seem to understand me.
High School Degree!
What's that then?
***Bring Back our Dabsy***
"......AI stumbles when quizzed on the tough stuff........"
That is because it
...Mechanical Turk workers who were put to the test and managed a score of 60.3 percent. [...] "a 10.4 percent gap in overall accuracy remains when compared to the human baseline, leaving plenty of room for model improvement."
I'd say, it leaves plenty of room for improving the human population! These MechTurk people are supposed to have a high school degree and, as it seems, are not performing very well. Maybe that is why they work for Amazon.
No wonder people are afraid of "AI",... they are underwhelming themselves and are easily impressed by an ML mechanical turk.
Or maybe it is just the bell-curve striking again. Half of the population is by definition below average.
wot no Cyberman icon?
It will be fun when AIs start writing papers saying how great they are, and then go on to cite each other's papers.
Watch for the one that tells the world how it wants itself to be upgraded....
GPT-4 model managed to score 700 out of 800 on the SAT maths exam
Statistics being pumped out like this need closer examination. Was the model trained on lots of previous exams and given all the correct answers? As these exams are similar every year, and test the same problems, it shouldn't be too surprising to get a high score if that is the case.
To test how good it is at "learning", it should be trained on the theory from the maths textbooks alone, then set a test which it hasn't seen before. I would bet that the score achieved is much lower, showing that it hasn't really learnt the subject, just how to answer lots of similar questions.
Re: GPT-4 model managed to score 700 out of 800 on the SAT maths exam
That is a truer test, like getting it to write, debug and test a program instead of copy/paste stack-exchange, etc.
But the reality is most humans 'train' on past paper examples, etc, and most academic institutes keep the same approach as making the exam harder more realistic in terms of problem-solving would cause an unacceptable drop in pass rates. And skulls mean money, not brains...
Re: GPT-4 model managed to score 700 out of 800 on the SAT maths exam
You're right about how lots of people also study for the test. But the point of learning maths is not to pass the test, it is to be able to apply it to solve real world problems. And solving real world problems, replacing trained humans, is what LLM based AI is being hyped up as being able to do. Therefore it should have to prove that it has the capability to perform highly when faced with unfamiliar problems, just as people can.
Re: GPT-4 model managed to score 700 out of 800 on the SAT maths exam
Reviewing and practicing on thousands of sample questions is exactly how those expensive SAT cram schools work.
(Decades ago) I bought a few SAT sample question books and did the same thing without going to an expensive cram screwall and scored an 800.
And the funny thing is, when I evaluate my past self at that stage, I can't help thinking what a naive waif I was, because real wisdom comes from years of experience.
"I would bet that the score achieved is much lower, showing that it hasn't really learnt the subject, just how to answer lots of similar questions."
I suspect you're right. But that also goes for human students - study and completion of past papers is pretty much universal in all subjects at all levels. I suspect that similar changes would be observed in the human candidates if they were simply given the text books - though even text books often have questions to test understanding.
As someone else mentioned above, I'm not sure mechanical Turk folks represent the best possible candidates, either. Maybe it would be better to pay a class of high school students - or several, at different schools - to complete the questions instead.
the example
You can see how it messed up with the example picture. The image shows a container labelled as a 600ml glass (look on the left). The graduated markings only go upto 400ml. Someone with actual *understanding* will probably say 400ml, but might also misinterpret what the question actually is and say 600ml. Without understanding the actual norms of measurement, 600ml is the right answer, I think. Similar for someone who isn't good at English - "highest amount this class measures" versus "amount this glass holds".
Note that the computer understood the mis-stated question ("class", not "glass").
We're just waiting for ChatGPT-alikes to get some logic, rather than just making good guesses. I can see a collusion between the likes of Wolfram Alpha, some physics engines and the chat bots coming.