Even complex AI models are failing 5th grade science
- Reference: 1647457214
- News link: https://www.theregister.co.uk/2022/03/16/scienceworld_ai_benchmark/
- Source link:
Researchers from the University of Arizona, Microsoft, and the Allen Institute for AI tested several different state-of-the-art agents and found them readily able to answer the "what" of a situation, but incapable of determining the "how" of them.
The agents were put to the test using a benchmark built especially for the task that the researchers called [1]ScienceWorld . ScienceWorld will be immediately familiar to anyone who's played an old-school, text-based MUD: it has multiple rooms, objects that can be interacted with, and tasks to be performed. In this case, it's not so much killing goblins, but completing the equivalent of elementary-level science projects.
[2]
ScienceWorld has simulation engines for thermodynamics, electrical circuits, matter and chemistry reactions, and biological processes. The researchers got their list of experiments for the agents by turning typical science test questions into experiments, and then testing the agents to see if they could reason to the answer.
[3]
[4]
The fact that agents can quickly answer "whats," but not "hows," raises "the question of whether current models are simply retrieving answers by way of seeing a large number of similar input examples or if they have learned to reason about concepts in a reusable manner," the researchers said.
Not to spoil it, but there's no reasoning going on inside those digital brains.
[5]
ScienceWorld's builders were looking for digital agents to do one thing in particular – "combine declarative scientific and world knowledge with the procedural knowledge required to correctly complete the experiment."
In one experiment, agents were tested to see if they could identify a fork, find the necessary materials needed to test it for conductivity, and then put it in the correct box.
[6]Machine-learning models more powerful, toxic than ever
[7]Research finds data poisoning can't defeat facial recognition
[8]Googlers and co offer video dataset-generating Kubric
[9]Cerebras brings wafer-size AI chips to medical data analysis
Another experiment had agents trying to determine if an ice cube would melt on a stove. Again, this requires the agent to identify, pick up, and manipulate several objects inside ScienceWorld. As an added level of challenge in all situations, various properties of objects inside ScienceWorld (location, color, etc.) change each time the simulation is started to prevent agents from simply memorizing a sequence.
Scoring of the 30 different tasks in ScienceWorld is based on a scale of 0.00, a total failure, to 1, indicating perfect performance. The highest score for any AI under test was 0.54, and that was on one of the simplest: identifying a non-living thing. For the ice, the best was 0.04. In fact, a random-action generator stood out, with 0.63 for identifying a non-living thing. Building circuits was also abysmal.
This led the academics to conclude:
Agents for text-based games as well as novel models adapted from transformer-based scientific question-answering solvers perform poorly on tasks (such as melting ice) that 5th grade science students can perform with ease
"Overall, these tasks are challenging for current models, with the best model (DRRN) achieving an average score of 0.18 across all 30 subtasks," the paper said. So, which models performed best? Even that's a tricky question to answer.
The researchers found models using valid action detection aid tended to perform better than those that must first learn to generate valid actions, and models that used large language model components for action selection tended to perform more poorly.
[10]
Interactive reinforcement learning models were able to quickly identify and classify objects, but had difficulty picking objects up and putting them in the right box. Open-ended tasks, like those requiring the bot to change the state of an object, were difficult for all the models.
The biggest takeaway from the project comes from another finding – that agents with larger models don't necessarily perform better. The DRRN model only had 1.5 million parameters, which was four orders of magnitude fewer than the pair of T5 models used in the experiment, yet DRRN performed better.
"Our results also suggest that agents that learn interactively in a grounded environment are more sample and parameter efficient than large language models that learn offline by reading text from static sources," the report concludes. ®
Get our [11]Tech Resources
[1] https://arxiv.org/abs/2203.07540
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YjJsJMVStxc40tk0LBl7-wAAAMw&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YjJsJMVStxc40tk0LBl7-wAAAMw&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YjJsJMVStxc40tk0LBl7-wAAAMw&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YjJsJMVStxc40tk0LBl7-wAAAMw&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2022/03/16/ai_index_report_2022/
[7] https://www.theregister.com/2022/03/15/research_finds_data_poisoning_cant/
[8] https://www.theregister.com/2022/03/15/deep_learning_vision_kubric/
[9] https://www.theregister.com/2022/03/14/cerebras_ai_chips/
[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YjJsJMVStxc40tk0LBl7-wAAAMw&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[11] https://whitepapers.theregister.com/
Re: "there's no reasoning going on inside those digital brains"
It's about as intelligent as a bumblebee.
I'm sure that the bumblebee is several orders of magnitude more intelligent than any AI currently available. I have not yet heard of or seen an AI that is actually self-sufficient.
But you are right, current AI systems are primarily statistical inference machines. They have their uses, but calling them "intelligent" would be a real slap in the face of that bumblebee. I'm sure the bumblebees perform better at solving the 5th grade science problems than the tested models.
Re: "there's no reasoning going on inside those digital brains"
Exactly this. There is zero understanding , so there's no intelligence . Artificial Idiocy (as I like to call it) has a ridiculously long way to go before it is capable of being able to "reason".
@Mike 137 - Re: "there's no reasoning going on inside those digital brains"
Even more depressing (I would say frightening) is the fact that AI is pushed or adopted in real-life scenarios where life of human beings is at stake, with no oversight and with no possibility of remediation. You must be at fault because AI said so and it can't possibly be wrong.
Re: @Mike 137 - "there's no reasoning going on inside those digital brains"
I've read quite a few SciFi stories based on that premise. Scary ones!
Re: @Mike 137 - "there's no reasoning going on inside those digital brains"
Recommendations welcome...
Re: @Mike 137 - "there's no reasoning going on inside those digital brains"
Ai, it's just a few short instructions:
input from IR
input from movement sensor
is IR input moving?
Aline weapon
input from Lidar
is weapon line obstructed?
change position to clear weapon line
discharge weapon
play recording " kill all humans"
close loop
restart loop
change position.
The next question...
..would you let a 5th grader drive your car?
Asimo
As anyone who remembers Asimo's debut demo will tell you, it aced the first step up a set of stairs, managed the second step, and fell over on the third step.
This is AI's downfall. Ask it a single question and its Artificial Guess can be accurate. But the second Artificial Guess takes the first Artificial Guess as a given, which introduces more uncertainty, and each subsequent Artificial Guess compounds it.
IRL, an AI can recognize potential areas of concern in a mammogram better than the average doctor. But if you tried to teach it to plan out the treatment, I eoud not be surprised to see it recommending mastectomies just in case.
Re: Asimo
" see it recommending mastectomies just in case "
Google "Ian Paterson"...
"there's no reasoning going on inside those digital brains"
We've known this for ages - it's logically inevitable. The current mechanism of AI is adaptive statistical template matching - i.e. lookup to find patterns maximally similar to input stimuli, potentially adjusting for majority trends. It's about as intelligent as a bumblebee (which also uses adaptive template matching to learn). That's not to say current AI doesn't have its uses - particularly where the range of alternatives to select from is huge - but it's quite unrealistic to call what it does 'intelligence' in human terms. Useful tools maybe, but probably misnamed, and certainly not safely applicable to situations where it's faced with the entirely unexpected.
However I'm a bit depressed that the 'projects' tested are classified as 5th grade - they seem rather elementary for that.