News: 1647457214

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Even complex AI models are failing 5th grade science

(2022/03/16)


Think your AI agents are actually learning to solve problems? A new benchmark sheds light on what is real when it comes to sophisticated AI.

Researchers from the University of Arizona, Microsoft, and the Allen Institute for AI tested several different state-of-the-art agents and found them readily able to answer the "what" of a situation, but incapable of determining the "how" of them.

The agents were put to the test using a benchmark built especially for the task that the researchers called [1]ScienceWorld . ScienceWorld will be immediately familiar to anyone who's played an old-school, text-based MUD: it has multiple rooms, objects that can be interacted with, and tasks to be performed. In this case, it's not so much killing goblins, but completing the equivalent of elementary-level science projects.

[2]

ScienceWorld has simulation engines for thermodynamics, electrical circuits, matter and chemistry reactions, and biological processes. The researchers got their list of experiments for the agents by turning typical science test questions into experiments, and then testing the agents to see if they could reason to the answer.

[3]

[4]

The fact that agents can quickly answer "whats," but not "hows," raises "the question of whether current models are simply retrieving answers by way of seeing a large number of similar input examples or if they have learned to reason about concepts in a reusable manner," the researchers said.

Not to spoil it, but there's no reasoning going on inside those digital brains.

[5]

ScienceWorld's builders were looking for digital agents to do one thing in particular – "combine declarative scientific and world knowledge with the procedural knowledge required to correctly complete the experiment."

In one experiment, agents were tested to see if they could identify a fork, find the necessary materials needed to test it for conductivity, and then put it in the correct box.

[6]Machine-learning models more powerful, toxic than ever

[7]Research finds data poisoning can't defeat facial recognition

[8]Googlers and co offer video dataset-generating Kubric

[9]Cerebras brings wafer-size AI chips to medical data analysis

Another experiment had agents trying to determine if an ice cube would melt on a stove. Again, this requires the agent to identify, pick up, and manipulate several objects inside ScienceWorld. As an added level of challenge in all situations, various properties of objects inside ScienceWorld (location, color, etc.) change each time the simulation is started to prevent agents from simply memorizing a sequence.

Scoring of the 30 different tasks in ScienceWorld is based on a scale of 0.00, a total failure, to 1, indicating perfect performance. The highest score for any AI under test was 0.54, and that was on one of the simplest: identifying a non-living thing. For the ice, the best was 0.04. In fact, a random-action generator stood out, with 0.63 for identifying a non-living thing. Building circuits was also abysmal.

This led the academics to conclude:

Agents for text-based games as well as novel models adapted from transformer-based scientific question-answering solvers perform poorly on tasks (such as melting ice) that 5th grade science students can perform with ease

"Overall, these tasks are challenging for current models, with the best model (DRRN) achieving an average score of 0.18 across all 30 subtasks," the paper said. So, which models performed best? Even that's a tricky question to answer.

The researchers found models using valid action detection aid tended to perform better than those that must first learn to generate valid actions, and models that used large language model components for action selection tended to perform more poorly.

[10]

Interactive reinforcement learning models were able to quickly identify and classify objects, but had difficulty picking objects up and putting them in the right box. Open-ended tasks, like those requiring the bot to change the state of an object, were difficult for all the models.

The biggest takeaway from the project comes from another finding – that agents with larger models don't necessarily perform better. The DRRN model only had 1.5 million parameters, which was four orders of magnitude fewer than the pair of T5 models used in the experiment, yet DRRN performed better.

"Our results also suggest that agents that learn interactively in a grounded environment are more sample and parameter efficient than large language models that learn offline by reading text from static sources," the report concludes. ®

Get our [11]Tech Resources



[1] https://arxiv.org/abs/2203.07540

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YjJsJMVStxc40tk0LBl7-wAAAMw&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YjJsJMVStxc40tk0LBl7-wAAAMw&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YjJsJMVStxc40tk0LBl7-wAAAMw&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YjJsJMVStxc40tk0LBl7-wAAAMw&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://www.theregister.com/2022/03/16/ai_index_report_2022/

[7] https://www.theregister.com/2022/03/15/research_finds_data_poisoning_cant/

[8] https://www.theregister.com/2022/03/15/deep_learning_vision_kubric/

[9] https://www.theregister.com/2022/03/14/cerebras_ai_chips/

[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YjJsJMVStxc40tk0LBl7-wAAAMw&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[11] https://whitepapers.theregister.com/



"there's no reasoning going on inside those digital brains"

Mike 137

We've known this for ages - it's logically inevitable. The current mechanism of AI is adaptive statistical template matching - i.e. lookup to find patterns maximally similar to input stimuli, potentially adjusting for majority trends. It's about as intelligent as a bumblebee (which also uses adaptive template matching to learn). That's not to say current AI doesn't have its uses - particularly where the range of alternatives to select from is huge - but it's quite unrealistic to call what it does 'intelligence' in human terms. Useful tools maybe, but probably misnamed, and certainly not safely applicable to situations where it's faced with the entirely unexpected.

However I'm a bit depressed that the 'projects' tested are classified as 5th grade - they seem rather elementary for that.

Re: "there's no reasoning going on inside those digital brains"

b0llchit

It's about as intelligent as a bumblebee.

I'm sure that the bumblebee is several orders of magnitude more intelligent than any AI currently available. I have not yet heard of or seen an AI that is actually self-sufficient.

But you are right, current AI systems are primarily statistical inference machines. They have their uses, but calling them "intelligent" would be a real slap in the face of that bumblebee. I'm sure the bumblebees perform better at solving the 5th grade science problems than the tested models.

Re: "there's no reasoning going on inside those digital brains"

heyrick

Exactly this. There is zero understanding , so there's no intelligence . Artificial Idiocy (as I like to call it) has a ridiculously long way to go before it is capable of being able to "reason".

@Mike 137 - Re: "there's no reasoning going on inside those digital brains"

Anonymous Coward

Even more depressing (I would say frightening) is the fact that AI is pushed or adopted in real-life scenarios where life of human beings is at stake, with no oversight and with no possibility of remediation. You must be at fault because AI said so and it can't possibly be wrong.

Re: @Mike 137 - "there's no reasoning going on inside those digital brains"

Will Godfrey

I've read quite a few SciFi stories based on that premise. Scary ones!

Re: @Mike 137 - "there's no reasoning going on inside those digital brains"

heyrick

Recommendations welcome...

Re: @Mike 137 - "there's no reasoning going on inside those digital brains"

very angry man

Ai, it's just a few short instructions:

input from IR

input from movement sensor

is IR input moving?

Aline weapon

input from Lidar

is weapon line obstructed?

change position to clear weapon line

discharge weapon

play recording " kill all humans"

close loop

restart loop

change position.

The next question...

Andy 73

..would you let a 5th grader drive your car?

Asimo

HildyJ

As anyone who remembers Asimo's debut demo will tell you, it aced the first step up a set of stairs, managed the second step, and fell over on the third step.

This is AI's downfall. Ask it a single question and its Artificial Guess can be accurate. But the second Artificial Guess takes the first Artificial Guess as a given, which introduces more uncertainty, and each subsequent Artificial Guess compounds it.

IRL, an AI can recognize potential areas of concern in a mammogram better than the average doctor. But if you tried to teach it to plan out the treatment, I eoud not be surprised to see it recommending mastectomies just in case.

Re: Asimo

heyrick

" see it recommending mastectomies just in case "

Google "Ian Paterson"...

One measure of friendship consists not in the number of things friends
can discuss, but in the number of things they need no longer mention.
-- Clifton Fadiman