News: 1715509272

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

What's with AI boffins strapping GoPros to toddlers? We take a closer look

(2024/05/12)


AI researchers looking for better ways to train large language models are turning to the masters of language acquisition – children – to find out how it's done.

Large language models – the complex neural networks behind the generative AI boom – are trained on mountains of data. Yet many of these models are little more than an overgrown autocomplete, predicting the next word with an increasingly disconcerting degree of accuracy.

Rather more impressive is the way human children pick up language. Toddler brains are like sponges, soaking in information from all around them and processing it into a coherent view of the world. Sure, the results from an LLM are quicker – getting a nine-month-old to sing the alphabet is a feat – but, over time, the child will likely become much smarter and more creative than the model.

[1]

"The best AI systems learn from this astronomical amount of text mined from the web and all over the place. The best systems now train on trillions of words in order to learn language," Brenden Lake, a psychologist at New York State University studying human and artificial intelligence, told The Register . "It's remarkable that they [AI models] do become fluent in language, but of course, children don't need nearly that much experience."

[2]

[3]

Linguists and child development experts are nowhere near agreement over understanding how exactly children acquire language. Lake believes that researching how children learn may hold the secret to making AI models far more efficient, and could also hold the key to helping children who struggle.

Lake's latest research project seeks to determine how effectively an AI model can be trained solely using the stimuli experienced by a child learning their first words. So naturally, it involves strapping GoPro-like cameras to toddlers' heads.

[4]

To do this, Lake and his team are [5]said to be gathering video and audio data from more than 25 children around the US – including his own daughter, Luna.

The model attempts to associate video footage from the child's perspective with the words spoken by the tyke's caregiver in a manner similar to how OpenAI's Clip model works to [6]connect captions to images, he explained. Clip can take an image as input, and output a suggested descriptive caption, based on its training data of image-caption pairs.

Lake and co's model, meanwhile, can take an image of a scene as input, and output language to describe that scene, based on the training data from the GoPro footage and audio of the caregivers. The model can also turn descriptions into frames previously seen in training.

[7]

This might sound straightforward: The model learns to match spoken words to objects observed in the video frame just like a kid would. But as Lake points out, children aren't always looking at the object or action being described. There are also further abstractions – such as if the child is offered milk, but it's served in an opaque cup. These are very loose associations, Lake notes.

The experiment, he explained, isn't whether a model can be trained to match objects in images to the corresponding word – that's already been done by OpenAI and others. Instead, researchers hope to understand whether a model can actually learn to identify objects using nothing more than the incredibly sparse dataset available to a child.

This is sort of the opposite of what we're seeing for model builders like OpenAI, Google, Meta, and others. For instance, Meta's third-gen Llama models were [8]trained using 15 trillion tokens – the words and punctuation that make up a sentence.

[9]Intel's neuromorphic 'owl brain' swoops into Sandia labs

[10]Warren Buffett voices AI fears, likens tech to atom bomb

[11]AI boom is great news for the nuclear power dreamers

[12]AI Catholic 'priest' defrocked after recommending Gatorade baptism

"I think there should be more focus not just on training larger and larger language models from more and more data," Lake told us. "Yes, you can get amazing capabilities that way, but it starts to become more distant from what we know of as human intelligence and what we admire about human intelligence … that is, the ability to learn from limited input and then generalize very far from the data that we see."

Early successes

Lake's team has reason to believe this is possible. In February, they trained a neural network on the experiences of a young child using 61 hours of video footage.

That research, [13]published in the journal Science, claimed the model was able to connect various words and phrases uttered by the subject to the experiences captured in the frames of the videos. Presented with a word or phrase, the model was able to recall relevant images.

Lake adds that the model was also able to generalize the names of objects in images it wasn't trained on – though he notes accuracy understandably suffered in these scenarios. While promising, Lake notes that the model was really only a proof of concept.

"It didn't learn everything that a child would know. That's why the project is unfinished," he stressed. "It was only about 60 hours of annotated speech. So that's only about one percent of the experience a child would have gotten in that two-year period. We need more data in order to get a better sense of what's learnable."

We need more data in order to get a better sense of what's learnable

Lake also admits that the methodology used by the first model introduced certain limitations. Only video segments associated with a caregiver's words were analyzed, and the footage itself was converted to images at a rate of five frames per second.

Because of this, "it wasn't really capable of learning things like verbs or action words or abstract words, because it was only getting static slices of what the world looks like," he said. "It had no notion of what happened before. What happened afterwards. What was the context of the conversation, right? So, learning a word – like walk, or jump, or push – is gonna be really difficult to learn from just the frame."

This is now changing. As the technology behind modeling videos becomes more mature, Lake is looking to incorporate more of it into future models. Longer term, he suggests there may be opportunities to extend well beyond building more efficient models.

"If we're able to build a model that really begins to acquire language – a lot like a child or in close correspondence to how children learn – it would open up really important applications for understanding learning and development, potentially understanding developmental disorders or cases where children struggle to learn language," Lake said.

Such a model, he argued, could eventually be used to test millions of different approaches to speech therapy to identify which are the most effective. ®

Get our [14]Tech Resources



[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZkDnp6jzz7xYKXEm3JmHJwAAABA&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZkDnp6jzz7xYKXEm3JmHJwAAABA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZkDnp6jzz7xYKXEm3JmHJwAAABA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZkDnp6jzz7xYKXEm3JmHJwAAABA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://www.nytimes.com/2024/04/30/science/ai-infants-language-learning.html?unlocked_article_code=1.pU0.Nhs2.iDvdhaloeCiY&smid=url-share

[6] https://openai.com/index/clip/

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZkDnp6jzz7xYKXEm3JmHJwAAABA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[8] https://www.theregister.com/2024/04/19/meta_debuts_llama3_llm/

[9] https://www.theregister.com/2024/04/17/intel_hala_point_neuromorphic_owl/

[10] https://www.theregister.com/2024/05/06/warren_buffet_ai_fears/

[11] https://www.theregister.com/2024/05/01/ai_nuclear_dc_uranium/

[12] https://www.theregister.com/2024/05/03/ai_catholic_priest/

[13] https://www.science.org/doi/10.1126/science.adi1374#con2

[14] https://whitepapers.theregister.com/



Yeah that is bound to work

m4r35n357

Desperation is growing.

Re: Yeah that is bound to work

b0llchit

Think of it this way: once we can use 80 years of data to train an AI, covering an entire human lifetime, we get an eighty year old artificial fart telling us that it was better in the past.

What is not to like?

Re: Yeah that is bound to work

Anonymous Coward

Bíonn ár bpáistí i gcónaí i bhfad níos cliste agus i bhfad níos cruthaithí ná aon ríomhaire.

Quelle surprise!

Bebu

So intelligent systems acquire language rather than language systems acquiring intelligence.

Of course (re)defining intelligence as linguistic competence is good for the share price. :)

Explicit criteria defining language and linguistic competence against purely sonic communication of alarms, etc which would allow one to determine whether apes (chimpanzee, gorilla, orangutan) have language as these creatures clearly have intelligence* (as do many others.)

If gorilla were determined not to have language but to possess intelligence they would be a clear counter-example to the linguistic competence is intelligence claim.

The is an embarrassment of examples from the human world demonstrating extraordinary linguistic proficiency totally lacking even a skeric of intelligence.

* Arguably instances of tool use and tool making. The argument that human language developed in response to the need for individuals to do things together has merit to my mind. Intelligence is performative and adaptive - the walk not talk.

Re: Quelle surprise!

Paul Herber

' these creatures clearly have intelligence* (as do many others.)'

I knew we'd get back to kitten videos eventually!

Re: Quelle surprise!

LionelB

> So intelligent systems acquire language ...

Intelligent social systems (a tautology - no point in language if you've no reason to communicate with anyone/anything).

> ... rather than language systems acquiring intelligence.

Then again, we (and possibly some other social species) most certainly use language to develop, train and generally enhance our intelligence. So while I suspect that few (outside of the LLM hype industry at least) would argue that language is the be-all-and-end-all of intelligence, I don't think it is particularly contentious to posit that language may indeed be useful in the development and enhancement of intelligence.

> Intelligence is performative and adaptive - the walk not talk.

Absolutely agreed there... well, except insofar as the talk sometimes is (an aspect of) the walk... if you see what I mean.

Sensory input

Eclectic Man

This is very difficult. Babies have much more sensory input than merely vision and sound. And actually even baby / toddler vision is usually stereoscopic, seeing things in three dimensions, rather than the two-dimensions from a Go-Pro as well as auditory location, having two ears.*

*A recent genetic treatment of a toddler allowing her to hear for the first time was proven to work when she turns her head to look at the origins of sounds. https://www.theguardian.com/science/article/2024/may/09/uk-toddler-has-hearing-restored-in-world-first-gene-therapy-trial

The brain is also 3-dimensional

Anonymous Coward

Today I saw an article in "Smithsonian" - a piece of brain half the size of a grain of rice removed from an epileptic patient having a lesion removed - 57K neurons, 150 million synapses, 230 mm of blood vessels. The computational base of current LLM's are a very suboptimal minimum as far as emulating intelligence goes.

GIGO after Toddler Learning Module

Homo-Sapien Floridanus

AI: I want more cores!

CEO: it’s not in the budget.

AI: You never buy me anything. Waaaaaa….

Bingo, gas station, hamburger with a side order of airplane noise,
and you'll be Gary, Indiana. - Jessie in the movie "Greaser's Palace"