News: 1715239510

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Big brains divided over training AI with more AI: Is model collapse inevitable?

(2024/05/09)


AI model collapse – the degradation of quality expected from machine learning models that recursively train on their own output – is not inevitable, at least according to 14 academics.

The risk that ongoing generative AI output, known as synthetic data, will dilute human-created organic data and impair the performance of models trained on this increasingly fabricated corpus was highlighted by a separate group last year, [1]in a paper titled: "The Curse of Recursion: Training on Generated Data Makes Models Forget."

Ilia Shumailov, lead author of that paper, spoke to The Register earlier this year about this phenomenon, which has been documented in [2]other [3]studies .

[4]

Now another set of boffins – Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, Daniel Roberts, Diyi Yang, David Donoho, and Sanmi Koyejo – contend that the problem of training AI on AI-made data isn't significant, given the way that model training is actually done.

[5]

[6]

This latest baker's dozen plus one – from Stanford, AI safety group Constellation, the University of Maryland at College Park, MIT, and Sequoia Capital – make the case for not worrying in [7]a paper titled: "Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data."

It's worth noting that some of these boffins acknowledge support through grants from commercial entities including OpenAI and Google, although the authors insist their research results do not necessarily reflect the positions or policies of their funders.

[8]

Gerstgrasser, a postdoctoral research associate at Harvard SEAS and visiting postdoctoral scholar at Stanford, [9]outlined on social media the argument he and his colleagues want to make.

"As AI-generated content becomes more prevalent on the internet, there's a growing concern that future AI models will be trained on this 'tainted' data," he asserted. "It's like a virus that could infect the entire AI ecosystem!

"Many experts have warned that this could lead to a doomsday scenario for AI. If models keep getting worse and worse with each generation, we could face an 'AI apocalypse'! But don't panic just yet …"

[10]AI is going to eat itself: Experiment shows people training bots are using bots

[11]Prompt engineering is a task best left to AI models

[12]What's up with AI lately? Let's start with soaring costs, public anger, regulations...

[13]AI agents can copy humans to get closer to artificial general intelligence, DeepMind finds

Gerstgrasser argued that while previous studies have warned about this "doomsday scenario," all that research relies on the assumption that each succeeding generation of AI would train exclusively on the synthetic data produced by the previous generation model.

He argues that legacy data won't just be discarded. Instead of being replaced every generation, it's more likely to accumulate – the synthetic data will just get mixed with the organic data, and the resulting model will continue to perform.

[14]

"Our findings extend these prior works to show that if data accumulates and models train on a mixture of 'real' and synthetic data, model collapse no longer occurs," Gerstgrasser et al declare in their "Is Model Collapse Inevitable?" paper.

"[T]hese results strongly suggest that the 'curse of recursion' may not be as dire as had been portrayed – provided we accumulate synthetic data alongside real data, rather than replacing real data by synthetic data only."

But the authors of a related [15]paper – Elvis Dohmatob, Yunzhen Feng, and Julia Kempe – titled, "Model Collapse Demystified: The Case of Regression," disagree that synthetic data can be added to model training without consequence.

All about scale

Julia Kempe, professor of computer science, mathematics and data science at the New York University Center for Data Science and Courant Institute of Mathematical Sciences, told The Register the "Is Model Collapse Inevitable?" paper is misguided in its conclusions – noting that it largely relies on the work that she and her colleagues did.

"Usually, when you train a model on lots of data, it gets better and better the more data you train on," Kempe explained. "This relation is called a 'scaling law' and has been shown to hold both empirically in many settings, and theoretically in several models.

"In our paper we show that when a model is trained on synthetic data that comes from a previous model that itself was generated on data from a previous model and so on, for a number of times (let us call the number of times n), then its performance does not obey the usual scaling laws; rather, it behaves effectively as if it had only been trained on an n-fraction of original data.

"For example, if we iteratively train and synthesize ten times, and then use the data from the last model to train, then we only get the performance we would get had we trained on 1/10 th of the original data, so much worse!"

Yunzhen Feng, a doctoral student in data science at New York University and one of Kempe's co-authors, also disagreed with the "Is Model Collapse Inevitable?" paper and its suggestion that model collapse can be discounted.

If the objective is to maintain a good performance, it might be preferable to consistently use the original dataset

"If the objective is to maintain a good performance, it might be preferable to consistently use the original dataset, which is already stored and selected prior to introducing synthetic data," Feng explained.

"Our aim is to keep the scaling benefits," Feng continued. "In the scaling regime, using clean data to increase the dataset size tenfold results in better scaling. Conversely, using synthetic data not only forfeits these benefits but also introduces a performance degradation. Therefore, we disagree with them."

Feng also pointed to [16]another paper – by Dohmatob, Feng, Pu Yang, Francois Charton, and Kempe – titled, "Tale of Tails: Model Collapse as a Change of Scaling Laws," and told The Register : "We argue that model collapse in AI data, from a scaling perspective, is twofold: It involves losing the performance benefits that additional human data would normally provide, and it results in recursive degradation across generations and retraining on AI data."

Feng noted that while there are various strategies that can be implemented to halt recursive degradation, there are performance consequences: "I believe most people do not regard solving only the second issue as sufficient to claim avoidance of model collapse."

Counterpoint

It's worth saying that Shumailov and his colleagues – Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and the late Ross Anderson – weren't really pitching the idea that AI is doomed to devour itself in their "Curse of Recursion" paper. Their conclusion was more subtle: That the model collapse can be mitigated by spending money to assure data quality – something big companies will find easier than small ones.

Asked about the findings from Gerstgrasser et al , Shumailov replied, "In principle it does not really invalidate anything we showed. With simple models, they show they can attenuate some effects. Do note that this comes with ever increasing cost and doesn't solve any of the problems for common users, who will have no ability to keep data long term."

AI collapse isn't inevitable – but neither is model performance. ®

Get our [17]Tech Resources



[1] https://www.theregister.com/2024/01/26/what_is_model_collapse/

[2] https://arxiv.org/abs/2311.16822

[3] https://arxiv.org/abs/2311.09807

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZjyewPRDlZcGfHvZCovvywAAAAY&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZjyewPRDlZcGfHvZCovvywAAAAY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZjyewPRDlZcGfHvZCovvywAAAAY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[7] https://arxiv.org/abs/2404.01413

[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZjyewPRDlZcGfHvZCovvywAAAAY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[9] https://x.com/MGerstgrasser/status/1785732983238111425

[10] https://www.theregister.com/2023/06/16/crowd_workers_bots_ai_training/

[11] https://www.theregister.com/2024/02/22/prompt_engineering_ai_models/

[12] https://www.theregister.com/2024/04/15/stanford_report_ai/

[13] https://www.theregister.com/2023/11/28/ai_agents_can_copy_humans/

[14] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZjyewPRDlZcGfHvZCovvywAAAAY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[15] https://arxiv.org/abs/2402.07712

[16] https://arxiv.org/abs/2402.07043

[17] https://whitepapers.theregister.com/



Strange you should say that....

The Dogs Meevonks

Just the other week I posted the comment below.

"As everyone (inc me keeps repeating) who will bother reading what no one could be bothered to write?

Oh... other AI's reading other AI's in a never ending loop of regurgitated ever more inaccurate bollocks until it collapses under it's own ineptitude.... at least... that's my hope. let it come crashing down as swiftly as NFTs did."

Filippo

From what I understand, they can't claim that there is no problem. They may claim that instead of "model collapse" we'll get just some degree of "model degradation". That's still a pretty big problem, considering that current models are already largely only good for party tricks, and fairly crap at serious tasks.

I still think that the current main dangers to AI research are not "model collapse", but (1) overhype, and (2) focusing resources on superficially promising avenues that turn out to be dead ends.

Tomi Tank

Why worry if it is over-hyped? So was the internet and that sorted it self out.

It seems that most criticisms against AI come from those who missed the boat. Yes, it is in a bubble at the moment, but it seems the greybeards dont seem to understand that the only thing that really matters nowadays is to make you stash and run - as I have done.

Who cares? Why care? No point. Take the money and run. Im off to the Cayman Islands soon and I wont be looking back.

Ive spent a good part of my career - like many here - working to improve people's lives through the Internet. And I've learnt it is mostly a waste of time because stupid always finds a way.

"considering that current models are already largely only good for party tricks, and fairly crap at serious tasks."

AI is much more than ChatGPT3. In my own field of predictive human behaviour it is a gold mine and it works better than humans. It is being tested by a mental health service in south east London on their patients and has proved wildly successful in trails. It is being used extensively in sport - especially tennis and football. It has value and offers what was never possible b4. It's being pitched for prisons, police stations, general hospitals - anything that needs observation. It is only a matter of time before this tech filters down to your CCTV. It is being used extensively throughout Japan by Lawsons and 7-11 to predict shoplifters. The Chinese (who are world leaders in this particular field) use it throughout their systems. To give u some perspective of how far the West is behind, DeepMind's recent football set-piece prediction is 2 years behind the work we did in China.

Even in the current state, predictive AI surrounds your existence. AI will completely dominate your future in a way that makes the internet looks puny. And that is a nightmare! A true horror story awaits us - I know cause a built a tiny part of it. The only chance I could see of living a normal life of sorts was to get a load of money. I recommend you do the same.

refitman

I'm going out on a limb and going to say that you went big in on crypto being "the next big thing that will revolutionise our lives"?

"it behaves effectively as if it had only been trained on an n-fraction of original data"

Mike 137

I can't comment on the numerical specifics without sight of the paper, but the principle is obvious. The potential variety of output depends on the diversity of the training data set. Recursive training on output data inevitably tends to homogenise output as no new information points are generated, only reorganisations of the original corpus. It would be nice to have a link to the Kempe paper.

Satire is what closes Saturday night.
-- George Kaufman