Should US Open-Weight AI Labs 'Distill' Frontier Models Too? (techcrunch.com)
- Reference: 0185640232
- News link: https://news.slashdot.org/story/26/09/13/2334206/should-us-open-weight-ai-labs-distill-frontier-models-too
- Source link: https://techcrunch.com/2026/09/11/y-combinators-garry-tan-wants-u-s-open-weight-ai-labs-to-distill-frontier-models-too/
> Distillation is the process of using the outputs of a more capable AI model to train a smaller or less capable one, sometimes illicitly... Anthropic [2]has accused Chinese companies such as Moonshot AI, DeepSeek, and MiniMax of the practice, while [3]OpenAI believes DeepSeek's V3 and R1 model architectures were distilled from its own GPT-4 and GPT-4o models. In the midst of this, the U.S.'s National Security Agency, Cybersecurity and Infrastructure Security Agency, and Federal Bureau of Investigation [4]released an [5]official cyber security advisory warning on the topic on Tuesday...
>
> But Tan believes regulators should focus less on curbing distillation and more on creating an equilibrium between open weight models and frontier models — as long as frontier models retain a price premium that allows their business model to remain feasible. "This is actually the ideal case. You want open weight models to give people freedom and access," he explained. "If I were a regulator, that's what I would go after." Tan acknowledged that this is a hard balance to strike, calling it "a tightrope." Nevertheless, he says it's a balance worth pursuing — saying it "could result in the best possible outcome."
Tan later told TechCrunch he'd like to see America with more open-weight options that aren't Chinese, built by smaller U.S. open-weight AI labs using those same training techniques on products from America's frontier AI labs:
> Anthropic CEO Dario Amodei had previously [6]publicly called on U.S. regulators to crack down on distillation. It's notable that the commander of Silicon Valley's prestigious and prolific startup accelerator doesn't agree.
>
> To be clear, Tan isn't advocating for American AI labs to use stolen credentials to distill. He wants them to be free to come in the front door. In fact, his argument is twofold. He feels it's an overreach for AI labs to dictate what their customers can do with the information their models share with them. He also notes that the proprietary AI labs didn't ask permission when they vacuumed up as much human knowledge as they could to train their models. They famously ingested plenty of copyrighted material [7]without the permission of those intellectual property holders . "Controlling what users and customers do with API calls to closed weight models feels constraining, and there's a role government can play here to normalize the fact that access to intelligence that was trained on broad public access data should itself also be more a form of a public good than something locked away behind restrictive terms of service," he told TechCrunch when asked why American labs should be free to distill, too...
>
> To him, the true AI doomer scenario is for all the immense power of frontier AI to wind up in the hands of a single powerful, proprietary provider. "The nightmare scenario, the doomer scenario for AI is that there's just one company," he said. "It has the best access to capital. It has the best AI researchers. It runs away with it and suddenly there's one company that's monolithic. And that would be bad."
[1] https://www.cnbc.com/2026/09/11/y-combinator-garry-tan-says-do-nothing-about-distillation.html
[2] https://www.cnbc.com/2026/09/03/anthropic-distillation-battle-turns-to-dark-web-china-concerns-swell.html
[3] https://www.cnbc.com/video/2025/02/21/how-deepseek-supercharged-ais-distillation-problem.html
[4] https://www.nsa.gov/Press-Room/Press-Releases-Statements/Press-Release-View/Article/4592113/nsa-and-others-warn-china-based-ai-companies-are-distilling-us-frontier-ai-mode/
[5] https://slashdot.org/story/26/09/08/220242/feds-accuse-china-of-systematic-distillation-of-us-ai-models
[6] https://www.anthropic.com/news/position-open-weights-models
[7] https://techcrunch.com/2026/07/20/anthropics-landmark-1-5b-copyright-settlement-is-approved/
Until it's Free as in Linux (Score:5, Insightful)
"Frontier" models trained on everyone else's data, according to the companies that made them it's perfectly legal. Thus it's perfectly legal to train off those same models, obeying terms of service and robots.txt is as perfectly optional for US open models as it is for these supposedly trillion dollar companies. Train away until it's actually as open source and free as in Linux and I never have to hear about Sam fucking Altman ever again.
Re: (Score:3)
Distillation on the frontier... sure sounds like making moonshine out in the woods... illegal, but it'll be done anyway. Might as well accept that some inbreds will be distillin'
Stealing copyright is the business model (Score:2)
The business model of AI companies is to steal copyrighted material as well as private data. The argument against distilling makes no sense to me. You have taught your AI drones this is the new Napster era payed for by Wall Street.
When I read AI tech discussions ... (Score:3)
... I know how my parents felt when I tried to explain programming too them. It all makes grammatical sense but the actual meaning eludes me.
Is it just me? Am I just getting old and heading for the get-off-my-lawn stage or are there any other people who like to think they're tech savvy but are totally confused by AI models and how they really work?
They distilled human knowledge (Score:1)
I suggest that banning distillation can only be considered if Dario, Sam, Sundar and Elon make sure they negotiate deals and pay for all copyrighted material that they used for training models in the past.
Lets just take the minimum statutory damages for copyright infringement, $700 per work.
Still interested Dario? Dario? Where did you go?
Re: They distilled human knowledge (Score:4)
I think we should just have the same laws for everyone. That's never going to happen though.
Re: (Score:3)
I believe restricting data is the answer to make a much better model. "AI", is just a search and compile engine. It doesnt think like people assume, it doesnt rationalize, it doesnt even understand what it is looking at.
For example, researching RNA, if you "teach" it the basics, and see if it "learns" on its own to get itself to where we are today, it wont know what to do. It cant come up with a theory and then logic through how to validate.
It CAN however, link likenesses together after learning how t
Re: (Score:2)
More to the point, Anthopic is trying to kneecap the Chinese models so that American frontier model companies can bask in warm glow of their exorbitant pricing.
Re: (Score:2)
Now that makes more sense.