News: 0185539530

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Feds Accuse China of 'Systematic' Distillation of US AI Models

(Tuesday September 08, 2026 @11:30PM (BeauHD) from the frontier-copying dept.)


The NSA, CISA, and FBI are [1]accusing (PDF) several Chinese AI companies of [2]carrying out "industrial-scale" distillation campaigns against leading U.S. models such as ChatGPT, Claude, Gemini, and Grok. Since at least 2024, the companies have allegedly routed millions of requests across accounts, APIs, proxies, cloud providers, and third-party aggregators to extract capabilities for their own models. "China-based artificial intelligence companies are conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies' models through industrial-scale knowledge distillation campaigns that form the core -- not merely a supplement -- of their AI development strategy," the agencies wrote. CyberScoop reports:

> DeepSeek, for example, distilled frontier U.S. models to generate synthetic training data for its R1 and R3 models, including four different versions of Claude, two versions of Gemini, five versions of ChatGPT and Grok 4. Those models helped train DeepSeek's capabilities in areas like agentic functioning, question and answer optimization, creative and occupational writing and others.

>

> Another Chinese company, Moonshot AI, allegedly distilled 18 different U.S. models -- including Fable 5, Anthropic's current, most advanced commercially available model -- to train its Kimi-K2 and Kimi K3 models. The company used millions of queries meant to extract enhanced capabilities in areas like agentic reasoning, coding and data analysis, computer vision, larger logical frameworks, visual processing and others.

>

> Chinese AI companies manage a sophisticated set of tools and systems that route requests and prompts through multiple pathways to avoid detection. The advisory lists common tactics observed by Chinese companies, including spreading requests across different accounts, models and platforms, using native APIs, remote cloud providers, and third-party aggregators to obfuscate user metadata, and leveraging proxies and gray tech markets to get around geographic restrictions, terms of use and safeguards built into frontier models.



[1] https://media.defense.gov/2026/Sep/08/2003992823/-1/-1/1/CSA_CHINA_BASED_AI_COMPANIES_MALICIOUS_DISTILLATION_AGAINST_US.PDF

[2] https://cyberscoop.com/us-accuses-chinese-ai-companies-distillation/



Lolol (Score:4, Insightful)

by dotslashdot ( 694478 )

You mean they are stealing the copyrighted proprietary works of others without compensation or warning? Say it is not so!

Re:Lolol (Score:5, Informative)

by ceoyoyo ( 59147 )

They're not stealing, and its not copyrightable. They're signing up, paying the requested amount, sending queries and accepting the responses.

Re:Lolol (Score:4, Informative)

by geekmux ( 1040042 )

> They're not stealing, and its not copyrightable. They're signing up, paying the requested amount, sending queries and accepting the responses.

This.

And if they’re worried about [some_country] getting into the precious Intelligence and infecting it, then turn on the fucking firewall. I’m assuming people still remember what those are for.

Oh. I’m sorry. Did that logic trip over someones stock price? Gee, can’t imagine what the problem is here..

Re: (Score:3)

by memory_register ( 6248354 )

Obligatory Princess Bride reference:

"You are trying to kidnap what I have rightfully stolen!"

Oh no! (Score:5, Insightful)

by paul_engr ( 6280294 )

Don't let them steal the shit that was made by stealing shit!

So what? (Score:5, Insightful)

by kertaamo ( 16100 )

So the accusation is that the Chinese obtained all the information they required from the US LMM creators.

But those US LLM creators obtained all the information they need from all of the users of the internet. Without any permission or compensation or even credit.

I don't think they have a leg to stand on in that accusation.

Re: (Score:2)

by HiThere ( 15173 )

Sorry, but just because A is guilty of something doesn't mean that B can't have done it to A.

OTOH, it's been ruled that the output of AIs isn't copyrightable, so...

Re: (Score:3)

by hey! ( 33014 )

It doesn't make it *right*, but it does make claiming it is *wrong* inconsistent with their own behavior. This could prevent the US companies from suing the Chinese companies seeking an injunction (due to the "unclean hands" doctrine), and probably blocks them from seeking monetary damages in most US jurisdictions.

And suing may undermine the US companies own intellectual property claims by exposing their shaky foundations. An AI model isn't *expression*, so it can't be copyrighted. Insofar as the service

So AI Finances dependant on Chinese accounts? (Score:5, Interesting)

by ealbers ( 553702 )

Does that mean most of the paid accounts of openai and anthropic are Chinese accounts?

Not a good look for a IPO

Re: (Score:2)

by sg_oneill ( 159032 )

It'd be a pretty amusing outome if the chinese scrapers all said "Fine, you know what.... we'll stop using you", and the income of Anthropic and OAI completely collapsed.

Re: (Score:2)

by postbigbang ( 761081 )

And.... look what pile of crap they scraped, distilled, and the output will be... crap.

Sure, this is about income, but little of this stuff works, even now.

Despite exhortations that New Advances Are Made Every Day and Look AI Employment Is Actually Up!, it's still highly-paid marketing people babbling BS about the power and integrity of their advancements.

Let the Chinese scrape! More is better! Why? It's like the old game of Telephone. The further away from cogent source material, the worse it will inevitab

Re: So AI Finances dependant on Chinese accounts? (Score:2)

by paul_engr ( 6280294 )

Oh man kevin oleary is going to explode when it's revealed that the hype was bankrolled by those pesky /distillers/ and there's no real demand for this trash

Accused? (Score:2)

by ThurstonMoore ( 605470 )

I assumed this was a given like water being wet, they would be stupid not to. Doesn't everyone do this?

Re: (Score:2)

by sg_oneill ( 159032 )

> I assumed this was a given like water being wet, they would be stupid not to. Doesn't everyone do this?

Yes. Early versions of both llama and claude would occasionally claim to be GPT, the smoking gun sign of distillation. And OAI has heavily implied that the "improper" use of ChatGPT by XAI was Distillation.

It has long considered a fairly normal and standard part of AI training by researchers.

Re: (Score:2)

by WaffleMonster ( 969671 )

> Yes. Early versions of both llama and claude would occasionally claim to be GPT, the smoking gun sign of distillation.

Distillation is about transferring behaviors not so much the absorption of information. There was probably just a lot of chatter about ChatGPT in its training set and /w typical system prompt blurting out ChatGPT organically isn't terribly surprising. This is no different than asking a model what it does for a living and where it works. Whatever it says chances are if you clear context and ask again it will say something different.

no honor among thieves (Score:3)

by sdinfoserv ( 1793266 )

It's so hard to feel bad for these thieving companies complaining about theft. Add on top increased utilities the data centers are costing us, the jobs they keep promising to take from us, and the risks of AI they continually warn about - despite marching stead fast towards the cliff. phuck'em.

Billion of $ of work sure is easy to steal (Score:2)

by locater16 ( 2326718 )

Unlike every other piece of software in history our AI model business is totally uncopyable, so long as you would pretty please stop copying it. Did I mention we're IPOing for infinity trillion dollars soon?

And the punishment? (Score:3)

by fahrbot-bot ( 874524 )

Ignoring the irony of claiming IP theft of AI data, that (basically) used IP stolen from everyone to build it up. What's the Administration going to do? Put tariffs on Chinese goods (that we pay) or impose a trade embargo on China (that we'll suffer)? 'Cause those have worked out so well before. /s Maybe start a real war with China? 'Cause the easy [1]"small potatoes" [apnews.com] "not a war" with Iran is going so well. /s I'm not really sure I see any good options for following up on these allegations, even if they're not completely bogus. Otherwise, it just sounds like it's setting up a pretext for casting blame if/when the U.S. falls behind 'cause we did something stupid.

[1] https://apnews.com/article/trump-iran-vance-war-23a45a2c45c048a9e894baeab89c7a2f

"Accuse China" (Score:1)

by Anonymous Coward

They accuse certain companies, not a whole country. Framing it like this is bad framing, maybe even bad faith.

Does anyone think any of the companies isn't doing (Score:2)

by allo ( 1728082 )

Does anyone think any of the companies isn't doing that? I bet as soon as a new DeepSeek, GLM, or other large model comes out OpenAI and Anthropic test what it is good at and how it can be distilled. That everyone is training on everyone's models is a open secret. Just try to let Claude write some stories and look for slop words that originally were only present in DeepSeek models.

Re: Does anyone think any of the companies isn't d (Score:2)

by paul_engr ( 6280294 )

Or don't look and LLMs as an industry will collapse and we can all move on

Re: (Score:2)

by allo ( 1728082 )

If one can make any prediction about the future of AI, than it's that LLM won't go away anymore.

Wow. Bummer. (Score:4, Insightful)

by ndykman ( 659315 )

Companies that ran roughshod over copyright and IP in order to create AGI (whatever the heck that means this month) are unhappy that China, who fundamentally doesn't see copyright and IP in the same way at all, are stealing things to undercut models they shouldn't have access too.

The same companies that wanted to eliminate as many creative or coding jobs as possible.

Oh no. How completely unexpected.

"Accuse" (Score:5, Interesting)

by reanjr ( 588767 )

The word "accuse" is being used as if it's some sort of crime or unethical to engage in this behavior. It is not. This is the inevitable result of a retarded business model pushed by well-connected tech morons like Sam Altman.

Boo hoo hoo... (Score:2)

by jenningsthecat ( 1525947 )

> ... conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies' models ...

How much "systematic extraction of proprietary" data did these same U.S. AI companies perform in order to train their models?

How telling it is of them - and how pathetic - to whine "You stole the stuff we stole!". Snowflakes and grifters, the whole damned lot.

Oh really? (Score:2)

by thedarb ( 181754 )

How can we help make it even easier for them?

Thieves outraged they got robbed. (Score:2)

by Wokan ( 14062 )

In other news, sky is still blue.

The young lady had an unusual list,
Linked in part to a structural weakness.
She set no preconditions.