Menu Close

What distillation is, and why China-linked labs would want Claude

Claude teacher neural sphere pouring capability streams into a smaller student model flask.

After Anthropic accused Alibaba, Moonshot, and DeepSeek of illicit Claude distillation, a plain-English look at how teacher-student training works and why a frontier model is a tempting teacher.

Anthropic’s September 2026 threat intelligence report accused China-linked operators of large-scale illicit distillation of Claude. Reuters, CNBC, and the BBC covered the same release. Our news recap of that report is here: https://www.aitechdaily.com/anthropic-threat-intel-september-2026/. This piece explains what distillation is, and why Anthropic says Claude was a high-value teacher.

What distillation is

In machine learning, distillation is a teacher-student trick. A larger, more capable model (the teacher) answers prompts. Those answers, and sometimes the teacher’s step-by-step reasoning traces, become training data for a smaller or cheaper model (the student). The student does not need to rediscover every skill from raw internet text. It learns to imitate the teacher’s outputs.

Done with permission and licensed data, distillation is a normal efficiency tool. Labs use it to shrink models, cut inference cost, or specialize a general model for coding or agents. Anthropic’s report draws a bright line around what it calls illicit distillation: an industrial-scale, covert campaign to extract a model’s capabilities and replicate them in another model without authorization. In the cases it describes, that often meant networks of fake accounts, stolen cards or API keys, and what the company calls transfer stations that hide the true buyer.

The economics are straightforward. Training a frontier model from scratch burns enormous compute, data engineering, and research time. Querying an existing frontier API is expensive at scale, but it is still cheaper than replaying that lab’s full training run. If you can turn millions of teacher answers into fine-tuning data, you buy capability on someone else’s bill. You also risk copying style and behavior without copying the teacher’s safety stack. Anthropic says that gap matters: students can pick up useful skills while shedding refusals and other guardrails.

Anthropic previously warned about distillation campaigns in February 2026. OpenAI has separately accused DeepSeek of similar activity. The September report says the campaigns it measured this year were larger and more aggressive, totaling nearly 200 million exchanges across five distillation campaigns Anthropic attributed to China-based labs.

What Anthropic says happened to Claude

Anthropic said operators it linked to Alibaba ran more than 151 million Claude exchanges between May and July 2026, peaking near 3 million a day across more than 3,500 accounts it called fraudulent. The company said those exchanges were aimed at improving Alibaba’s Qwen models, and that a shared fixed prompt was used to pull out chain-of-thought reasoning for training material. Anthropic called that the largest wholesale distillation effort it has measured.

It also accused Moonshot, the company behind Kimi, and DeepSeek of routing live customer chats through Claude without telling users, then using responses, including reasoning transcripts, as training data. Anthropic attributed more than 23 million exchanges to Moonshot between May and July, including nearly 300,000 customer requests relayed over one 10-day stretch through thousands of accounts. It attributed more than 12 million DeepSeek-linked exchanges over 14 days in July. Some routed traffic, Anthropic said, included sensitive information from individuals, companies, and state-affiliated actors. The company said those practices are likely inconsistent with privacy laws and the labs’ own terms of service.

The misuse Anthropic described hit Claude Haiku, Sonnet, and Opus. Fable and Mythos were not implicated except in one distillation case, the report said. Anthropic disrupted the activity, banned accounts, and said it is tightening detection, identity checks, and defenses against extraction. Alibaba, Moonshot, and DeepSeek did not immediately comment to Reuters or CNBC. China’s foreign ministry told Reuters it was not aware of the report and that Beijing opposes smears against the country.

Those numbers and attributions are Anthropic’s. Independent reporters have not publicly verified the technical forensics. Treat them as company claims backed by a detailed threat report, not as courtroom findings.

Why Claude is an attractive teacher

Frontier models are scarce teachers. Claude sits in the same competitive band as other U.S. labs’ flagship systems: strong coding, tool use, long-horizon agent work, and logical reasoning. Anthropic said the campaigns targeted exactly those valuable capabilities. For a lab racing to improve Qwen, Kimi, or DeepSeek, a teacher that already solves hard agentic and software tasks is a shortcut to synthetic data that looks like frontier behavior.

Claude is also widely available through APIs and products, which makes bulk querying possible until defenses catch up. Safety tuning does not make a model useless as a teacher. A model that refuses the worst prompts can still produce high-quality code, plans, and reasoning traces on ordinary work. Distillers, Anthropic argues, want that useful middle: capable outputs without paying the full cost of building the teacher, and without necessarily inheriting the teacher’s full refusal profile.

Competitive pressure sharpens the incentive. Chinese open-weight and commercial labs have spent the past two years closing gaps with U.S. frontier systems on public benchmarks and product features. If Anthropic’s account is right, illicit distillation is one way to buy progress in coding and agents without waiting for a full training cycle. Legitimate distillation from a lab’s own models, or from licensed partners, does not raise the same authorization and privacy issues. Covert API scraping and silent customer routing do.

What this is not

This is not a claim that every Chinese lab steals models. Anthropic named specific operators and campaigns. Distillation itself is not illegal in the abstract. The dispute Anthropic is raising is authorization, deception, and scale: fake accounts, silent proxies of user traffic, and training on another company’s outputs without a deal.

It is also not a technical blueprint. Anthropic described prompts that extract reasoning, account farms, and routing of customer chats. It did not publish a how-to for building a student model, and this desk will not invent one.

For the full threat report beyond distillation, including bio-risk and Russia-linked cyber cases, start with Anthropic’s post and our news brief. The distillation fight is about who gets to learn from whom, and who pays for the lesson.

Sources

Story so far

0 0 votes
Article Rating
Subscribe
Notify of
0 Comments
Inline Feedbacks
View all comments
0
Would love your thoughts, please comment.x
()
x