The Truth About China Distilling U.S. AI Models
Dan PetersonSynopsis — AI-drafted from Dan's notes
An eight-minute video that grants the premise and disputes the conclusion. Chinese labs did distill US models: Microsoft and OpenAI said so about DeepSeek in January 2025, and Anthropic’s February 2026 disclosure attributed more than 16 million exchanges, through about 24,000 fraudulent accounts, to DeepSeek, Moonshot and MiniMax. Anthropic’s own figures put MiniMax at about 13 million of those and DeepSeek at about 150,000, and say MiniMax moved nearly half its traffic to a new Claude model within a day of its release. The video’s case is that copying a teacher’s outputs can only bring a student up to the teacher, citing the 2025 paper “SFT Memorizes, RL Generalizes”, so distillation cannot explain how the gap closed.
It credits the labs’ own research instead: DeepSeek’s GRPO and efficient mixture-of-experts architecture, Moonshot’s Kimi K2 (384 experts, trained with the MuonClip optimizer on 15.5 trillion tokens), and Alibaba’s QwQ reasoning model. Two of these details need correcting. GRPO was introduced in DeepSeekMath (February 2024), not with R1, and what it drops is the separate critic (value) model rather than a step-by-step grader. And QwQ’s preview (November 2024) did not precede the American reasoning models: OpenAI’s o1-preview shipped that September, and QwQ was benchmarked against it. The Kimi K2 figures match Moonshot’s technical report.
For the gap itself it cites the Stanford AI Index 2026, which does say the US–China gap has effectively closed. The 2.7% is an Arena leaderboard margin in March 2026 (Claude Opus 4.6 over ByteDance’s Dola-Seed-2.0 Preview), and the earlier 17.5 to 31.6 point gap is from May 2023, nearly three years before. It also reads Chinese labs’ retreat from open weights as a sign they no longer need the shortcut; Alibaba’s record in 2026 is mixed, with its Plus and Max flagships proprietary while smaller Qwen models still ship under Apache 2.0. Its explanation for the accusations is political: a story about stolen intellectual property is easier to take to Congress in the case for chip export controls, and Anthropic’s February post does tie distillation to its support for those controls.
What the video leaves out is that the accusation had widened by the time it was posted. On 8 September 2026 the NSA, CISA and FBI named six Chinese firms, Alibaba among them, and called distillation “the core—not merely a supplement” of their strategy (advisory AA26-251A), though without publishing forensic evidence. Two days later Anthropic’s September threat report attributed 151 million exchanges between May and July 2026 to an Alibaba campaign feeding Qwen, the line the video treats as homegrown. These remain allegations. EdgeTheory’s assessment (31 August 2026) found the public evidence does not show that Chinese frontier models were primarily built by distillation, which is the narrower claim the video can defend.
Connections
Links
- https://www.youtube.com/watch?v=zTHaF8s0tTw
- https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks
- https://www.cnbc.com/2026/02/24/anthropic-openai-china-firms-distillation-deepseek.html
- https://techcrunch.com/2025/01/29/microsoft-probing-whether-deepseek-improperly-used-openais-api/
- https://arxiv.org/abs/2501.17161
- https://arxiv.org/abs/2402.03300
- https://arxiv.org/abs/2507.20534
- https://hai.stanford.edu/ai-index/2026-ai-index-report
- https://thenextweb.com/news/stanford-ai-index-2026-china-us-performance-gap
- https://en.wikipedia.org/wiki/Qwen
- https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a
- https://www.nbcnews.com/tech/tech-news/us-accuses-china-ai-developers-deepseek-alibaba-copying-american-ai-rcna596696
- https://techcrunch.com/2026/09/10/anthropic-details-distillation-campaigns-from-alibaba-moonshot-ai-and-deepseek/
- https://edgetheory.com/resources/ai-distillation-attacks-china-us-ai-race