US agencies call distillation the core of China's AI strategy
Two intelligence agencies and CISA put their names on an accusation the industry has made privately for a year — and attached a sanctions threat to it. Elsewhere, JD Cloud sketched out a 100,000-card cluster built entirely on Chinese silicon.
The NSA, FBI and CISA released a joint advisory on Tuesday accusing Chinese AI companies of running "industrial-scale" distillation campaigns against U.S. frontier models, and Treasury Secretary Scott Bessent immediately put sanctions on the table. The language is what matters here: the agencies write that the "sheer scale of these campaigns and their sophistication indicate that distillation is not a supplement to these companies' AI model development, but the critical core of it." DeepSeek and Alibaba are named directly, along with a description of the plumbing involved — requests routed through native APIs, remote cloud providers and third-party aggregators that strip identifying metadata, plus a gray market of resellers the advisory calls "transfer stations" that resell frontier model access at a fraction of the official price. Bessent's line was blunt: where distillation crosses into IP theft, "sanctions and Entity List designations will be on the table."
The accusation that will travel furthest is the one about DeepSeek's origin story. The advisory claims DeepSeek used distillation to generate the synthetic training data behind its models, which would make its celebrated claim of frontier performance on trivial compute spending false — and that claim, remember, is what wiped hundreds of billions off infrastructure stocks in early 2025. The recommended countermeasures are quieter than the rhetoric: subtly degrade responses to suspected distillation traffic so the stolen capability is subtly wrong, watch for new accounts that immediately max out usage, and correlate activity across providers to spot distributed campaigns. We covered the commercial end of this fight in Anthropic locks Claude's thinking blocks to kill API distillation and the infrastructure behind it in Deep Dive — Anthropic goes public with the dark-web distillation pipeline. What's new is that this is now a national-security instrument rather than a terms-of-service dispute — and that the remedy the government proposes is deliberately poisoning the output.
JD Cloud says it will build a 100,000-card cluster on Moore Threads' Chinese-made GPUs, the first time domestic silicon has been picked for a leading Chinese AI cloud's largest training cluster. Announced at the JD Discovery conference on Wednesday by Cao Peng, chairman of JD's technology committee, the plan follows an existing 10,000-card domestic cluster built with the same partner, and targets model training, inference and embodied intelligence. Moore Threads founder Zhang Jianzhong framed it in the terms the whole industry argues about — "Scaling Law still holds, and 100,000-card clusters are the inevitable trend" — while Cao pitched domestic compute as the "core pillar" of JD's physical AI strategy.
Two details make this more than a procurement note. First, it lands days after China's Ministry of Industry and Information Technology called for the orderly deployment of 10,000- and 100,000-card clusters with greater use of domestic chips, so the plan is policy-aligned as much as it is commercial. Second, JD paired it with a model: JoyAI-Echo WM, a real-time interactive world model that it says scored 81.6 on the WBench Navigation benchmark, plus a 10-million-hour embodied data collection effort whose first open dataset, EgoLive, has drawn access requests from over 100 universities across eight countries. JD also shipped logistics and industrial models claiming a 96.7% task success rate in multi-task logistics scenarios; we covered the previous generation in JD Logistics' Super Brain 3.0 now runs its warehouses end to end. The throughline is that China's largest-commerce operator is trying to own the whole stack — chips, data, world models — inside one procurement cycle.
Arm used its Everywhere China event to launch an AI Portal, a catalogue of Arm-optimized models with latency, memory and accuracy figures so developers can compare candidates before deploying. The pitch, per Arm's developer-relations vice president Shantu Roy, is that model choice shouldn't be a small-versus-large argument but a task-first one: a phone notification summary and six-document strategy analysis are different jobs, and a model that looks fine on a laptop can collapse under a phone's thermal and memory limits. A companion profiling tool, Performix, feeds performance data back into coding agents so an assistant can read where the time goes. It's early-access and tied to Hugging Face, and it's more platform hygiene than news — but a searchable, benchmark-backed model catalogue aimed at agentic consumers rather than humans is the shape of how model selection will actually get done.
What to watch: whether any Chinese lab is formally named in sanctions, and whether the "transfer stations" the advisory describes start getting shut off.
Should degrading model output for suspected distillation traffic be standard practice, or does it poison the product for everyone? Tell us in the comments.
Sources: NSA — Advisory on China-based AI companies' distillation campaigns · CNN — US claims Chinese AI firms are carrying out 'industrial-scale' theft · The Register — US claims Chinese AI companies' core AI strategy is distilling American models · Guancha — 京东规划建设国产十万卡集群 · Sina Tech — 京东规划建设10万卡国产算力集群 推出JoyAI世界模型 · BigGo Finance — JD Cloud and Moore Threads to Build 100,000-Card China-Made GPU AI Cluster · Zhidx — Arm发了个"挑模型神器" · Arm Newsroom — Arm unveils Arm AI Portal