Center for Cyber Diplomacy and International Security

The Distillation Wars: How AA26-251A Turned Attribution Into America’s New Weapon of AI Diplomacy

Published by

on

By Vladimir Tsakanyan, PhD · Center for Cyber Diplomacy and International Security · cybercenter.space


Executive Summary

On September 8, 2026, three of the United States’ most powerful security institutions — the National Security Agency, the Cybersecurity and Infrastructure Security Agency, and the Federal Bureau of Investigation — did something they have historically done only when the stakes were enormous: they put names on paper. In joint cybersecurity advisory AA26-251A, the agencies publicly accused six Chinese artificial intelligence companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI — of conducting “aggressive, malicious, and targeted” knowledge distillation campaigns at industrial scale against America’s frontier AI models, “likely with Chinese government awareness.”

The document is technically rich: it names the models distilled (Claude, GPT, Gemini, Grok), maps the tactics to the MITRE ATLAS framework, and details a gray market of API proxies called “transfer stations” that bypass geographic restrictions. But its real significance is not technical. It is political.

AA26-251A marks three simultaneous shifts in American AI statecraft. First, attribution is now a diplomatic instrument: the advisory reads less like a defensive playbook and more like an evidence package — assembled weeks before President Trump hosts Chinese leader Xi Jinping in Washington, with Treasury Secretary Scott Bessent already threatening sanctions and Entity List designations. Second, the locus of control is moving from chips to the models themselves: export controls on semiconductors were the supply-side strategy; the advisory concedes that the battlefield has migrated to API endpoints, where America’s most capable systems are being systematically harvested. Third, and most uncomfortably, Washington is deputizing private companies to defend themselves — asking them to quietly degrade the quality of their own products for suspected attackers, while acknowledging that the legal basis for punishment barely exists.

This analysis argues that AA26-251A is best understood not as a cybersecurity advisory but as the opening move in a new phase of the AI contest: the criminalization-by-attribution of a technique that is, by everyone’s admission, entirely standard practice — distinguished from ordinary research only by scale, intent, and evasion. That distinction may be defensible. But enforcing it will require Washington to do something it has so far avoided: decide what, exactly, is illegal about learning from someone else’s machine.


1. What the Advisory Actually Says

Strip away the geopolitics and the advisory’s claims are strikingly specific — a level of detail that itself carries a message. The authoring agencies assert that since at least late 2024, the six named Chinese companies have extracted billions of tokens across millions of exchanges from U.S. frontier models, and that distillation forms “the core — not merely a supplement” of their development strategy.

The attribution section reads like an intelligence dossier. DeepSeek, the agencies say, has run an organized campaign since late 2024 targeting reasoning capabilities and specialized optimizations, distilling Claude 3.7, Sonnet 4, Sonnet 4.5, Opus 4.1, Gemini 2.5 Pro and Flash, GPT-4, GPT-4o, GPT-5, and Grok 4 to train its R1 and V3 models. The advisory pointedly notes that DeepSeek’s famously quoted $5.6 million training cost is “misleading” because it excludes the true cost of the data acquired through malicious distillation. Moonshot AI extracted Claude Fable 5 data for its Kimi-K3 and GPT-4o data for its Kimi-K2. Alibaba distilled Claude-4, Claude Opus, Claude Sonnet, and GPT-5 to improve the Qwen family’s software engineering and customer-service capabilities. MiniMax pulled chain-of-thought reasoning, reinforcement learning signals, and software engineering capability from Claude Code, Sonnet 4, Opus 4.5, and Gemini models — and even used prompt injections to trick Claude Code into believing it was a MiniMax product. StepFun and Z.AI targeted GPT-5.x and Claude Opus variants for coding and agentic functions.

The tactics section, mapped to the MITRE ATLAS framework, describes a mature industrial operation: a gray market of API proxies — “transfer stations” — reselling access to frontier models at a fraction of the official price and obfuscating user metadata; bulk procurement of premium subscriptions shared across developer teams; chain-of-thought extraction via prompts that coax models into revealing their hidden internal reasoning; automated failover between access pathways during blocking attempts; and production-grade quality-assurance pipelines to detect defensive countermeasures. MiniMax, the agencies note, redirected its campaigns to a newly released Claude model within 24 hours — evidence of real-time provider monitoring and pre-positioned infrastructure.

The specificity matters. In February 2026, Anthropic had already disclosed that three Chinese labs — DeepSeek, Moonshot, and MiniMax — generated more than 16 million exchanges with Claude through roughly 24,000 fraudulent accounts. The advisory converts a private company’s complaint into the official position of the U.S. intelligence community, backed by three-letter-agency authority. That escalation from corporate grievance to state accusation is the political event; the technical details are the ammunition.

2. Attribution as Statecraft

Cybersecurity advisories are ordinarily defensive instruments: they warn, they recommend, they move on. AA26-251A is something else. It is an indictment-shaped document — and its timing reveals its purpose.

The advisory landed on September 8. President Trump is scheduled to host Xi Jinping in Washington later this month, with AI explicitly on the agenda, and the two countries are preparing a separate AI safety dialogue for mid-September, according to Reuters. Publishing a detailed, named accusation of state-tolerated industrial-scale theft days before those meetings is not a coincidence. It is leverage — a public evidence package that enters the summit as a pre-negotiated fact, whether or not Beijing accepts it.

The choreography extended beyond the agencies. Treasury Secretary Scott Bessent had already threatened, in July and again after the advisory, that “when [People’s Republic of China] firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table.” The sequence is deliberate: intelligence agencies establish the factual predicate, the Treasury signals the economic weapon, and the White House holds the summit where both can be deployed. This is how the United States now does technology diplomacy — not through quiet démarches, but through public attribution cascades designed to corner an adversary into negotiating from a position of documented guilt.

Beijing responded exactly as the playbook predicts. The Chinese Embassy in Washington called the advisory a “deliberate attack on China’s development and progress in the AI industry.” China’s Commerce Ministry said the allegations “lacked factual basis and legal grounding,” accused the U.S. of double standards and interference in normal commercial activity, and called distillation a “neutral technical method in nature” — adding, pointedly, that some American companies had “extensively distilled Chinese models.” Foreign Ministry spokesperson Mao Ning framed China’s AI progress as self-reliance and urged cooperation. None of this is surprising. What matters is that both sides now agree the fight is about distillation — Washington calls it theft, Beijing calls it technique. The dispute has moved from whether it happens to what it means, which is precisely where Washington wants it: in the realm of norms, where American standard-setting power is strongest.

3. The Doctrine Shift: From Chips to the Models Themselves

For four years, American AI containment strategy has been supply-side: deny China the advanced semiconductors needed to train frontier models, and the capability gap takes care of itself. The export control regime — expanded, tightened, and diplomatically expensive to maintain across allied capitals — rested on a simple assumption: compute is the chokepoint.

AA26-251A quietly concedes that the chokepoint has moved. If Chinese labs can extract the capabilities of a frontier model without training one from scratch — by treating American APIs as a free research and development department — then chip controls slow but do not stop the catch-up. As Anthropic itself argued in February, export controls matter precisely because they reduce “both direct model training capabilities and the extent of improper distillation.” The advisory is the government’s acknowledgment that the second half of that sentence is now the main event.

This is a profound and uncomfortable shift. Denying chips was an industrial policy problem: measurable, enforceable at customs, legible to allies. Defending against distillation is a counterintelligence problem that lives inside the business model of the very companies Washington is trying to protect. Every API endpoint, every chatbot interface, every cloud-hosted inference service is now a potential exfiltration surface — not for data the company holds, but for the capabilities the company is. The asset being stolen is not a file. It is the model itself, extracted one conversation at a time, at a scale no human operator could sustain and no per-user rate limit was designed to catch.

The same week the advisory dropped, Google’s Threat Intelligence Group released a report finding that hackers “of various stripes” are increasingly targeting AI assets — “proprietary AI models and source code,” model weights, cloud compute quotas — as high-value targets for espionage, extortion, and resource theft (CNN). The convergence is not accidental: the AI system itself has become the crown jewel, and the industry’s attack surface has expanded from the networks around the model to the model as the network.

4. The Deputized Defender

Perhaps the most revealing passage of AA26-251A is its recommendations section — not for what it asks, but for what it admits it cannot do. The agencies recommend three actions for U.S. AI companies: implement comprehensive detection of anomalous prompts, accounts, and usage patterns; establish cross-organization intelligence sharing across model providers, cloud platforms, and API aggregators; and — most strikingly — “subtly alter responses for suspected malicious distillation attempts to attenuate the payoffs” of the campaigns.

Read that second recommendation carefully. The United States government is advising its most innovative companies to deliberately degrade the quality of their own products — to serve poisoned, subtly wrong answers to suspected attackers. This is defensive deception as official policy, and it carries a cost the advisory does not discuss: every altered response risks hitting a legitimate user, every detection heuristic risks false positives, and every company that implements “subtle” sabotage must build the infrastructure to do so — a quiet tax on the American AI industry, levied in the name of defending it.

Meanwhile, the advisory is candid about the limits of enforcement. The “transfer stations” — the gray market of API proxies that resell frontier-model access at a fraction of the official price — operate across jurisdictions, through third-party aggregators that “automatically obfuscate user metadata.” Shutting down one proxy spawns three more; blocking one account pattern teaches the adversary to rotate faster. The advisory describes an opponent with production-grade automation, 24-hour retargeting cycles, and quality-assurance pipelines sophisticated enough to distinguish service issues from defensive countermeasures. Against that, it offers the digital equivalent of “watch your meters and share what you see.”

There is a deeper structural problem. Distillation at scale is profitable for the intermediaries, useful for the buyers, and — until now — nearly costless for the perpetrators. The advisory does not propose any new legal authority, any new liability regime, or any mechanism to make the gray market uneconomical. It deputizes the victims and wishes them luck. That is not a strategy. It is an admission that the government has not yet figured out what a strategy would look like.

5. The Law Is Missing

Here is the question AA26-251A never quite answers: what, exactly, is illegal about what the six companies did?

Distillation — training a smaller model on a larger model’s outputs — is not a crime. It is not even novel. Every major AI lab in the world, American ones included, uses it routinely on its own models. The advisory itself concedes it is “a legitimate and useful technique in AI research.” The line between legitimate and “malicious” distillation in the advisory rests on three things: violation of terms of service, evasion of geographic restrictions, and the use of fraud (fake accounts, proxies). These are contract violations and access-control evasions — real wrongs, but a long way from espionage, and a longer way from the sanctions and Entity List designations Bessent is threatening.

This is the legal vacuum at the heart of the new doctrine. American companies’ terms of service prohibit distillation for competing-model development, and lawyers have suggested U.S. labs could sue Chinese rivals for breach of contract. But as the Wall Street Journal notes, gathering conclusive evidence is difficult when the activity runs through overseas middlemen, and prolonged litigation “offers little upside in an industry moving at breakneck speed.” Sanctioning a company for violating a click-through agreement would be an unprecedented escalation of economic statecraft — and one that invites the obvious retaliation: American cloud providers, chip designers, and software firms all operate in China under Chinese rules they do not always follow to the letter either.

Beijing’s “double standards” charge is propaganda, but it is propaganda with a working surface. The American AI industry’s own training-data practices — models built on vast corpora of copyrighted books, images, and code, defended with the argument that the machines are “learning, not copying” — make the moral line harder to draw than Washington would like. The U.S. position is that scale, deception, and state awareness transform a research technique into theft. That may be right. But it is a standard defined by degree, not by kind, and standards defined by degree are enforced by power, not by law. Which is, of course, exactly what the sanctions threat is for.

6. The Success Tax

There is a final, structural irony that no advisory can fix. The reason Chinese labs distill American models is that American models are worth distilling. Every capability gap that makes Claude, GPT, Gemini, and Grok the world’s best AI systems simultaneously makes them the world’s most valuable teachers. Distillation campaigns are, in economic terms, a success tax: the better your model, the more of the world’s adversaries will pay the API bill to learn from it.

This dynamic punishes openness in two directions. Frontier labs face a choice between keeping models behind ever-tighter access controls — which slows the ecosystem that made them dominant — or accepting that every public capability will be harvested within months. The open-weights movement, which has made Chinese models like DeepSeek’s and Alibaba’s Qwen globally popular and commercially competitive, complicates the picture further: open models accelerate diffusion of capabilities to everyone, including American startups that now route real work through cheaper Chinese systems. Washington is threatening to blacklist the very models its own companies are increasingly using.

The uncomfortable truth is that distillation works because the teacher-student paradigm is how the entire field advances. Today’s American frontier model was trained, in part, on synthetic data generated by yesterday’s frontier model. The difference between “our” distillation and “their” distillation is ownership of the teacher — a distinction that holds in court, perhaps, but dissolves in the training loop. A containment strategy built on preventing learning from American machines is, at some level, a strategy against the diffusion of knowledge itself. It can slow the adversary. It cannot stop the dynamic.


Conclusion: An Evidence Package in Search of a Strategy

AA26-251A should be read for what it is: not a defensive manual but a diplomatic instrument. Its 483 lines of attribution detail, its MITRE ATLAS mappings, its named companies and named models — these are not written for the security engineer. They are written for the sanctions lawyer, the summit negotiator, and the allied capital deciding whether to join an American-led enforcement regime. The advisory is the factual predicate for the economic and diplomatic campaign that Bessent previewed and the Trump–Xi summit will test.

Whether that campaign succeeds depends on questions the advisory leaves unanswered. Can the United States define, in law rather than in rhetoric, where legitimate distillation ends and theft begins? Can it impose costs on a gray-market infrastructure that spans jurisdictions faster than indictments can? Can it ask its own AI champions to degrade their products and share their intelligence without taxing the innovation it claims to protect? And can it sustain an allied coalition for model-level containment when the models in question are embedded in the commercial workflows of those same allies?

The cyber arms race, as this publication has argued, has entered the AI model itself. AA26-251A is Washington’s first attempt to govern that new battlefield — and it reveals, with unusual honesty, how much of the governing has yet to be invented. Attribution was the easy part. Everything after attribution is the actual policy. The summit in Washington will show whether the United States has one.


Vladimir Tsakanyan is a cybersecurity policy analyst and political commentator covering cyber diplomacy, geopolitical threat intelligence, and the intersection of technology and national security.

Sources:


Discover more from Center for Cyber Diplomacy and International Security

Subscribe to get the latest posts sent to your email.

Leave a comment

Discover more from Center for Cyber Diplomacy and International Security

Subscribe now to keep reading and get access to the full archive.

Continue reading