The United States’ NSA, CISA and FBI have issued a joint cybersecurity advisory alleging that several Chinese artificial-intelligence companies conducted model-distillation activity on an enormous scale. The advisory says the activity sought to extract proprietary functions and capabilities from leading American AI systems, framing the issue as both a commercial and cybersecurity concern.
The companies named in the warning are DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI. The agencies allege that, since 2024, these firms pulled billions of tokens through millions of exchanges or requests involving frontier models from Anthropic, OpenAI, Google and xAI. The systems cited include Claude, GPT, Gemini and Grok.
That is a sizeable allegation, and it is important to keep the wording straight: these are accusations set out by US agencies, not a public technical finding that independently establishes each claim. Still, the advisory puts unusual governmental weight behind a dispute that AI developers have increasingly raised themselves as powerful models become more accessible through consumer and enterprise interfaces.
What “distillation” means in this dispute
Distillation has a legitimate technical meaning in machine learning. Broadly, a more capable “teacher” model can provide outputs that help train a smaller, newer or more specialized “student” model. Used under authorized conditions, the approach can be a practical way to transfer useful behavior and make systems cheaper or faster to run.
The conflict arises when a company allegedly uses another provider’s model outputs at volume, without permission, to reproduce valuable behavior in its own product. In that scenario, the question is not merely whether an answer resembles another chatbot’s answer. It is whether huge numbers of interactions, prompts and responses were gathered in a manner intended to replicate model performance, capabilities or protective behavior.
For the companies building expensive frontier systems, the stakes are clear. Developing leading models can require substantial computing capacity, specialized research work and long training cycles. If a rival can systematically query an existing system and use the answers as training material, the original developer may view that as an attempt to bypass part of that investment.
There is also a security dimension. Advanced models may have specialized knowledge, safeguards, tool-use patterns and reasoning behaviors that their operators do not want extracted at scale. The joint advisory focuses on the alleged loss of proprietary capabilities, while also offering mitigations intended to help US organizations defend against what it calls malicious distillation campaigns.
The models and companies identified in the advisory
The advisory specifically alleges that DeepSeek used data and capabilities from multiple Claude, Gemini, GPT and Grok models in training its own systems, including DeepSeek R1. R1 is the company’s open-source reasoning model released in early 2025, and it became a major part of public discussion around Chinese AI development.
Moonshot AI is also singled out. The agencies allege that significant data was extracted from Claude Fable to train Kimi K3. Kimi K3 has been described as among the most advanced AI models made by a Chinese company. The warning additionally says that Moonshot’s Kimi K2 used data from GPT-4o.
Claude Fable requires a little context. Anthropic created it to make some capabilities of Mythos, its advanced cybersecurity model, available to the public. Mythos itself is reserved for Project Glasswing participants. That distinction matters because cybersecurity-focused capabilities can be particularly sensitive: a system built to help understand, analyze or respond to security work is not simply another general-purpose text generator.
The advisory also names Alibaba, MiniMax, StepFun and Z.AI in its wider list of alleged participants, while laying out the American models it says were targeted. The material supplied with the advisory does not provide a full public technical accounting of every individual allegation or a response from every named company. Readers should therefore avoid treating all of the claims as identical in scope or evidence merely because the companies appear in the same warning.
Why the token volume is central
The most consequential detail is the asserted scale: billions of tokens across millions of requests. Tokens are small chunks of text processed by AI models; they can be words, fragments of words, punctuation or other units. A few ordinary chats are one thing. A sustained operation involving millions of exchanges is something else entirely.
At that scale, an operator could potentially collect a broad training set spanning programming, explanation, writing, reasoning, multilingual requests and many other task types. Repeatedly varying prompts and comparing responses can also reveal how a system behaves under different conditions. The objective need not be perfect one-for-one duplication for the collection to be commercially useful. A large response corpus may be valuable for tuning another model toward a desired level of competence.
That helps explain why ordinary rate limits, account checks and anomaly detection matter so much to AI providers. The alleged campaigns described by US authorities are not about a person asking a chatbot to draft an email or solve a small coding problem. They concern patterns of access that, if accurately characterized, could turn a public model interface into a pipeline for training a competing system.
An argument that predates the advisory
American AI companies have been voicing concerns over suspected copying and distillation for some time. After DeepSeek rose to the top of the free iPhone app chart early in the previous year, OpenAI said it and Microsoft were banning accounts suspected of using its technology for distillation. DeepSeek was among the companies being investigated at that point.
Anthropic also accused DeepSeek, Moonshot and MiniMax earlier this year of wide-ranging distillation efforts. The new joint advisory does not emerge from nowhere; it places federal cybersecurity agencies alongside an argument that had already moved through corporate statements and industry debate.
It is also not an accusation that can be cleanly reduced to geography alone. In OpenAI’s announcement about removing its models from Cursor, the company noted that Elon Musk acknowledged using OpenAI output to train xAI models during cross-examination in the lawsuit he brought against OpenAI. The wider issue is therefore about the boundaries around model outputs, platform access and competitive training practices—not simply a single country or a single group of companies.
That wider context is useful when considering the policy implications. AI builders are competing in a field where model outputs are useful, accessible and often difficult for outsiders to trace once absorbed into a training pipeline. Policies that prohibit competitors from training on outputs may be clear on paper, but detecting and proving violations can be much harder than spotting conventional software copying.
What mitigation may look like
The advisory includes mitigations for US firms confronting suspected malicious distillation. Although the supplied information does not enumerate every recommendation, the basic defensive direction is easy to understand: providers need to identify automated or coordinated use patterns that appear designed to harvest outputs rather than serve normal customers.
- Account and access controls: Providers can scrutinize suspicious account behavior and take action against accounts that violate usage restrictions.
- Rate and volume monitoring: Unusual request frequency, sustained token consumption and highly repetitive patterns can be signs that warrant closer review.
- Detection of coordinated activity: A campaign may be distributed across accounts, infrastructure or prompt styles, making behavioral analysis more useful than looking only at one user.
- Protecting sensitive capabilities: Systems with advanced cybersecurity functions or other high-value behavior may need particularly careful access controls.
None of these measures is a magic shield. Overly aggressive limits can harm legitimate developers, researchers and businesses that use an AI service heavily. Too little scrutiny can make large-scale extraction cheaper and easier. The challenge for providers is to recognize the difference without creating a customer experience that feels like trying to enter a dungeon after forgetting every password, key and quest item.
The dispute also arrives while AI tools are becoming more common in everyday services. As assistants move into planning, shopping, operating-system features and workplace software, the value of a model’s distinctive capabilities rises alongside the temptation to reproduce them. For a more consumer-facing look at what responsible use of a major model can involve, see our take on Gemini as a travel-planning co-op partner.
A sharper line around AI competition
The joint NSA, CISA and FBI advisory is a notable escalation in the public framing of AI distillation. It takes a practice often discussed as a terms-of-service, intellectual-property or competitive issue and describes alleged activity at a scale serious enough to merit a cybersecurity warning.
The named companies, the models cited and the stated timeline will all attract scrutiny. So will the safeguards adopted by US AI companies in response. What is already apparent is that frontier-model competition is no longer only about chips, research talent and data centers. It is also about who can access a model, what can be learned from its answers and whether those answers can become the raw material for the next rival system.






