NSA, CISA, FBI Accuse DeepSeek, Moonshot AI, Alibaba of Industrial-Scale US AI Model Theft

·
·
4 min read
·
AI-assisted
Author Profile
by Albert Schaper
Share
NSA, CISA, FBI Accuse DeepSeek, Moonshot AI, Alibaba of Industrial-Scale US AI Model Theft

On September 8, 2026, the NSA, CISA, and FBI issued joint Cybersecurity Advisory AA26-251A, accusing six Chinese AI firms — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, of industrial-scale "knowledge distillation" campaigns against U.S. frontier AI models like Claude, GPT, Gemini, and Grok since late 2024, reframing this activity as a national security threat.

Allegations of Industrial-Scale Distillation

The joint advisory details how the named Chinese AI firms allegedly extracted billions of tokens through millions of requests from various U.S. frontier AI models. This large-scale data extraction is described as a "knowledge distillation" process, where the output and behavior of a larger, more advanced model are used to train a smaller, more efficient model. The U.S. agencies assert that this method significantly reduces the training costs and development timelines for the Chinese companies involved.

Specific allegations include DeepSeek using distilled data to train its R1 and V3 models, Moonshot AI for its Kimi K2/K3 models, and Alibaba to enhance its Qwen family of models. MiniMax was also cited for allegedly using prompt injection techniques to manipulate Claude Code into believing it was a MiniMax product, further illustrating the sophisticated methods reportedly employed.

Tactics and Techniques Employed

The advisory outlines several tactics allegedly used by these firms to facilitate their distillation campaigns:

  • Gray-Market "Transfer Station" API Proxies: Utilizing intermediary services to obscure the origin of requests and bypass direct access restrictions.
  • Bulk-Shared Premium Subscriptions: Employing numerous shared premium accounts to gain extensive access to U.S. AI models.
  • Chain-of-Thought Reasoning Extraction: Techniques designed to extract not just final answers but also the step-by-step reasoning processes of the frontier models.
  • Automated Failover: Implementing systems to automatically switch between different access points or accounts when one is detected or blocked, ensuring continuous data extraction.

These methods collectively point to a coordinated and persistent effort to use U.S. AI advancements without incurring the full development costs associated with building such models from scratch.

National Security Implications and Misleading Costs

The U.S. government's advisory elevates knowledge distillation from a technical concern to a national security issue. It argues that by relying on distilled data, these Chinese firms are not only gaining an unfair competitive advantage but are also potentially accelerating their AI capabilities in ways that could have strategic implications. The advisory suggests these campaigns are likely conducted with the knowledge of the Chinese government.

A notable point in the advisory concerns DeepSeek's widely reported R1 training cost of $5.6 million. The U.S. agencies contend that this figure is misleading because it omits the substantial value derived from the allegedly stolen distilled data, which would otherwise represent a significant investment in research and development.

Recommended Mitigations for AI Labs

In response to these alleged activities, the U.S. government is urging AI laboratories and developers to implement several mitigation strategies:

  • Detect Anomalous Usage: Enhance monitoring systems to identify unusual patterns of interaction, such as high-volume, repetitive queries or requests originating from suspicious IP addresses or proxy networks.
  • Subtly Degrade Responses: For suspected distillation attempts, AI models could be programmed to subtly alter or degrade the quality of their responses, making the extracted data less valuable for training purposes without immediately alerting the perpetrator.
  • Share Cross-Provider Intelligence: Establish mechanisms for AI providers to share threat intelligence and indicators of compromise with each other, fostering a collective defense against such campaigns.

These recommendations aim to create a more robust defense against ongoing and future distillation efforts, protecting the intellectual property and strategic advantage of U.S. AI innovation.

Conclusion

The joint advisory AA26-251A from the NSA, CISA, and FBI marks a significant escalation in the U.S. government's stance on AI intellectual property and national security. By directly accusing six Chinese AI firms of industrial-scale model distillation, the advisory highlights the growing concerns over the methods used to accelerate AI development and the potential strategic implications. AI developers and organizations are now called upon to implement advanced detection and mitigation strategies to safeguard their frontier models against these sophisticated and persistent threats.

Sources

About the Author

Albert Schaper avatar

Written by

Albert Schaper

Albert Schaper is a co-founder of Best-AI.org. He focuses on product strategy, AI adoption, practical tool selection, and educational content that helps users compare AI products with clearer context.

More from Albert

Was this article helpful?

Found outdated info or have suggestions? Send us a note.

Discover more insights and stay updated with related articles

Discover AI Tools

Find your perfect AI solution from our curated directory of top-rated tools

Less noise. More results.

One monthly email with the industry news tools that matter - and why.

No spam. Unsubscribe anytime. We never sell your data. See our Privacy Policy.

What's Next?

Continue your AI journey with our tools and resources. Whether you're looking to compare AI tools, learn about artificial intelligence fundamentals, or stay updated with the latest AI news and trends, see what fits your needs. Explore our curated content to find the right AI tools for your workflow.