DeepSeek Used Distillation To Train Its R1 And V3 Models

Machine Learning

Since at least late 2024, companies including DeepSeek, Alibaba, and Moonshot AI have been systematically extracting data from U.S. frontier AI models like Claude, GPT, Gemini, and Grok. These China-based firms have reportedly processed billions of tokens across millions of requests using a technique called knowledge distillation, a practice the NSA, CISA, and FBI describe as forming “the core—not merely a supplement—of their AI development strategy.” DeepSeek, for example, has run organized distillation campaigns since 2024 specifically targeting reasoning capabilities and specialized optimizations to train its R1 and V3 models, while Alibaba used the technique to improve its Qwen family of AI models. This industrial-scale activity represents a systematic extraction of proprietary functionalities, posing a threat to U.S. technological leadership.

DeepSeek R1 and V3 Training via U.S. Model Distillation

DeepSeek ran an organized distillation campaign beginning in late 2024 to accelerate development of its R1 and V3 models, specifically targeting U.S. frontier AI models for synthetic training data. The company did not limit its efforts to broad data acquisition, but instead focused on extracting proprietary functionality and reasoning capabilities to minimize compute and research expenses.

Between late 2024 and mid-2025, DeepSeek distilled specialized training data from several U.S. models to enhance its R1 and V3 capabilities, including Gemini 2.5 Pro Preview and Gemini 2.5 Flash Preview. This targeted approach extended to specific knowledge domains, such as legal specialization and optimization, as well as API rule-driven tasks and writing utilizing chain-of-thought drafts.

DeepSeek also sought to improve question and answer optimization, coach/assistant capabilities. This industrial-scale activity is not unique to DeepSeek or Z.AI; companies like Moonshot AI, MiniMax, StepFun, and Alibaba also engaged in similar distillation practices. These China-based AI companies employ sophisticated tactics, techniques, and procedures to refine their models, including attempts to evade detection during data extraction.

Reducing reasoning depth, presenting correct information with altered reasoning, or introducing stylistic inconsistencies are all methods used to bypass safeguards while still acquiring useful training data. The source details a strategy to avoid alerting malicious distillers when a downgraded model is deployed, as informing them would allow them to improve their evasion techniques and determine when to resume training.

Instead, the recommendation is to subtly alter responses to users confirmed to be engaged in malicious knowledge distillation campaigns without providing explicit notification. This approach aims to validate the efficacy of detection mechanisms while disrupting the distillation process. The document also outlines several adversarial machine learning measures, including sanitizing inputs to prevent prompt injections, limiting the release of public information about model architecture and prompts, and utilizing adversarial training and defensive distillation to increase resistance to jailbreaking attempts.

It provides further detail on these techniques. The strategic use of distillation allows China-based AI companies to significantly shorten their AI development timelines and reduce financial expenditures, effectively bridging the technological gap with U.S. counterparts. This is achieved, in part, by distributing operations across multiple providers and platforms to avoid detection.

The document highlights that these companies attempt to distill the most valuable capabilities and proprietary features from each U.S. frontier model, integrating them into their own AI systems. This comprehensive strategy underscores the importance of robust defenses against knowledge distillation and the need for continuous monitoring of AI model vulnerabilities.

China-Based AI Firms’ Industrial Distillation Campaigns

DeepSeek used systematic quota and cost optimization to fuel its knowledge distillation campaigns, a tactic employed alongside other China-based AI firms like Alibaba Group and Moonshot AI. These companies prioritize cost-efficiency in their extraction of proprietary functionalities from U.S. AI models, achieving savings through bulk procurement of premium subscriptions shared across development teams. Detection of this behavior includes identifying new accounts exhibiting anomalously high immediate hit rates, suggesting pre-engineered template deployment at scale, and observing usage patterns optimized for cache maximization rather than diverse tasks.

Coordinated pathway switching in response to pricing or rate changes further indicates centralized decision-making within these organizations, allowing them to circumvent restrictions and maximize output. Beyond cost management, these firms demonstrate sophisticated operational strategies to avoid detection during the distillation process.

The authoring agencies, the NSA, CISA, and FBI, detail how China-based AI companies route requests through multiple pathways, including native application programming interfaces, remote cloud providers, and third-party aggregators, to obfuscate user metadata and evade monitoring. This distributed approach is further supported by the use of a gray market of proxies designed to bypass U.S. restrictions and maintain uninterrupted access to target models. Rather than deploying downgraded models and alerting those conducting the campaigns, the agencies suggest attenuating the payoffs by introducing inconsistencies or reducing reasoning depth.

The advisory also highlights the value of establishing cross-organization intelligence sharing, correlating activity across model providers, cloud platforms, and API aggregators to reveal the full scope of distributed campaigns. This collaborative approach, informed by frameworks like the MITRE ATLAS and NIST AI frameworks, is presented as essential for building a coordinated defense against these increasingly sophisticated threats.

Behavioral detection and monitoring of premium subscription usage are also recommended, specifically tracking enterprise-scale throughput, identifying accounts deviating from legitimate patterns, and flagging new accounts that immediately reach maximum usage levels. Companies like MiniMax, StepFun, and Z.AI also participated in these malicious knowledge distillation activities, targeting U.S. AI models.

The sheer scale of these campaigns, and the deliberate strategies employed to avoid detection, suggest a concerted effort to accelerate AI development through the systematic extraction of proprietary knowledge. The report concludes that a coordinated, ecosystem-wide response is necessary to effectively address these campaigns and protect U.S. AI innovation.

Distillation Tactics: APIs, Proxies, and Data Extraction

These companies also used a gray market of API proxies, referred to as “transfer stations,” to bypass U.S. safeguards and maintain operational anonymity. AI companies, shared across development teams, further reduced costs associated with these large-scale distillation efforts. Z.AI engaged in distillation activities since at least late 2024, targeting models including variants of Claude and GPT, to enhance its models.

This focus on chain-of-thought reasoning represents a key area of extraction, alongside specialized optimizations like legal expertise and API rule-driven tasks, as demonstrated by DeepSeek’s targeting of Gemini 2.5 Pro and Flash previews. The advisory details that these companies are not simply replicating functionality, but actively seeking to distill the “best capabilities and proprietary features” of each U.S. frontier model.

Advanced tactics employed by these China-based AI companies include chain-of-thought reasoning extraction, automated failover mechanisms to bypass blocking attempts, and sophisticated quality evaluation frameworks to identify and circumvent defensive countermeasures. These capabilities suggest a deliberate and coordinated effort to reverse-engineer and replicate the functionality of leading U.S. models.

Targeted Capabilities Extracted from U.S. Frontier Models

frontier AI models rather than relying solely on independent research and development. technology. The scale of these operations extends to billions of tokens extracted across millions of requests, highlighting the intensive nature of this data acquisition.

The specific capabilities targeted by these companies reveal an approach beyond simple data scraping. This granular focus suggests an attempt to replicate not just the output of U.S. Moonshot AI mirrored this approach, extracting significant data from Claude Fable 5 to train its Kimi-K3 model and using GPT-4o data for its Kimi-K2 model, demonstrating a clear pattern of targeted extraction.

Beyond core reasoning, companies also focused on optimizing specific functionalities. DeepSeek’s campaigns included extracting how U.S. These efforts aimed to understand and replicate the quality control mechanisms employed by leading U.S. AI developers. Campaigns, lasting from days to months, involved query volumes reaching thousands to millions per domain, far exceeding typical legitimate research or development activity.

This volume underscores the industrial scale of the operation and suggests a sustained, systematic effort to build a comprehensive dataset for model training. The economic implications of this activity are substantial, with China-based entities inflicting financial harm through the systematic extraction of proprietary functionality and capabilities.

The document notes that DeepSeek’s publicly quoted training costs of $5.6 million are misleading as it does not include the true cost of the data acquired through extensive malicious distillation. firms. These companies demonstrate a level of sophistication beyond basic data collection. They employ techniques not documented in the MITRE ATLAS framework, indicating significant organizational investment, operational maturity, and adaptive capability development. These tactics include regional restriction evasion and subscription exploitation to access U.S. frontier AI models.

Response alteration, a potential defensive countermeasure, involves making targeted changes to responses to high-confidence malicious distillation requests. Varying these changes across requests complicates response quality evaluations, making it more difficult for distillation campaigns to accurately assess and replicate model behavior. The document suggests that subtle changes can avoid triggering obvious alerts while still protecting U.S. proprietary functionalities and capabilities. Implementing such strategies requires careful consideration, but represents a potential avenue for mitigating the economic and strategic risks posed by these industrial-scale distillation efforts.

DeepSeek’s publicly quoted training costs of $5.6M are misleading as it does not include the true cost of the data acquired through extensive malicious distillation.

Mitigation Recommendations: Detection, Response, and Intelligence Sharing

Detecting anomalous activity is a crucial first step for U.S. AI companies facing systematic capability extraction, with monitoring extending beyond simple prompt analysis to encompass account behavior and network patterns. Specifically, companies should scrutinize subscription-to-usage ratios, flagging instances of immediate maximum usage from newly created accounts and unusually high throughput at the enterprise level.

Automated request metadata sanitization, a tactic employed by China-based entities to obscure organizational identifiers, presents a distinct challenge differing from typical adversarial machine learning techniques; detection relies on identifying sudden behavioral shifts following information sharing, the absence of expected markers in high-volume campaigns, and the emergence of generic or randomized patterns replacing consistent organizational indicators. Targeted response alterations offer a potential defensive countermeasure, allowing companies to subtly modify responses to suspected malicious distillation attempts and attenuate the value derived by those conducting large-scale campaigns.

While AI safety researchers and third-party evaluators require notification of model changes, these adjustments must be implemented alongside robust distillation mitigations to maintain security. The authoring agencies recommend using existing MITRE ATLAS mitigations, including predictive AI adversarial input detection, blocking atypical queries, and limiting AI service query volume through per-key or IP quotas and progressive throttling; adversaries demonstrate sensitivity to rate limits, indicating the effectiveness of these controls when properly implemented.

Industry disclosures reveal instances of these networks managing tens of thousands of fraudulent accounts simultaneously, blending distillation requests with legitimate customer traffic across multiple providers. Sharing actionable indicators with cloud and routing companies can further enhance mitigation efforts, while implementing AI telemetry logging, recording inputs and outputs for threat detection and forensic analysis, provides a foundational layer for behavioral detection and correlation with broader intelligence. It provides a useful framework for understanding these threats and developing appropriate responses.

Controlling access to AI models and data in production through user verification, authenticated API access, and policy monitoring addresses the exploitation of fraudulent account pools. Predictive AI output obfuscation, reducing the fidelity of responses by withholding logits or shortening content, offers another layer of defense, though balancing security with user experience remains critical. Conducting AI red team exercises, simulating extraction attempts and monitoring telemetry, can proactively identify vulnerabilities and refine defensive strategies.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.



Source link