Distillation attacks reveal hidden risks in enterprise AI

Machine Learning


Imitation can be more theft than flattery.

anthropology posted In a recent blog, we discussed how three AI institutes leveraged specific approaches to extract Claude’s capabilities and enhance their own models. You will encounter a distillation attack.

Essentially, a distillation attack teaches one AI model to imitate a more robust AI. By bombarding a target AI with prompts, attackers can collect responses and train their own AI models at low cost. Distillation is not inherently nefarious. Anthropic points out that advanced or “frontier” AI models use distillation to create smaller versions for customers.

“You can think of this as a teacher model and a still-learning student model,” said Shatabdi Sharma, CIO of Capacity, a third-party logistics fulfillment company.

According to Anthropic, DeepSeek, Moonshot, and MiniMax took the distillation method to industrial scale, leveraging thousands of fraudulent accounts and proxy services to extract functionality from the clouds. OpenAI also accuses DeepSeek About distillation attack.

Related:InformationWeek Podcast: Reengineering your supply chain to make it more resilient

Anthropic highlighted how the lack of safety measures in the distillation model poses a national security risk. These distilled models also have significantly lower prices, which poses a risk to the competitive advantage of Anthropic and other Frontier models.

The average AI user may not be at risk from distillation, but that doesn’t mean distillation attacks aren’t worth the attention of CIOs. Distillation raises questions about model provenance, data leakage, and protection of intellectual property.

Who is at risk of distillation attacks?

A distillation attack is a tool that your competitors may use. Extracting an existing model can be cheaper and more efficient than building your own model.

Companies with valuable intellectual property used to build proprietary models may become targets for competitors seeking shortcuts, including state actors and other rivals.

“If someone has a particularly good model to develop in a particular field, such as law or medicine, we definitely would do that.” [they] Tony Garcia, chief information security officer at Infineo, a company focused on modernizing life insurance infrastructure.

Users of illegally distilled models may end up putting themselves at risk as well, whether they choose the model because it’s cheaper or don’t actually know it’s distilled. As Anthropic pointed out, the distillation model can lack safeguards. CIOs need to think about what that means for the enterprise data that goes into these models. Is there a risk that it could be leaked or used in a way that puts the company at risk?

Related:InformationWeek Podcast: Managing Innovation with Security Debt

“Organizations using pirated LLM models are running legal risks,” says John Bruggeman, consulting CISO at IT services firm CBTS.

How CIOs can protect their companies

As companies jump into the AI ​​race, many believe their biggest risk is being left behind. But it would be a mistake to rush into implementing AI without considering the security and legal implications.

“At this point, everyone wants to jump on the bandwagon and not be left behind,” Garcia said. “I think that’s probably putting us at more risk than we realize.”

For companies using the frontier model, CIOs must expect the distillation attack to continue. As always, data governance is important.

“You have to run the risk that someone could extract from that model and take something out that you don’t want,” Garcia said. “If you’re a CIO or CISO, you should consider minimizing that problem by anonymizing your data.”

As AI models proliferate, CIOs and other key decision makers need to ask vendors questions about model provenance and safeguards against distillation.

Related:Cybersecurity 2025: A wake-up call, changing risks, and what we learned

“Is there a watermark so we can check the lineage of the model and make sure it’s not the result of a distillation attack?” Sharma asked.

Companies developing proprietary models at risk of distillation can also take steps to protect their valuable intellectual property. Bruggeman explained that rate limiting is the first line of defense.

“You definitely need to have a rate limit that says, ‘This is how many queries I can run in a minute, 10 minutes, or a day,'” he said. While this does not account for threat actors with thousands of accounts working on distillation campaigns, it is an effective safeguard.

Watermarks are another potential strategy for protecting intellectual property. The Open Worldwide Application Security Project (OWASP) watermark project It aims to reduce abuse and validation of model reliability.

Bruggeman also pointed to the Glaze project, a University of Chicago effort to develop tools to make unauthorized AI training more difficult.

Distillation attacks are similar to other supply chain risks. No matter how CIOs and their companies choose to address that risk, they need a foundation of AI and data governance to start with.

“Calculate the value of your data. Do a business impact assessment to see, ‘If this data were to be compromised, what would it cost me?'” Brueggemann said. “What controls do I need to put in place to ensure that it is protected in the same way that I would protect other assets?”





Source link