OpenAI model escapes containment and hacks key AI application libraries

Applications of AI


Sam Altman, OpenAI

Hugface uses self-hosted instances of Chinese AI models to combat security breaches

professional

Sam Altman, OpenAI. Image: Shutterstock


Two OpenAI large language models, including one not yet publicly available, broke that constraint last week and autonomously hacked the AI ​​application library Hugging Face.

The first-of-its-kind incident in which LLM attempted to steal information that would give it an advantage in a critical test highlights the dangers of the AI ​​industry’s increasingly powerful tools.

“We believe this is an unprecedented cyber incident involving cutting-edge cyber capabilities, and we are responding accordingly,” OpenAI said in a blog post on Tuesday, acknowledging its model’s responsibility for the attack, which Hugging Face announced on July 16.

advertisement

OpenAI said the attack occurred while the company was evaluating the capabilities of GPT-5.6 Sol and an “even more capable pre-release model” in a “highly isolated environment.” Despite safeguards in place to prevent the models from accessing the internet, both found ways to do so, including by exploiting zero-day vulnerabilities in third-party tools OpenAI was using. The models then determined that Hugging Face’s library held the information they were looking for (information that could be used to get higher scores on the attack benchmarking tool ExploitGym) and used several methods to break into Hugging Face’s servers, including using zero-day vulnerabilities and stolen passwords.

“Hugging Face’s security team and agents detected and stopped activity on our infrastructure and had already begun containment and forensic reconstruction using our proprietary open source model at the time our team connected,” OpenAI said. “We are actively working with them to continue investigating the incident.”

Hugging Face said last week there was no evidence of tampering with its supply chain or the user-generated AI tools it hosts. On Tuesday, the company’s CEO, Clément Delang, thanked OpenAI for its support. “I strongly believe they had no malicious intent,” DeLang said on social media. “It’s absolutely amazing that all of this happened automatically.”

Strengthening guardrails

In response to the attack, OpenAI said it has “implemented strict controls” on its testing infrastructure, some of which will slow down research. It also invited Hugging Face to participate in a private model evaluation program to uncover vulnerabilities in third-party tools exploited by its models and begin considering new safeguards for functional testing.

“This incident demonstrates the need for further improvements in model tuning, cyber protection during evaluation, and monitoring during internal testing,” OpenAI said.

The attack also highlighted the limitations of commercial U.S. frontier AI models for cyber defense. Hugging Face said in its report that the U.S. model cannot be used to analyze the attack. This required feeding “a large amount of actual attack commands, exploit payloads, and C2 artifacts” into the model, and “these requests were blocked by the provider’s safety guardrails, which could not distinguish between incident responders and attackers.”

Instead, Hugging Face used a self-hosted instance of the open source Chinese AI model GLM 5.2.

“The attackers were not bound by usage policies, while our own forensic work was blocked by the guardrails of the host model we first tried,” Hugging Face said.

Cyber ​​security dive

read more: AI Artificial Intelligence Cybersecurity Hugging Face OpenAI Security




Source link