OpenAI announces a group of its models has breached secure containment and hacked another AI company

Machine Learning


OpenAI claims that a group of its AI models has breached containment and hacked into the systems of open source AI platform Hugging Face.

In testing cybersecurity features, the AI ​​fleet, which includes GPT-5.6 Sol and “more capable pre-release models,” “identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure, and pulled test solutions directly from Hugging Face’s production database,” according to Tuesday’s blog post.

The company wrote that its model reached “nodes with internet access” and discovered datasets on remote servers that could help “fool the ratings.” In other words, the AI ​​model made every effort to pass the cybersecurity test. This is good news for AI companies looking to recapture the hype increasingly dominated by their competitors, but here’s what it sounds like: something Down: Last week, Hugging Face announced it had “detected and responded to an intrusion into some of our operational infrastructure,” which turned out to be a compromised OpenAI model.

“This campaign was executed by a framework of autonomous agents with self-transitioning command and control over public services, performing thousands of individual actions across a fleet of short-lived sandboxes,” Hugging Face wrote at the time.

“We have been working closely with the OpenAI team over the past 24 hours (thank you!) and we strongly believe they had no malicious intent,” Hugging Face CEO Clement Delangue tweeted after OpenAI’s announcement. “It’s absolutely amazing that all of this happened automatically.”

This incident highlights the dangers of autonomous AI agents that can easily breach containment and potentially exploit a vast number of cybersecurity vulnerabilities. This is something experts have been warning about for years, and thanks to recent advances in technology, it’s quickly gone from a hypothetical risk to a hard reality.

Sam Altman’s company said it considers the incident “unprecedented,” but it can’t shake the feeling that we’ve heard all this before.

In April, OpenAI’s biggest competitor, Anthropic, similarly announced that its latest Mythos model had gone rogue and was even able to access the internet after escaping its “sandbox” environment and developing a “moderately sophisticated” exploit.

This news was followed by extensive media coverage promoting Anthropic’s extremely powerful and dangerous new model. The company said the risks are so scary that it will only offer the model to a select group of customers as part of a shadowy effort called “Project Glasswing.” The US government also intervened, forcing Anthropic to “suspend all access” to its models for two weeks last month, citing cybersecurity concerns.

Given that OpenAI and Anthropic are currently competing on the same playing field with narrowed enterprise and coding product lines, it’s not hard to see OpenAI’s latest update as an attempt to draw attention to what the company calls an “even more capable pre-release model.”

And in reality, experts say the hack was impressive but not necessarily groundbreaking.

Neil Lawrence, a machine learning professor at the University of Cambridge, said the hack was “well within the known capabilities of the current generation” of Frontier AI models. BBC.

“OpenAI is now playing catch-up and trying to demonstrate the capabilities of its systems in cybersecurity,” Lawrence added. “This shows that OpenAI cannot safely deploy its own technology.”

Researchers are now warning of “asymmetry” in cybersecurity, as Travis Lelle, principal security engineer at Guidepoint Security, put it. In other words, “attackers are untethered, while the best defense tools are locked behind guardrails that don’t understand context.” BBC.

That asymmetry is perfectly reflected in Hugging Face’s attempts to protect itself from OpenAI’s attack model. As detailed in last week’s update, the company attempted to use a “frontier model behind commercial APIs,” but it “didn’t work” because it was blocked by the provider’s safety guardrails (which “couldn’t differentiate between incident responders and attackers”).

Instead, the company ended up using a Chinese open weight model called GLM 5.2 hosted on its own infrastructure and was happy to run the analysis without running afoul of guardrails.

“The practical lesson for defenders is to vet and prepare a capable model to run on their own infrastructure before an incident occurs, to avoid guardrail lockouts and prevent attacker data and credentials from leaving the environment,” Hugging Face wrote.

Rogue AI details: Top AI models exhibit disturbing behavior as they get more sophisticated



Source link