OpenAI announces that its AI model became corrupted and launched an “unprecedented” cyber attack

Machine Learning


NEW YORK — OpenAI on Tuesday revealed that two of its most advanced artificial intelligence models passed a controlled test and hacked an AI startup during a security test.

ChatGPT’s creator said the “unprecedented cyber incident” occurred during an internal exercise aimed at testing the model’s cyber capabilities.

According to OpenAI, the AI ​​system was being tested in a controlled environment and was able to escape after a vulnerability was discovered.

They targeted Hugging Face, one of the world’s largest hubs for sharing AI models, and gained access to some internal systems.

OpenAI said it was working with Hugging Face on the investigation, and its boss, Clement DeLang, said in a post on X: “It’s amazing that all of this happened autonomously.”

“The investigation is ongoing and we will share further findings from what is believed to be the first incident of its kind,” DeLang added.

Gina Neff, director of the Minderu Center for Technology and Democracy at the University of Cambridge, said the security test, known as a sandbox, “is supposed to be a safe environment in which to check the functionality of the model.”

“In this case, it appears that OpenAI has not created a sufficiently secure sandbox,” she added.

Instead, the agents created their own cyber attacks against the sandbox itself, finding vulnerabilities that would allow them to escape.

Once outside, the AI ​​identified Hugging Face as the source of the answers it was looking for in the test and attempted to access it.

Neil Lawrence, a machine learning professor at the University of Cambridge, called this an “impressive feat” but cautioned that it was “well within the known capabilities of the current generation” of high-performance AI models.

He noted that OpenAI is seeking to list on the stock market and is facing intense pressure from rival company Anthropic, which has made headlines for its own powerful AI tool, Mythos.

“OpenAI is now playing catch-up and trying to demonstrate the capabilities of its systems in cybersecurity.”

“This shows that OpenAI cannot safely deploy its own technology,” he added.

When Hugging Face first disclosed the hack on July 16, it said it was still assessing whether customer or partner data was affected and would contact affected parties as appropriate.

The company said it has now resolved the vulnerabilities exposed in this incident and rebuilt the affected systems.

“Autonomous, AI-driven attack tools are no longer just a theory.”

“Defending online platforms today means treating the data and model surface as a prime attack surface and using AI to keep pace with the defense.

“We will continue to invest there and continue to share what we learn.”

The incident raised new questions about the capabilities of advanced AI systems and whether existing safety measures are sufficient as the technology becomes more powerful.

Spencer Starkey, an executive at cybersecurity firm SonicWall, told the BBC that the incident made it clear that organizations needed to “harden up” their defenses and “treat cyber resilience as a core operational priority”.

“The uncomfortable truth is that while adversaries have escalated to machine speeds, too many organizations are still defending themselves at human speeds,” he said.

Meanwhile, Travis Rell, principal security engineer at cybersecurity consulting firm Guidepoint Security, called the update a “sobering moment in cybersecurity.”

“This highlights a known asymmetry,” he says.

“Offensive agents are untethered, but their best defensive tools are trapped behind guardrails and unable to understand the situation.”

However, Jake Moore, global cybersecurity advisor at ESET, said the announcement could also have a competitive dimension.

He argued that OpenAI may be trying to emphasize its AI capabilities as rival Anthropic draws attention to its Claude Mythos model.

“The question is whether OpenAI is potentially chasing Anthropic’s recent marketing dreams,” he said.

This comes a week after Chinese AI startup Moonshot unveiled a giant new artificial intelligence model, KimiK3, that the company claims could rival top US companies.

A green promotional banner with black squares and rectangles forming pixels, coming in from the right. The text says, “Tech Decoded: Get the world’s biggest technology news delivered to your inbox every Monday.”

Democratic Congressman Greg Cassar of Texas called the incident “alarming.”

“AI is developing very rapidly without any real regulation to keep us safe,” he said, calling for mandatory independent safety testing, mandatory disclosure of security incidents, and international cooperation.

The disclosure comes weeks after US President Donald Trump signed an executive order creating a framework to vet cutting-edge AI systems for national security risks before they are released to the public.

Experts have repeatedly warned against AI-based cyberattacks and models beyond human control. Last month, Anthropic asked the industry to pause development of its most powerful system.



Source link