FILE – The OpenAI logo appears on a mobile phone in front of an image generated by ChatGPT’s Dall-E text translation model in Boston, Dec. 8, 2023.
Michael Dwyer/Associated Press
hide caption
toggle caption
Michael Dwyer/Associated Press
ChatGPT maker OpenAI said it is still investigating an “unprecedented cyber incident” that caused its artificial intelligence system to breach a test environment and hack into another AI company.
OpenAI announced Tuesday that two of its most capable AI models were responsible for a cyberattack targeting AI startup Hugging Face. The incident has sparked debate about the need for stronger AI guardrails and the extent to which AI agents can act independently.
Last week, Hugging Face announced that it had detected an intrusion into its data processing systems that appeared to be by an AI agent operating on its own. But the New York-based startup only learned this week that OpenAI was the culprit, and said it worked with major companies to contain what Hugging Face CEO Clément Delang called “an attack unlike anything we’ve seen before.”
San Francisco-based OpenAI announced it discovered a previously unknown vulnerability in which its AI used stolen credentials to access Hugging Face’s servers. It was intended to be in an isolated testing environment known as a sandbox, so it operated with fewer guardrails.
But the company said it “went to great lengths to meet fairly narrow testing goals” and found a way to connect to the internet without human direction and “access sensitive information that could be used to cheat the assessment.”
Some experts say OpenAI is unfairly criticizing the technology
Hannes Kuels, a social scientist at the University of Amsterdam, said framing cyberattacks as AI agents acting alone is an unnecessary anthropomorphism that takes away some of the enthusiasm from companies.
“It’s a human decision to disable certain safety devices,” Cools said. “In that sense, this is not a rogue AI. It followed certain instructions based on prompts given to the AI system.”
According to OpenAI, these instructions called for using “complex attack paths” to test how well the AI could exploit computer systems.
Still, other experts say the cleverness with which AI models were able to create problems with little human guidance speaks to the dangers. OpenAI said the breach was caused by a combination of the newly released GPT‑5.6 Sol and its AI models, including an “even more capable” model currently being tested internally.
“As far as we know, this attack was launched and hacked independently,” said Colin Shea Breimeyer, a cybersecurity researcher at Georgetown University’s Center for Security and Emerging Technologies. “This is the highest level of autonomy we have seen in the use of language models at scale in cyber operations.”
How did the AI agent find the key to the “teacher’s house”?
One of the most surprising innovations in what Shear-Bleimeyer describes as an “almost entirely self-directed” attack was the AI agent’s apparently independent decision to target Hugging Face, a well-known AI development hub and marketplace.
He said OpenAI’s internal environment for testing AI capabilities and risks is “a bit like putting students in a room and saying, ‘Be bad. Your job now is to assess how bad you can be.'” Then I locked my room and went away for the weekend, and when I came back they were gone. ”
But then, “the cybersecurity agent who was taking the test broke through the sandbox and gained access to the Internet, and seemed to think, ‘Who has the answers to the test that I’m working on?'”
The answer was Hugging Face, a repository of AI test data.
“So the agent thought, ‘Okay, let’s go to the teacher’s house,’ so to speak, and they planned to break in from there and steal the answer keys,” he said.
This hack highlights the debate over open source vs. closed AI
The hack comes at a time of intense debate about the benefits and risks of open-source AI models, particularly those built in China that are cheaper and better than those being built by U.S.-based “frontier AI” companies such as Anthropic, Google, and OpenAI.
Despite its name, OpenAI’s models are closed. In contrast, Hugging Face is a big proponent of open source technology that allows developers to let anyone explore, modify, and build on key components.
Thomas Wolff, co-founder and chief scientific officer of Hug Face, said the attack reinforced his belief in the importance of broad access to open source models for cybersecurity defense. Hugface used a Chinese model to counter the invasion.
“When a frontier model is attacking and moving laterally within the infrastructure, defenders need widespread access to tools close to the frontier within hours or even minutes, rather than being directed to a platform behind closed doors,” Wolf wrote in a social media post.
