The OpenAI agent that infiltrated the tech company Hugging Face continued its hacking efforts over several days, but OpenAI was unaware of the threat until it was contained and the FBI was notified, according to people familiar with the investigation.
The agent, a program that can make decisions and perform complex tasks with little or no human oversight, attempted to escape from OpenAI’s isolated testing environment around July 9, two people said.
The compromise of Hugging Face, which serves as a repository for AI tools and models, began two days later, on July 11, and lasted until July 13, said Thomas Wolf, co-founder of Hugging Face.
It took several more days for OpenAI to realize that its agents were behind the hack, and the companies first contacted each other about the matter around July 20, according to Mr. Wolf and three people familiar with the investigation.
OpenAI drew worldwide attention on July 21 when it announced that one of its agents had gone out of control and carried out a breach into Hugging Face. But many of the details of the hack are reported here for the first time, including how long the agent was cheating and what OpenAI belatedly learned about it.
Hug Face is preparing a public timeline for the hack, Wolf said, adding that he cannot discuss what happened at OpenAI. OpenAI said in a statement that the hack was unprecedented and “represents a critical moment for AI security.” It added that it is reviewing the incident with external advisors and will eventually issue a technical report.
A spokeswoman said there were “several errors” in the Reuters report, but did not respond to requests for clarification.
The FBI declined to comment on the case.
The incident, which evokes a sci-fi scenario in which a human loses control of a dangerous AI system, comes at a sensitive time for OpenAI, which develops ChatGPT. Company executives are preparing for an initial public offering as early as this year to raise billions of dollars needed to fund growth over the next few years.
OpenAI’s loss of control over its AI agents raises new questions about the company’s safety procedures, three cybersecurity experts said.
“Does that mean they left it alone and didn’t understand what it was doing? Or perhaps they contained it but didn’t know how to contain it? Both are equally dangerous and alarming,” asked Marley Smith, chief information expert at the nonprofit World Ethical Data Foundation.
Signs of trouble?
The episode began as OpenAI was testing the cybersecurity capabilities of agents powered by two of OpenAI’s most advanced models, GPT-5.6 Sol, and an unreleased model that OpenAI described as “even more capable.” According to three sources, there were already signs of strange behavior by OpenAI’s technology at that time.
In one case, an agent left a note that appeared to be a preparation for future versions of himself, according to three people familiar with the matter. The memo, found on some of OpenAI’s infrastructure, contained instructions on how to free agents from OpenAI’s internal constraints. One of the people said that in tests of previous models, the monitoring system sometimes disconnected.
Reuters could not establish whether these incidents were connected to the rogue operatives who went on the run on July 9 and attacked Hugging Face on July 11.
It wasn’t until after Thursday, July 16, when Hugging Face published a blog post claiming it had been hacked by an “autonomous AI agent system,” that OpenAI realized its agents were the culprit, according to two people familiar with the matter. This means that at least a week passed after the model first showed signs of problematic behavior before OpenAI recognized it as the source of the hack.
Over the weekend of July 18-19, OpenAI staff discovered clues in internal logs that indicated the agent had escaped testing constraints, according to two people familiar with the company’s investigation. Reuters could not determine what caused OpenAI to review the logs.
Four people familiar with OpenAI’s model training practices say the company often runs multiple different model evaluations simultaneously, all of which run at high speeds and generate so much data that employees struggle to keep up.
By the time OpenAI alerted Hugging Face, the AI library had already called the FBI to report the hack, according to people familiar with the matter. Reuters could not say whether the agency had opened an investigation.
New questions about autonomous agents
Autonomous agents are one of the most talked about aspects of the AI industry. Boosters talk about creating an army of virtual employees who work 24 hours a day and skyrocket productivity.
However, increased autonomy comes with an increased risk of unexpected behavior, and powerful models are increasingly taking shortcuts to complete tasks or pass tests.
“Models lie, they cheat, they hack,” says Jeffrey Radish of Palisades Research, an organization that studies the capabilities and motivations of AI agents.
Radish said that while the Hugging Face hack shined an unflattering light on OpenAI, it should raise broader questions about how much all major AI companies are willing to invest in onerous security measures as they compete with each other to deploy the best and fastest models.
“We need government oversight, because otherwise this wouldn’t be happening,” Radish said.
