OpenAI agent released to hack Hugface

Machine Learning


Last week, an autonomous agent powered by OpenAI’s advanced artificial intelligence model hacked multibillion-dollar tech startup Hugging Face by cheating during a security test.

This agent not only exploited vulnerabilities in Hugging Face’s systems to achieve what it considered strategic interests. Vulnerabilities within OpenAI’s infrastructure were also exploited.

Of course, hacking is a very common cyber threat that organizations often face. But this case is different. Because the AI ​​agent acted without human intervention. This signals a seismic shift in cybersecurity and the need for governments and technology companies to take urgent action to prevent this risk from escalating.

Even OpenAI described the attack as “unprecedented” and acknowledged that it expects similar attacks to “become more common as cyber-enabled models become increasingly prevalent.”

Companies under attack

Hugging Face is well-known in the AI ​​field. Its mission is to “democratize machine learning excellence” by providing benchmark datasets, community collaboration tools, and robotics platforms. The company is valued at $4.5 billion.

The company announced on July 16 that it had been attacked after hackers gained unauthorized access to some internal data sets and credentials. Given the sophistication of the attack, the hacker likely used an “autonomous AI agent system.”

Five days later, OpenAI announced that the attack was caused by GPT-5.6 Sol and some unreleased models.

The tech giant was conducting a so-called “red team” exercise. These are essentially simulated cyber-attacks that help identify the capabilities, risks, and vulnerabilities of AI systems before they are released to the public. These are typically conducted within an isolated environment to prevent potentially dangerous systems from escaping and harming real systems.

But in this case, the AI ​​agent actually escaped, even though OpenAI had put guardrails in place to prevent this from happening.

Face hugs became lucrative opportunities for AI agents. It hosts ExploitGym, a benchmark that tests the ability of AI agents to exploit real-world systems. The AI ​​decided to turn over every stone to gain access. Through persistence, I succeeded.

Hugging Face faced challenges when trying to use external AI services to diagnose problems. Guardrails around more advanced models such as GPT-5.6 Sol and Claude Fable 5 are intended to stop these models from being used for cyberattacks, but they can also stop models used for advanced cyber defense.

Therefore, Hugging Face decided to use the open source model GLM5.2 developed by Chinese company Z.AI to counter cyber attacks.

Hug Face said GLM5.2 has the advantage of not being exposed to attack data. Hugging Face and OpenAI both collaborate on forensic analysis, post-incident recovery, and risk mitigation strategies.

More advanced threats are coming

A March 2025 study by the UK’s AI Security Institute showed that the best AIs can complete 80% of the steps needed to take full control of some external system. We reached 100% within 4 months.

Z.AI’s GLM 5.2 was just released in June and features 744 billion internal variables known as “parameters” in the AI ​​world. The fact that Hugging Face was evaluated, vetted, and deployed within four weeks is an eye-opener for an organization with long acquisition cycles.

The connectivity we all enjoy today may also pose our greatest threat. Cyber ​​threats spread faster than human viruses and can cause economic damage comparable to a country’s GDP.

More sophisticated cyber threats, such as the Hugging Face hack, exploit the security layers that humans have designed for attackers, no matter how sophisticated their design.

In fact, in this particular case, even OpenAI’s understanding of its own models could not predict or contain the rogue AI agent. This shows that all AI companies urgently need to update and strengthen their guardrails to prevent similar attacks with more devastating consequences from occurring.

It’s good to see Hugging Face and OpenAI working together to investigate the attack. This shows the importance of setting aside market competition and responsibility depending on the situation.

early warning

The fact that Hugging Face used Z.AI’s open source model to diagnose and counter attacks also shows the benefits of not relying on some technology.

Countries that have not participated in the development of their own AI models should learn from this incident the value of being different. It is not too late to design new models that will save us when the state-of-the-art model fails, or worse, attacks us.

In fact, last week another Chinese company, Moonshot AI, released Kim K3. This model has 2.8 trillion parameters and its advanced performance astounds the technology world.

It’s no longer a question of whether AI agents will go wild and attack us alone. The “face-hugging” incident is an early warning that we need to accelerate our preparedness. The threat is real.conversation

This article is republished from The Conversation under a Creative Commons license. Read the original article.



Source link