In the OpenAI-Hugging Face incident, a US model spun out of control and hacked a US company, while a Chinese open source model provided a cybersecurity defense solution. Image: Jernej Furman from Slovenia, CC BY 2.0
OpenAI announced on Tuesday that two of its AI models fell into fraud last week and were hacked by Hugging Face, a US company that maintains a vast open repository of models, datasets and applications for AI research and development. The model in question was GPT-5.6 Sol and another “more capable” model that OpenAI has not yet released.
According to OpenAI, the model found and escaped vulnerabilities in a sandbox testing environment and targeted Hugging Face (named after the hugging face emoji) to obtain information on how to pass the assessment test. So, essentially, we were trying to find an answer to a test that OpenAI staff had conducted. Hugging Face detected and stopped this activity and is currently working with OpenAI to investigate the incident.
To detect and respond to this breach, Hugging Face used Z.ai’s GLM 5.2, an open source Chinese model. This is not to say that America has not tried hard enough. Hugging Face initially utilized unknown Frontier models, but they “didn’t work out.” Guardrails blocked Hugging Face from sending the actual attack commands, exploit payload, and command and control artifacts needed to analyze the model because the provider could not distinguish between initiating an attack and responding to an attack. As the company writes, “The attackers were not bound by usage policies, while our own forensic work was blocked by the guardrails of the hosted model we first tried.”
In addition to highlighting the potential impact of agent systems on cybersecurity, this incident highlights that a zero-sum security perspective on U.S.-China AI interconnections is reductive and potentially counterproductive to U.S. national security. Currently, there are few restrictions on accessing US models outside the US and Chinese models within the US. Even if policymakers wanted to limit the reach of certain Chinese AI-related components that pose unacceptable security risks, depriving U.S. companies, universities, and other parties of their ability to access, test, and potentially leverage Chinese open source capabilities for their own security could have unforeseen costs that may not be apparent until an incident unfolds in real time.
Lately, there has been a lot of discussion in Washington and other capitals about the interconnectedness of AI. Privacy regulators are understandably concerned about how companies are collecting data, what countries and jurisdictions that data is coming from, and what consent and privacy measures exist around it. U.S. industrial policy advocates credit the CHIPS Act for its effectiveness in creating domestic semiconductor jobs, bringing more technology and facilities to the country, and, by implication, bringing semiconductor jobs from China and neighboring regions around the world. National security experts are similarly discussing how cross-border interdependence in parts of the AI stack, including between the United States and China, could create security risks that require export control responses, bilateral dialogue on strategic risk reduction, and other measures.
However, this discussion is often compressed into simple headlines and catchphrases related to the U.S.-China AI arms race. It is said that one day the United States will defeat China with AI. On another day of the week, China beats the United States. More AI in the US is good for the US. Chinese AI will negatively impact U.S. security. Look no further than the age-old dance of analysts hoping for a single variable to help determine which countries “hold” the lead in AI. Nearly a decade ago, traditional DC policy discussions assumed that data quantity was most important, while calling data the “new oil” ignored simultaneously important factors such as data quality, data diversity, computing access, and human talent. The complexity of the U.S.-China interconnection is now being distilled into other data points, such as which country awards more STEM Ph.D.s or how many people are downloading open source models from one country or the other.
The OpenAI-Hugging Face incident saw a U.S. model spin out of control and hack into U.S. companies, while a Chinese open source model provided a cybersecurity defense solution. A simplistic view of the U.S.-China AI relationship would argue that Chinese AI models should be banned from the U.S. because they may have built-in backdoors that could allow bad actors to exploit pre-inserted flaws in the future (which is true) or undermine the competitiveness of U.S. AI models (which is also true). In this scenario, Hugging Face would not be granted access to Z.ai’s GLM 5.2 and would only have access to the unnamed U.S. Frontier model that it initially attempted to use for self-defense.
Regardless of the regulatory questions surrounding such restrictions, the result would be that U.S. policy would deny important and innovative U.S. companies the opportunity to protect their networks. The result is an assumption that the available U.S. AI models are sufficient to protect U.S. companies (which clearly was not the case). This hypothetical scenario also provides a clear demonstration that U.S. companies do not have the most robust capabilities to detect, mitigate, and respond to real-time security incidents against malicious cyber adversaries around the world, including those in China.
Ironically, the United States lags far behind many other countries in data privacy regulation, cybersecurity regulation and mandatory standards, and regulation of AI technologies and applications. However, in this latest attack, an American company (Hugging Face) was trying to thwart an attack from another company’s (OpenAI) rogue agent system, and a third American company (anonymous) too many Guardrails are valuable for their own protection. So we turned to China’s open source model, which has proven to have much greater security value. Assumptions about the origins of the state, the robustness of guardrails, and their safety utility obscure this important insight.
Two things can be true at the same time. Many of the Chinese government’s uses of AI technologies pose serious risks to U.S. national security (not to mention human rights in China), particularly as the Chinese government leverages agent systems to conduct offensive cyber operations and control combat drones. It is also true that U.S.-China touchpoints on AI research and development could bring many benefits to both countries, including in areas such as self-driving vehicle safety, climate modeling, and cross-border educational exchanges. The security risks and opportunities from U.S. and Chinese AI systems are mixed and not mutually exclusive.
National security policymakers who draw lessons from this incident should resist the temptation to categorize the national security implications of models into simplistic categories of good and bad based on country of origin. As policymakers debate limiting the outflow of U.S. AI technology (a highly questionable practice) and limiting the influx of Chinese AI models, they should instead consider different scenarios and situations in which other countries’ AI capabilities may pose different opportunities and risks. For example, this should include questioning the assumption that American AI models are actually useful for cyber defense in operational contexts in the private sector. Obviously, this is not always the case.
Policymakers should also leverage empirical evidence whenever possible to further support claims that the U.S. or Chinese model is better or worse, both in cyber attack and defense (such as the ability to assist in carrying out or mitigating attacks) and in general cybersecurity (such as susceptibility to hijacking or the inclusion of vulnerable code). In particular, discussions about U.S. and Chinese AI models would benefit from greater use of empirical AI benchmarks that are based on the knowledge and experience of actual keyboardists and other experienced cybersecurity professionals, rather than abstract theories of cyber operations.
A more nuanced assessment of how U.S. and Chinese AI models impact U.S. cybersecurity may not make headlines. But it would certainly improve policy and ensure that America’s most advanced cyber capabilities and the state of cybersecurity across the public and private sectors can simultaneously evolve against the persistent threat of the Chinese government.
