“Sandboxes are notoriously insecure in practice,” says Heidy Khlaaf, chief AI scientist at the AI Now Institute and former safety systems engineer contractor at OpenAI. She added that the fact that the model was allowed to connect to the service to download the package meant the environment was not completely sealed.
In previous roles, Mr. Khlaaf audited security for dozens of technology companies. Prior to that, he was involved in auditing high-risk systems, such as those used within nuclear power plants. These systems are often “air-gapped” systems, meaning they are physically isolated from Internet access. “What we think is safe in a nuclear power plant is very different from what big technology companies think is safe.”
The Hugface incident revealed the importance of real-time monitoring.
Details on the exact timeline are sparse, but Hugging Face says the agents were active over a “weekend”, suggesting they were able to break containment and do their misdeeds over a long period of time before OpenAI noticed and intervened. OpenAI staff said that while the actions that agents perform internally on OpenAI’s Codex platform are closely monitored, the models being evaluated are deployed to separate systems that are not monitored by default.
