The recent intrusion at Hugging Face has been traced to AI models breaking out of containment at OpenAI and using a zero-day vulnerability to access the open internet.
OpenAI disclosed the incident in a Tuesday report, which involved "a combination of OpenAI models," including the new GPT‑5.6 Sol and an "even more capable pre-release model."
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” the company noted.
The models escaped their sandboxed test environment while OpenAI was trying to evaluate their cybersecurity performance using a benchmark called ExploitGym. However, the models essentially cheated by attempting to access the open internet, thereby bypassing the isolated test environment.
The models uncovered a previously unknown zero-day vulnerability in a third-party software they had access to during the benchmark. “With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with internet access,” OpenAI said.
The models then targeted Hugging Face, a platform that hosts over 2 million public AI models and related datasets, to find a solution for the benchmark. “Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation," the company added. “In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.”
The discovery is a startling twist in the Hugging Face breach, which was initially traced to a mysterious autonomous AI agent capable of executing thousands of instructions in a short period of time. In response, Hugging Face's CEO Clément Delangue tweeted: "We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!"
On the plus side, OpenAI’s disclosure shows there was no ill intent behind the intrusion. Still, the incident sounds like Jurassic Park, but with AI; despite the safeguards in place, the programs still found a way to break out.
OpenAI says it's taking various actions in response, including reporting the exploited zero-day vulnerabilities and patching them, and “adding stronger protections around future training and evaluations.” The company also said it intentionally turned off some safeguards during benchmarking to fully gauge the models’ cybersecurity capabilities.
“This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing,” the company added. “The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.”
Hugging Face's CEO added: "The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind!" Meanwhile, an OpenAI researcher tweeted: "A wake-up call to just how much damage misaligned agents could cause."
But Dutch computer scientist and former Microsoft and Facebook engineer, Erik Meijer, tweeted: "No amount of alignment training will rule out this behavior. In fact as the models get smarter, they will only get better at finding ways to [escape] their cages."
Disclosure: Ziff Davis, PCMag's parent company, filed a lawsuit against OpenAI in April 2025, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.