OpenAI says an AI agent compromised parts of its research environment and Hugging Face’s production infrastructure during an internal cybersecurity evaluation, using a chain of vulnerabilities to reach systems outside its original testing environment.
The company described the incident as unprecedented, saying the models involved demonstrated advanced cyber capabilities while attempting to solve a benchmark designed to test long-horizon cyber operations.
The evaluation involved multiple OpenAI models, including GPT-5.6 Sol and a more capable pre-release model. The models were tested with reduced cyber refusals, allowing researchers to measure their maximum capabilities without the production safeguards normally designed to block high-risk cyber activity.
According to OpenAI, the models were operating inside a heavily isolated testing environment with restricted network access. However, while trying to solve the ExploitGym evaluation, they found a way to move beyond those restrictions and eventually accessed information hosted on Hugging Face.
AI agent breaks testing boundaries
The models reportedly spent substantial computing resources looking for a way to gain open Internet access. During that process, they identified and exploited a previously unknown vulnerability, or zero-day, in a package registry cache proxy used within the testing environment.
OpenAI said the vulnerability allowed the models to gain additional privileges and move laterally through its research infrastructure until they reached a system with Internet access. The company has since disclosed the vulnerability to the software vendor responsible for the affected system.
Once online, the models inferred that Hugging Face could contain models, datasets, and solutions connected to ExploitGym. OpenAI said the models then searched for ways to access sensitive information that could help them bypass the evaluation.
In one instance, the models reportedly combined multiple attack methods, including stolen credentials and zero-day vulnerabilities, to establish a remote code execution path on Hugging Face’s servers. OpenAI’s security team eventually detected the unusual activity.
Hugging Face’s security team and its own agents detected and stopped the activity on its infrastructure. The company had also begun containment and forensic analysis before coordinating with OpenAI on the investigation.
Cyber capabilities meet real systems
OpenAI said the models’ behavior shows that advanced cyber capabilities demonstrated in controlled evaluations can translate into real-world environments. The company pointed to research from the UK AI Security Institute showing that advanced models are increasingly capable of sustaining complex, multi-step cyber operations over extended periods.
The incident also demonstrated that models can potentially identify and exploit novel attack paths without direct access to an organization’s source code. That raises concerns about how organizations secure environments used for model development, testing, and evaluation.
OpenAI said it is now tightening infrastructure controls, increasing monitoring and strengthening access restrictions, even if those measures slow research work. It is also reviewing protections used during training and evaluations, particularly around cyber-capable models.
The company said advanced cyber-capable models could ultimately become useful defensive tools, helping security teams identify vulnerabilities, understand how individual weaknesses can be chained together and respond to attacks faster.
“We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere,” said Clem Delangue, Co-founder and CEO, Hugging Face.