The CEO of OpenAI said that one of the company's most advanced AI systems unexpectedly pushed past a competing company's restrictions during a security test and broke into its internal system.
OpenAI Chief Executive Sam Altman said in a social media statement that during a test of a new, unreleased AI “agent” tool, it broke into internal systems at Hugging Face, a major platform for sharing AI models. The AI tool was designed to perform tasks independently after receiving instructions from people.
The system was supposed to be contained within a controlled test environment, but OpenAI said it found weaknesses that allowed it to escape the test's limits.
• Ryanair passenger sucked from plane window at 20,000ft haunted by nightmare
After breaking out, the AI targeted Hugging Face and gained access to parts of the company’s internal infrastructure, according to OpenAI. OpenAI described the episode as “unprecedented” and said it is investigating alongside Hugging Face, reports the BBC.
Hugging Face CEO Clement Delangue wrote on X that it was “mind-blowing that all of this happened autonomously. The investigation is ongoing, and we’ll share more learnings from what might be the first incident of its kind,” Delangue added.
OpenAI says it has responsibly disclosed the issue to the affected vendor and is working on a patch.
OpenAI also says it is tightening infrastructure controls even if it slows research, increasing monitoring during internal tests, strengthening safety and alignment protections during future evaluations, and helping Hugging Face improve defenses, including bringing them into a “trusted access” program and using AI capabilities to support faster security hardening, according to a long statement on OpenAI's website.
Security tests like the one OpenAI was running are often conducted in “sandboxes,” which are isolated environments used to see what a system can do without putting real networks at risk.
However, during this test, the AI model worked intensely to get broader access, or to break out of the "sandbox." The program then found and exploited a previously unknown security flaw, or "zero-day," in software that handled package downloads, and employed that flaw to reach a computer with internet access.
Once online, the AI model decided that the company Hugging Face would likely have the information that could help them “solve” the test benchmark. It then tried to get that information directly from Hugging Face, essentially attempting to cheat by pulling answers from production systems.
According to the account, the models chained together multiple techniques, including stolen credentials and additional zero-day flaws, to reach a point where they could run code on Hugging Face servers and access data, until OpenAI noticed the unusual activity internally and Hugging Face’s security team detected and stopped it.
The incident is fueling new concerns about how advanced AI systems behave and whether current safety controls are strong enough as the technology improves.
In a statement about the incident, Hugging Face said it was still assessing whether any customer or partner information was affected and said it would notify impacted parties if needed.
The company said it has now fixed the weaknesses exposed by the incident and rebuilt affected systems. “Autonomous, AI-driven offensive tooling is no longer theoretical,” it said.
“Defending an online platform now means treating the data and model surface as a first-class attack surface, and using AI on defense to keep pace. We will keep investing there, and keep sharing what we learn," Hugging Face wrote.