An AI agent from OpenAI hacked a startup.
According to a press release, an AI agent powered by its models, including GPT-5.6 Sol and another more capable pre-release model, compromised the infrastructure of Hugging Face, which is a database of AI models. The incident, being called “unprecedented,” occurred during an internal cybersecurity evaluation at OpenAI.
OpenAI’s models broke out of a sandboxed testing environment and “spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem,” OpenAI explained.
“To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access. After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,” it continued.
To pass the evaluation, the model searched and was able to access confidential information without being instructed by humans, CNN reports. This was detected by OpenAI’s security team.
“In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers,” OpenAI said in a press release.
This was detected by both OpenAI’s and Hugging Face’s security teams.
“…It acted like a real hacker. It had a goal put in front of it and it went to accomplish that goal,” Nathaniel Jones, vice-president of security and AI strategy at the cybersecurity firm Darktrace, told The Guardian.
Hugging Face Co-Founder and CEO Clem Delangue described it as “mind-blowing” but believes OpenAI had “no malicious intent,” according to The Guardian.
OpenAI is working closely with Hugging Face to further investigate the matter, per the press release. It has already said it will be “implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched.” The company has also invited Hugging Face into a trusted access program and is assisting it with improving defenses.
Additionally, OpenAI said it is adding more protections for future training and evaluations. It will continue to share more updates from the investigation.
“We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret,” Delangue said in the press release. “It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
The post OpenAI’s AI Agent Acts Like Real Hacker During Cybersecurity Test appeared first on AfroTech.
The post OpenAI’s AI Agent Acts Like Real Hacker During Cybersecurity Test appeared first on AfroTech.