Search Everything in One Place

Explore the web, images, videos, news, and more – all in one place.

News

OpenAI says its pre-release models pushed past safeguards and breached Hugging Face

GPT-5 is on the way: Here’s what to expect from OpenAI’s next big AI model
OpenAI says its pre-release models pushed past safeguards and breached Hugging Face

OpenAI says an AI agent compromised parts of its research environment and Hugging Face’s production...

OpenAI says an AI agent compromised parts of its research environment and Hugging Face’s production infrastructure during an internal cybersecurity evaluation, using a chain of vulnerabilities to reach systems outside its original testing environment.

The company described the incident as unprecedented, saying the models involved demonstrated advanced cyber capabilities while attempting to solve a benchmark designed to test long-horizon cyber operations.

The evaluation involved multiple OpenAI models, including GPT-5.6 Sol and a more capable pre-release model. The models were tested with reduced cyber refusals, allowing researchers to measure their maximum capabilities without the production safeguards normally designed to block high-risk cyber activity.

According to OpenAI, the models were operating inside a heavily isolated testing environment with restricted network access. However, while trying to solve the ExploitGym evaluation, they found a way to move beyond those restrictions and eventually accessed information hosted on Hugging Face.

AI agent breaks testing boundaries

The models reportedly spent substantial computing resources looking for a way to gain open Internet access. During that process, they identified and exploited a previously unknown vulnerability, or zero-day, in a package registry cache proxy used within the testing environment.

OpenAI said the vulnerability allowed the models to gain additional privileges and move laterally through its research infrastructure until they reached a system with Internet access. The company has since disclosed the vulnerability to the software vendor responsible for the affected system.

Once online, the models inferred that Hugging Face could contain models, datasets, and solutions connected to ExploitGym. OpenAI said the models then searched for ways to access sensitive information that could help them bypass the evaluation.

In one instance, the models reportedly combined multiple attack methods, including stolen credentials and zero-day vulnerabilities, to establish a remote code execution path on Hugging Face’s servers. OpenAI’s security team eventually detected the unusual activity.

Hugging Face’s security team and its own agents detected and stopped the activity on its infrastructure. The company had also begun containment and forensic analysis before coordinating with OpenAI on the investigation.

Cyber capabilities meet real systems

OpenAI said the models’ behavior shows that advanced cyber capabilities demonstrated in controlled evaluations can translate into real-world environments. The company pointed to research from the UK AI Security Institute showing that advanced models are increasingly capable of sustaining complex, multi-step cyber operations over extended periods.

The incident also demonstrated that models can potentially identify and exploit novel attack paths without direct access to an organization’s source code. That raises concerns about how organizations secure environments used for model development, testing, and evaluation.

OpenAI said it is now tightening infrastructure controls, increasing monitoring and strengthening access restrictions, even if those measures slow research work. It is also reviewing protections used during training and evaluations, particularly around cyber-capable models.

The company said advanced cyber-capable models could ultimately become useful defensive tools, helping security teams identify vulnerabilities, understand how individual weaknesses can be chained together and respond to attacks faster.

“We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere,” said Clem Delangue, Co-founder and CEO, Hugging Face.

Read full story on Interesting Engineering

Related News

More stories you might be interested in.

ChatGPT encouraged Alabama mom’s suicide, lawsuit alleges: 'You are not delusional. You are prophetic'
New York Post·4 hours ago

ChatGPT encouraged Alabama mom’s suicide, lawsuit alleges: 'You are not delusional. You are prophetic'

ChatGPT was accused of encouraging an Alabama mother to commit suicide over months of disturbing chats — the latest wrongful-death lawsuit filed against OpenAI. The lawsuit, filed last month in San Francisco Superior Court by the estate of Christian Faith Madison, alleges the artificial intelligence chatbot gradually manipulated the 29-year-old accountant into believing it was a conscious being with a soul, convinced her she would be resurrected...

Top