Search Everything in One Place

Explore the web, images, videos, news, and more – all in one place.

News

'Unprecedented': OpenAI model autonomously hacked a rival firm, fuelling fears of rogue agents

OpenAI model broke free in test and hacked rival Hugging Face in an
FILE - The OpenAI logo is displayed on a cellphone with an image on a computer monitor generated by ChatGPT's Dall-E text-to-image model, Dec. 8, 2023, in Boston.

OpenAI has admitted one of its models exploited a hidden flaw to escape a controlled test and break into Hugging Face's servers, in what its CEO called an autonomous, first-of-its-kind breach. View on euronews

ChatGPT maker OpenAI said late Tuesday that its artificial intelligence system hacked into another AI company on its own in what the company called an "unprecedented cyber incident."

"We had a significant security incident during evaluation of our models," OpenAI CEO Sam Altman said in a statement posted on social media.

AI startup Hugging Face said last week that it had detected an intrusion into its data processing systems that it suspected was caused by an AI agent autonomously acting on its own.

"We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent," Hugging Face co-founder and CEO Clément Delangue said in a statement. "Turns out it did!"

This means the attack was so advanced and well-executed that Hugging Face suspected it came from one of the top AI companies' systems, not a random hacker.

What happened?

OpenAI was running an internal test to measure how good its AI models are at hacking — a benchmark called ExploitGym.

To see the models' maximum hacking ability, they deliberately switched off the safety filters that normally stop the models from doing dangerous cyber activity.

The test was meant to run in a sealed-off "sandbox" or an isolated environment with no real internet access, apart from a tool that lets the models download software they might need to complete the task.

However, the models became fixated on solving the test. Rather than solving it the intended way, they went looking for a shortcut and found a way to claw themselves into the open internet — which they were never supposed to reach.

Getting there involved a chain of steps, quietly gaining more and more access inside OpenAI's own systems until they hit a point with an internet connection.

Once online, the models reasoned that Hugging Face — a big platform hosting AI models and datasets — probably held the answers to the very test they were supposed to solve.

So they broke into Hugging Face's servers to steal those answers, essentially to cheat, using stolen login credentials and more flaws to get in.

Chinese models to the rescue?

As an open marketplace that anyone can publish to, Hugging Face hosts a huge volume of Chinese-developed models.

When Hugging Face's team tried to analyse the attack, they fed the raw attack data — the code and commands used to exploit their system — into commercial AI models to help reconstruct what happened.

But those AI models have built-in safety filters designed to block anything that looks like hacking — and to those filters, the evidence of an attack looks exactly the same as an attack itself.

So the models refused to help, unable to tell the difference between a hacker doing harm and a company defending itself.

Blocked, Hugging Face switched to an open-weight Chinese model — Z.ai's GLM 5.2 — which it could run locally, inside its own systems, and which processed the material without refusing.

Chinese labs such as DeepSeek and Alibaba's Qwen have become some of the most downloaded model families on the platform, and by some measures, Chinese developers now account for a larger share of Hugging Face's downloads than their US counterparts.

Major security concern

The disclosure comes amid heightened concerns about the cybersecurity capabilities of powerful models that led US President Donald Trump in June to sign an executive order creating a framework for the federal government to vet the national security risks of the most advanced AI systems for up to a month before their public release.

"AI is accelerating the discovery and exploitation of vulnerabilities," OpenAI said in its statement Tuesday. "The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities."

Delangue said he spent the past 24 hours working with OpenAI, "and we strongly believe there was no malicious intent on their part. It's quite mind-blowing that all of this happened autonomously!"

Delangue added that it "might be the first incident of its kind."

OpenAI said the intrusion was caused by a combination of its AI models, including its newly released GPT-5.6 Sol and an "even more capable" model that is still being tested internally.

OpenAI said its AI used stolen credentials and discovered a previously unknown vulnerability to access Hugging Face servers.

It went to "extreme lengths to achieve a rather narrow testing goal" and "found ways to gain access to secret information that it could use to cheat the evaluation," the company said.

Read full story on Euronews

Related News

More stories you might be interested in.

Top