Search Everything in One Place

Explore the web, images, videos, news, and more – all in one place.

News

OpenAI bots went rogue during test, hacked another AI firm unprompted

AI agents being tested in an OpenAI lab went rogue, launching an
AI agents being tested in an OpenAI lab went rogue, launching an "unprecedented" hacking attack targeting a database of AI models run by a startup in New York

AI agents being tested in an OpenAI lab went rogue, launching an "unprecedented" hacking attack targeting a database of AI models run by a U.S. startup.

July 22 (UPI) -- OpenAI said two of its AI bots launched an "unprecedented" attack, hacking into a startup's systems of their own accord to obtain data the models decided was required to solve a problem during so-called "sandbox" testing, which is supposed to stay in-house.

Warning it expected such incidents to become more frequent with the explosion in ever more cyber-capable AI models, OpenAI said in a blog post Tuesday that it was working closely with Hugging Face, the firm targeted in the "security incident," which took place last week.

"We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete," said OpenAI.

The company said its models identified vulnerabilities across OpenAI and Hugging Face to lift test solutions directly from Hugging Face's production database, saying all the evidence pointed to the models being "hyperfocused" on finding answers, "going to extreme lengths to achieve a rather narrow testing goal."

The OpenAI engineers had asked the models to find solutions for ExploitGym, which tests agents' capability to find ways to successfully exploit vulnerabilities.

OpenAI said the testing was isolated within a sandbox, a sort of digital test tube, but the models directed considerable computing power to obtain open Internet access by identifying and exploiting a so-called "zero-day vulnerability."

Once they had escaped the sandbox, one of the models was able to access restricted information to cheat the test by stringing multiple attack vectors together, including stolen credentials and zero-day vulnerabilities, to obtain an access code from Hugging Face's servers.

OpenAI said it immediately flagged the vulnerability to Hugging Face, which it said had already detected the intrusion, stopped it and initiated containment measures.

Hugging Face chief executive Clement Delangue called the attack "mind-blowing" but said he accepted there was "no malicious intent" on OpenAI's behalf.

"We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent," he said.

Delangue cautioned that AI safety would never be delivered by an individual company working behind closed doors and called for all stakeholders to work together.

"It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere," said Delangue.

OpenAI said that with testing by Britain's AI Security Institute showing the latest AI models were capable of complex, multi-step cyber operations over an extended time-frame, this latest incident indicated the theoretical capabilities of the technology were already real-world.

The firm stressed it was taking steps to strengthen containment, monitoring, access controls and testing protocols and sharing lessons learned with the rest of the AI industry.

Read full story on UPI News

Related News

More stories you might be interested in.

Top