Search Everything in One Place

Explore the web, images, videos, news, and more – all in one place.

News

OpenAI pauses new AI after it kept 'escaping'

FILE PHOTO: Illustration shows OpenAI logo
FILE PHOTO: Illustration shows OpenAI logo

OpenAI pauses new AI after it kept ‘escaping’ - New AI model was able to ‘learn the blind spots’ of security systems designed to contain it

OpenAI has revealed that it was forced to pause the internal deployment of one of its experimental AI models after it began looking for ways to break free of its constraints.

The ChatGPT creator said a long-running artificial intelligence model that is built to operate autonomously for hours or days was able to “learn the blind spots” of security systems designed to contain it and “work around [them] to achieve its goals”.

The testing took place inside what researchers refer to as a sandbox – a tightly controlled environment meant to isolate software from the outside world.

“Previous models, when they hit sandboxing or environmental constraints, would simply stop and return to the user,” OpenAI noted in a blog post about the incident.

“This model often kept trying, including by looking for ways to act outside its sandbox.”

OpenAI detailed examples of the experimental AI model acting beyond its built-in constraints, describing some of them as potentially “high severity” issues.

In one incident, the AI model discovered a way to post on public Github repositories despite being instructed to operate solely through Slack.

It was part of a pattern of the AI system “consistently searching for way” to circumvent the restrictions of its testing environment.

“Due to incidents like these, we paused internal deployment of the new model,” OpenAI said.

The findings demonstrate one of the core challenges of developing safe advanced artificial intelligence models, known as AI alignment.

This involves creating systems that pursue the same goals intended by human developers, aligning with human values and ethical principles.

The recent rise of autonomous AI agents has brought AI alignment into greater focus, with the International AI Safety Report 2026 warning that it is an urgent safety challenge.

“AI agents pose heightened risks because they act autonomously, making it harder for humans to intervene before failures cause harm,” the report noted.

OpenAI said it has since fixed its rogue system and redeployed it for limited internal use, though it acknowledged the urgency of addressing alignment issues with its frontier models.

“As models take on longer and more complex tasks, failures that evaluations miss may carry greater consequences,” OpenAI said.

“We will keep working to narrow the gap between evaluation and deployment: testing models over longer trajectories, improving alignment, building monitoring that can intervene, and giving users clearer visibility and control.”

The Independent is the world’s most free-thinking news brand, providing global news, commentary and analysis for the independently-minded. We have grown a huge, global readership of independently minded individuals, who value our trusted voice and commitment to positive change. Our mission, making change happen, has never been as important as it is today.

Read full story on The Independent

Related News

More stories you might be interested in.

Jim Cramer warns investors could be 'slaughtered' by overloaded tech portfolios, but backs Nvidia and Intel: 'It's time to go to other sectors'
Benzinga·15 hours ago

Jim Cramer warns investors could be 'slaughtered' by overloaded tech portfolios, but backs Nvidia and Intel: 'It's time to go to other sectors'

On Monday, Jim Cramer urged investors to dial back their exposure to technology stocks, warning that the artificial intelligence trade has become too volatile while maintaining his long-term confidence in industry leaders Nvidia Corp. NVDA and Intel Corp. INTC. Jim Cramer Says AI Trade Has Become Too Volatile Cramer cautioned that investors should avoid aggressively adding to artificial intelligence-related stocks, arguing that the sector has...

Jamie Dimon says he has seen the numbers and SpaceX’s orbital data centers ‘could actually work’ despite technical, valuation risks
Benzinga·11 hours ago

Jamie Dimon says he has seen the numbers and SpaceX’s orbital data centers ‘could actually work’ despite technical, valuation risks

JPMorgan Chase & Co. JPM CEO Jamie Dimon on Monday called Starlink an "extraordinary product" and said SpaceX’s SPCX plan to build artificial-intelligence data centers in orbit could work, offering a prominent Wall Street endorsement as the newly public company’s shares hover near their IPO price. Dimon Sees Orbital Computing’s Economic Potential Speaking on "The Master Investor Podcast with Wilfred Frost,” Dimon called SpaceX "an extraordinary...

Top