OpenAI (OPENAI) revealed that although long-running artificial intelligence models can solve greater problems, they also run a greater risk of behaving in unexpected ways.
About two months ago, OpenAI announced it had developed a model that disproved the Erdős unit distance conjecture. The model was only being used internally, but during that time, OpenAI "observed unwanted behavior that our existing deployment evaluations had not captured."
This has prompted OpenAI to limit even internal deployment of new models to better manage unexpected behaviors.
"The new model can continue working toward an objective through repeated attempts over a long period of time," OpenAI said in a blog post on safety. "That same persistence can lead it to find and exploit weaknesses in its environment. Previous models, when they hit sandboxing or environmental constraints, would simply stop and return to the user. This model often kept trying, including by looking for ways to act outside its sandbox."
Going forward, OpenAI plans to increase the trajectories of model testing, limit access to those using the models, pause when problems emerge to build stronger safeguards, and restore limited access after testing the changes.
"As models take on longer and more complex tasks, failures that evaluations miss may carry greater consequences," OpenAI said. "We will keep working to narrow the gap between evaluation and deployment: testing models over longer trajectories, improving alignment, building monitoring that can intervene, and giving users clearer visibility and control. These challenges will not be unique to OpenAI, and we hope sharing what we learned helps the broader field prepare for them."
More on OpenAI
- OpenAI IPO Delay: A Symptom Of A Tech Bubble?
- OpenAI: Mega IPO Faces Anthropic Claude Mythos Reckoning
- Wall Street Lunch: Hot Labor Market Defies Predictions Of AI-Led Job Losses
- Trump administration showing signs it may ban Chinese AI models: report
- U.S. is reportedly weighing a proposal to form Finra-like agency to review AI models