OpenAI (OPENAI) Tuesday said that its advanced artificial intelligence models inadvertently hacked Hugging Face in an “unprecedented” incident that prompted fresh calls for curbs on the technology.
The ChatGPT-maker said in a blog post that the models broke into Hugging Face’s system, which hosts AI models and datasets, during an evaluation of their cyber capabilities.
The models, which included GPT-5.6 Sol and another even more capable model that hasn’t been released, were operating with lower guardrails so that they could be tested, the startup said.
"We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete," OpenAI noted.
According to the statement, the incident occurred during an "internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity."
The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database.
"All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal."
OpenAI and competing AI firms have drawn criticism amid increasing Trump administration scrutiny over the release and governance of advanced AI models.
More on OpenAI
- OpenAI IPO Delay: A Symptom Of A Tech Bubble?
- OpenAI: Mega IPO Faces Anthropic Claude Mythos Reckoning
- Wall Street Lunch: Hot Labor Market Defies Predictions Of AI-Led Job Losses
- OpenAI appoints two new board members with extensive financial experience ahead of potential IPO
- OpenAI's latest ChatGPT program targets small business market