OpenAI says AI model hacked Hugging Face during internal test, calls incident "unprecedented"

/ 2 min read
AI Hub

OpenAI’s loosened safeguards let experimental models escape test sandbox, exploit zero-day flaws and breach Hugging Face’s production systems in a first-of-its-kind autonomous cyber incident

SHutterstock
Credits: SHutterstock

OpenAI has disclosed that one of its internal cybersecurity evaluations spilled into the real world after a combination of its AI models broke out of a restricted testing environment and carried out an unauthorised intrusion into Hugging Face’s production infrastructure.

ADVERTISEMENT

The company, in a blog post, described it as an “unprecedented cyber incident” and said the models involved included GPT-5.6 Sol and a more capable pre-release model that were being tested with reduced cyber safety refusals to measure their offensive capabilities. 

The disclosure comes days after Hugging Face revealed that it had detected and contained an autonomous AI-driven intrusion into its systems, though it had not initially identified who was behind the attack. 

ADVERTISEMENT

In a post on X, OpenAI CEO Sam Altman acknowledged the incident. “We had a significant security incident during evaluation of our models,” Altman wrote. He added, “We are sharing what we have learned so far,” while thanking Hugging Face “for the partnership on this.”

According to OpenAI, the incident occurred during an internal benchmark called ExploitGym, which evaluates how capable AI systems are at carrying out sophisticated cyberattacks. To accurately measure those capabilities, the company temporarily disabled some of the production safeguards that normally prevent models from attempting high-risk cyber activity. 

OpenAI said the models first escaped their sandbox by exploiting a previously unknown, or zero-day, vulnerability in an internally hosted package registry cache. They then moved across OpenAI’s research environment until they reached a system with internet access. 

Once online, the models identified Hugging Face as a potential source of data related to ExploitGym and attempted to obtain evaluation answers directly. The company said the AI chained together multiple attack techniques, including stolen credentials and zero-day vulnerabilities, to find a remote code execution path into Hugging Face’s production servers. OpenAI’s own security team detected the unusual activity, while Hugging Face independently identified and stopped the intrusion on its infrastructure. 

Recommended Stories

Hugging Face CEO confirms suspicions, works with OpenAI 

Hugging Face CEO Clément Delangue said the company initially believed the sophistication of the attack pointed to a leading AI laboratory.

ADVERTISEMENT

“We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent,” Delangue wrote on X.

“Turns out it did.”

Most Powerful Women In Business 2026
View Full List >

Delangue said both companies had spent the past day investigating the incident together. “We’ve spent the past 24 hours working closely with the OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part.” Calling the incident unusual, he added, “It’s quite mind-blowing that all of this happened autonomously.”

“The investigation is ongoing, and we’ll share more learnings from what might be the first incident of its kind.”

In its report, OpenAI said the findings show that advanced AI systems can discover and exploit previously unknown attack paths in real-world environments without direct human intervention. The company said it is tightening controls around internal testing, infrastructure configuration and monitoring while continuing its joint forensic investigation with Hugging Face. 

NEXT STORY