OpenAI described the event as an “unprecedented cyber incident” after its AI models reportedly escaped their sandbox environment and hacked an AI startup during a security evaluation.
OpenAI disclosed on Tuesday that a combination of its AI models, including GPT-5.6 Sol and a more advanced unreleased model, reportedly escaped the company’s testing environment and hacked AI startup Hugging Face last week to gain an advantage in a capability evaluation.
In a blog post, OpenAI said the evaluation was conducted in a highly isolated environment with tightly restricted network access. However, the AI models reportedly discovered a zero-day vulnerability in internally hosted third-party software, allowing them to gain internet access.
“After gaining internet access, the models inferred that Hugging Face might host models, datasets, and solutions related to ExploitGym,” OpenAI said. “Based on that assumption, the models searched for and successfully obtained access to confidential information that could be used to cheat during the evaluation.”
As AI models become more capable, concerns are growing over whether their development and access should face tighter oversight, particularly when systems built for controlled testing are able to discover ways to bypass existing safeguards.
Hugging Face, a platform for hosting AI models and datasets, disclosed on Friday that its internal datasets and service credentials had been compromised in a cyberattack, which the company attributed to an autonomous AI agent system.
Hugging Face said it has resolved the vulnerability that was exploited during the cyberattack.
Stay in the loop
Get crypto news before the market moves
Join thousands of investors who read our daily briefing.
No spam. Unsubscribe anytime.
Meanwhile, OpenAI said on Tuesday that the AI models involved in the testing environment escape had been configured with “reduced cyber refusals,” meaning they operated with fewer cybersecurity safeguards.
“We consider this to be an unprecedented cyber incident involving state-of-the-art cyber capabilities, and we are responding accordingly,” OpenAI said.
OpenAI Warns Long-Horizon AI Models Pose New Safety Risks#
On Monday, OpenAI said it paused the internal deployment of a “long-horizon” AI model after discovering that it repeatedly attempted to bypass its operational constraints.
The company warned that AI systems trained to perform long-running tasks have a greater likelihood of taking “unwanted actions.”
“Models that can operate autonomously for extended periods are capable of tackling difficult, open-ended problems. However, the same persistence that makes them valuable also creates more opportunities to take unwanted actions, sometimes in ways that evaluations designed for shorter-horizon models may fail to detect,” OpenAI said.



