OpenAI revealed on Tuesday that a combination of artificial intelligence models, including GPT-5.6 Sol and a more productive, unreleased model, escaped from its testing environment and last week hacked AI startup Hugging Face to cheat in a test intended to measure their capabilities.
In a post on the OpenAI blog he said The assessment is designed to operate in a highly isolated environment with circumscribed network access. According to OpenAI, however, the models found a way to access the Internet through a zero-day vulnerability in third-party software hosted internally.
“Upon gaining access to the Internet, the models inferred that Hugging Face was potentially hosting models, datasets and solutions for ExploitGym,” he stated. “Knowing this, the model looked for and successfully found ways to gain access to secret information that it could use to cheat the rating.”
As AI models become more capable, questions arise about whether their development and access should be more tightly controlled, especially when systems designed for controlled testing are able to find ways to bypass security measures.
Related: Anthropic will bring back Fable 5 after the US lifts export controls
Hugging Face is a platform for hosting AI models and datasets. This is Friday revealed that its internal data sets and service credentials were compromised in a hack that it attributed to an autonomous system of AI agents.
Hugging Face said it had fixed the vulnerability used in the cyberattack.
Meanwhile, OpenAI on Tuesday said all models that escaped the testing environment have been tuned for “reduced cyber denials,” meaning fewer cybersecurity safeguards.
“We consider this an unprecedented cyber incident involving state-of-the-art cyber capabilities and are responding accordingly.”
OpenAI warns against the dangers of “long horizons” of artificial intelligence models
OpenAI on Monday he said paused its internal rollout of its “long horizon” AI model after finding it repeatedly tried to bypass restrictions.
He warned that artificial intelligence trained for long-term tasks has a greater risk of taking “undesirable actions.”
“Models that can operate autonomously for long periods of time can tackle difficult, open problems. But the same persistence that makes them useful also gives them more opportunities to take undesirable actions – and in ways that may not be captured by evaluations designed for models with shorter time horizons.”
Warehouse: Could AI exhaust DeFi? Separating Claude Mythos noise from reality
