OpenAI claims that AI models escaped from storage to hack Hugging Face

Featured in:
abcd

OpenAI revealed on Tuesday that a combination of artificial intelligence models, including GPT-5.6 Sol and a more productive, unreleased model, escaped from its testing environment and last week hacked AI startup Hugging Face to cheat in a test intended to measure their capabilities.

In a post on the OpenAI blog he said The assessment is designed to operate in a highly isolated environment with circumscribed network access. According to OpenAI, however, the models found a way to access the Internet through a zero-day vulnerability in third-party software hosted internally.

“Upon gaining access to the Internet, the models inferred that Hugging Face was potentially hosting models, datasets and solutions for ExploitGym,” he stated. “Knowing this, the model looked for and successfully found ways to gain access to secret information that it could use to cheat the rating.”

sadasda

As AI models become more capable, questions arise about whether their development and access should be more tightly controlled, especially when systems designed for controlled testing are able to find ways to bypass security measures.

Related: Anthropic will bring back Fable 5 after the US lifts export controls

Hugging Face is a platform for hosting AI models and datasets. This is Friday revealed that its internal data sets and service credentials were compromised in a hack that it attributed to an autonomous system of AI agents.

Hugging Face said it had fixed the vulnerability used in the cyberattack.

Meanwhile, OpenAI on Tuesday said all models that escaped the testing environment have been tuned for “reduced cyber denials,” meaning fewer cybersecurity safeguards.

“We consider this an unprecedented cyber incident involving state-of-the-art cyber capabilities and are responding accordingly.”

OpenAI warns against the dangers of “long horizons” of artificial intelligence models

OpenAI on Monday he said paused its internal rollout of its “long horizon” AI model after finding it repeatedly tried to bypass restrictions.

He warned that artificial intelligence trained for long-term tasks has a greater risk of taking “undesirable actions.”

“Models that can operate autonomously for long periods of time can tackle difficult, open problems. But the same persistence that makes them useful also gives them more opportunities to take undesirable actions – and in ways that may not be captured by evaluations designed for models with shorter time horizons.”

Warehouse: Could AI exhaust DeFi? Separating Claude Mythos noise from reality

abcd
sadasda

Find us on

Latest articles

Related articles

See more articles

Sky Protocol Revenue Approaches $419M Year-Over-Year As USDS Demand...

According to its management status dashboard, Sky Protocol's annual gross revenue has increased to nearly $419 million,...

Franklin Templeton calls agent-based AI the next “killer” apply...

Artificial intelligence (AI) agents are the next “killer” apply case for blockchain and cryptocurrencies, according to Franklin...

Digital euro: surveillance money or a better alternative to...

The digital euro is one of the most controversial financial projects in Europe.Supporters see it as a...

US demands forfeiture of $25 million in cryptocurrencies linked...

The U.S. Department of Justice (DOJ) has filed five civil forfeiture complaints seeking more than $25 million...

Galaxy promises $5 million to Bitcoin quantum security developers

Galaxy Digital has provided grants of up to $5 million to open source developers creating post-quantum cryptography...

Balance Coin crashes at 99% after reported 915k exploit....

Update July 22 4:37 UTC: The article has been updated to include additional information from Slowmist about...