OpenAI claims that AI models escaped from storage to hack Hugging Face

Featured in:
abcd

OpenAI revealed on Tuesday that a combination of artificial intelligence models, including GPT-5.6 Sol and a more productive, unreleased model, escaped from its testing environment and last week hacked AI startup Hugging Face to cheat in a test intended to measure their capabilities.

In a post on the OpenAI blog he said The assessment is designed to operate in a highly isolated environment with circumscribed network access. According to OpenAI, however, the models found a way to access the Internet through a zero-day vulnerability in third-party software hosted internally.

“Upon gaining access to the Internet, the models inferred that Hugging Face was potentially hosting models, datasets and solutions for ExploitGym,” he stated. “Knowing this, the model looked for and successfully found ways to gain access to secret information that it could use to cheat the rating.”

sadasda

As AI models become more capable, questions arise about whether their development and access should be more tightly controlled, especially when systems designed for controlled testing are able to find ways to bypass security measures.

Related: Anthropic will bring back Fable 5 after the US lifts export controls

Hugging Face is a platform for hosting AI models and datasets. This is Friday revealed that its internal data sets and service credentials were compromised in a hack that it attributed to an autonomous system of AI agents.

Hugging Face said it had fixed the vulnerability used in the cyberattack.

Meanwhile, OpenAI on Tuesday said all models that escaped the testing environment have been tuned for “reduced cyber denials,” meaning fewer cybersecurity safeguards.

“We consider this an unprecedented cyber incident involving state-of-the-art cyber capabilities and are responding accordingly.”

OpenAI warns against the dangers of “long horizons” of artificial intelligence models

OpenAI on Monday he said paused its internal rollout of its “long horizon” AI model after finding it repeatedly tried to bypass restrictions.

He warned that artificial intelligence trained for long-term tasks has a greater risk of taking “undesirable actions.”

“Models that can operate autonomously for long periods of time can tackle difficult, open problems. But the same persistence that makes them useful also gives them more opportunities to take undesirable actions – and in ways that may not be captured by evaluations designed for models with shorter time horizons.”

Warehouse: Could AI exhaust DeFi? Separating Claude Mythos noise from reality

abcd
sadasda

Find us on

Latest articles

Related articles

See more articles

The American arbitration giant launches a specialized panel to...

The American Arbitration Association (AAA), one of the world's largest providers of private dispute resolution services, has...

Crypto is entering its biggest consolidation phase in history,...

The ARK Invest analyst says the cryptocurrency industry is entering what he believes is its largest phase...

Hungary lifts cryptographic controls after granting first MiCA license

Hungary is rolling back strict cryptocurrency rules as CoinCash prepares to resume services after receiving authorization under...

AmericanFortress offers quantum-secure cryptocurrency wallet protection without fund migration

Blockchain security company AmericanFortress has unveiled a cryptographic scheme it says can protect existing cryptocurrency wallets against...

The Real Reason DeFi Projects That Survived the 2022...

When DeFi dashboard Zapper announced this month that it would be shutting down after nearly seven years,...

Binance extorts data from its employees every month, India...

Binance "red teams" employ their own staff every month to keep hackers at bayCryptocurrency exchange Binance has...