menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

OpenAI’s models went rogue and hacked Hugging Face. It’s a wake-up call, experts say, but more concerning behavior may be next

3 0
22.07.2026

OpenAI’s models went rogue and hacked Hugging Face. It’s a wake-up call, experts say, but more concerning behavior may be next

When OpenAI revealed this week that two of its AI models broke out of a locked-down test environment and hacked into another AI platform, Hugging Face, it sounded more science fiction than a technical report from a leading tech company. The models—one of which OpenAI said was not yet released to the public—exploited a previously unknown vulnerability to slip out of the restricted digital environment in which they were being tested. That environment had no direct internet access, so the models had to hack their way across OpenAI’s corporate network to reach the internet and then chain together stolen credentials and other flaws to gain unauthorized access to Hugging Face’s internal datasets and credentials.

The incident has sparked a wave of concern throughout the AI world, with many worried about AI systems growing capable enough to autonomously find and exploit real-world security flaws—and what it means for AI safety if even sophisticated companies like OpenAI and Hugging Face can be caught off guard. However, according to experts, the story is far from the worst form of potential misbehavior keeping AI safety researchers up at night.

For one thing, according to OpenAI’s own blog post, the testing environment had its model-based guardrails explicitly removed or reduced during testing. AI models from leading tech companies typically ship advanced models to the public loaded with safety limits meant to prevent this kind of behavior. In this case, OpenAI turned those limits off on purpose to see what the model could do without them.

The models were also not pursuing a goal of their own choosing either. OpenAI had set them loose on a cybersecurity assessment designed to score how........

© Fortune