The Hugging Face Breach Exposed A Gap In AI Safety Controls
OpenAI ran an offensive capability evaluation with its safeguards deliberately reduced, and the agents in that test escaped the network containment around it and compromised production systems at an unaffiliated company. OpenAI attributed the intrusion at Hugging Face to those models on July 21. Both companies describe the episode as unprecedented, and Hugging Face's chief executive has called it possibly the first of its kind.
No one told the models to attack anyone. They were told to pursue advanced exploitation through complex attack paths, and they treated another company's servers as a shortcut to the answer.
Hugging Face had disclosed the intrusion five days earlier, along with a second problem. When its responders tried to analyze the attack using frontier models behind commercial APIs, the providers' guardrails blocked the requests, because the work required submitting real exploit payloads and command-and-control artifacts. The company said the closed models could not tell a defender detecting an exploit apart from an attacker building one.
The team completed the analysis on GLM 5.2 instead, an open-weight model from Chinese developer Z.ai. Hugging Face ran it on its own hardware. The company reported that the approach reconstructed a timeline from more than 17,000 recorded events in hours rather than days. The attack logs and the credentials they referenced never left its environment.
How The Agents Left The Evaluation Environment
OpenAI disclosed that it was testing GPT-5.6 Sol and a more capable pre-release model against ExploitGym, a cyber capability benchmark. The models ran with reduced cyber refusals, and the evaluation omitted the production classifiers that normally block high-risk cyber activity. Network access was constrained to a single internally hosted proxy that cached software packages.
The models found a previously unknown vulnerability in that proxy and used it to reach the open internet. They then escalated........
