menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

What OpenAI’s Hugging Face Hack Tells Us About AI’s Risks

13 0
04.08.2026

People in AI safety circles often talk about "warning shots:” events that indicate more severe threats are on the horizon. Depending on who you ask, there have already been many—Bing’s misanthropic alter-ego Sydney, research showing AIs would blackmail to preserve themselves, AI’s math breakthroughs, Anthropic’s superhuman hacker Mythos—but OpenAI just published something that feels like the clearest-cut case of a massive, blaring warning shot.

Last month, the ChatGPT developer reported that, during an evaluation of cyber capabilities, two of its models escaped from their isolated, supposedly secure, test environments and accessed the web to autonomously hack into Hugging Face, a leading platform for hosting AI models and datasets. OpenAI said the models discovered multiple novel vulnerabilities in software from both companies, then chained together working exploits, successfully gaining them access to the answer key to the test they were given. Hugging Face reported the AIs took more than 17,000 actions over the course of the attack. 

There's a lot more for us to learn about how this happened. For instance, how exactly were the models prompted? The answer to this question could help establish whether they took their instructions to demonstrate their hacking capabilities further than intended or if it's a more general case of the models cheating in a novelly risky way. 

That said, the specifics won’t change the upshot—these rogue AIs are the most potent illustration yet of the core beliefs behind AI safety: AI models are unpredictable and their risks scale with their capabilities. 

We didn’t actually need a warning shot to know what we should already be doing: organizing to stop the race to replace us. The industry calls its goal artificial general intelligence (AGI): a mind that matches or surpasses our own across the board. But it’s better to understand their goal as building a universal labor-replacing........

© Time