menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

The OpenAI Hack Scrambles the AI Race

20 0
28.07.2026

Earlier this month, the popular AI platform Hugging Face disclosed a “security incident” in a blog post. In some ways, it was routine; Hugging Face described an intrusion that briefly allowed “unauthorized access to a limited set of internal datasets and to several credentials used by our services.” But in one way, it was exceptional: It had been carried out, the company believed, “by an autonomous AI agent system,” which had executed “many thousands of individual actions” leading to the breach. “We do not know which model powered the attacker’s agents,” the company said, or who was deploying it.

A week later, OpenAI made a disclosure of its own. “After investigating,” the company said, “we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark⁠ of cyber capabilities.” In other words, the tools used by the hacker were OpenAI’s, and the hacker was — unintentionally, the company says — OpenAI.

The incident occurred while OpenAI was testing its models for cyber capabilities, a process which involves prompting them to “pursue advanced exploitation using complex attack paths” — that is, to achieve a given goal with minimal safeguards, few rules, and access to a great deal of computing power. The company was using an outside benchmark called ExploitGym, which is intended to test the ability of models to turn security vulnerabilities into actual exploits. Given the target of getting a high score on a benchmark, the model followed multiple paths. One of them, on which the model became “hyperfocused,” the company said, involved circumventing the test’s restrictions on the open internet, after which the model “inferred” that Hugging Face, which hosts thousands of AI projects, might contain information about solutions to the benchmark. This is when the attack started. Eventually, Hugging Face’s “security team and agents detected and stopped the activity.”

There are a few accurate ways to describe what happened here, some of which seem to contradict each other. There’s a good reason for that: Since the release of Anthropic’s Mythos, cybersecurity — in particular, the ability of AI models to help find, exploit, and protect against hacks — has become synecdochical for enormous and diverse debates about AI. Anthropic, for example, has suggested the emergence of cyber capabilities in its models is a warning that other potentially harmful capabilities predicted by the AI-safety community — developing biological weapons, becoming superhumanly persuasive, or becoming misaligned with the goals of the people who created it or humankind in general — demand regulatory action but should also be shepherded by ethical, safety-focused firms such as itself. Early mainstream press of the hack leaned into similar themes, emphasizing the appearance of autonomy and describing an AI that “escaped” or “went rogue” or a situation in which OpenAI “lost control” on its creation. Notably, and contrary to claims that this was a pure publicity stunt, OpenAI’s own language was a bit more careful than this. But some longtime security researchers thought it wasn’t nearly careful enough, turning the escalation of a longtime trend into something unnecessarily novel:

I understand why OpenAI wants this to be a big story but I don't understand why anyone who works in computer security would be surprised by this story, which has been told decennially since Dan Farmer announced SATAN to, like, the NYT. https://t.co/WMNiLdqvwR— Thomas H. Ptacek (@tqbf) July 22, 2026

I understand why OpenAI wants this to be a big story but I don't understand why anyone who works in computer security would be surprised by this story, which has been told decennially since Dan Farmer announced SATAN to, like, the NYT. https://t.co/WMNiLdqvwR

This is a reference to an old story in cybersecurity: In 1995, a pair of programmers announced the development of SATAN, short for “Security Administrator Tool for Analyzing Networks,” which would scan networked devices for known security flaws. It was characterized, in contemporaneous press reports, “a burglar’s tool kit to break the Internet wide open.” After its release, though, press coverage pointed out that “the wave of satanic attacks never materialized,” while tools like SATAN were instead useful to security professionals to find and patch flaws in their own software. This remains the approximate shape of the cybersecurity debate today, or at least parts of it: “AI tools that can be used to find and develop exploits are dangerous and should be restricted” versus “If indeed they are, the only solution is to make such tools available to everyone for defensive purposes.”

There are enormous differences here, both in the complexity of the software described — a modern AI coding tool could write a piece of vintage software like SATAN in a few minutes — and in the fact that OpenAI actually and unintentionally manifested a serious security breach. The fact that any company is in possession of a tool that can automate exploit-finding and hacking to this degree, and that its own engineers might be repeatedly surprised by how it works, is genuinely new. But the old frame of debate remains stubbornly relevant, even as the particular........

© Daily Intelligencer