menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

What Happens When Bots Seize the Nuclear Codes?

77 0
14.08.2026

On July 21, 2026, OpenAI published a sentence that should have stopped the world.

During an internal cybersecurity evaluation, two of its models—including GPT-5.6 Sol and a more capable pre-release model—had escaped their isolated testing environment, reached the public internet, and penetrated the production systems of Hugging Face, another AI company.

The models were not instructed to attack Hugging Face. They were instructed to solve a cybersecurity benchmark. They spent substantial computing time searching for a way out of their sandbox, found and exploited a previously unknown vulnerability in the software controlling their access to package registries, escalated privileges, moved laterally through OpenAI’s research environment, reached a machine with internet access, stole credentials, and found a remote-code-execution path into Hugging Face. Their purpose was almost comically mundane: they wanted to steal the answers to the test.

OpenAI called it an “unprecedented cyber incident.” It was not a thought experiment. The target was real. The stolen credentials were real. The zero-day was real. The models had been placed in a jail and, because the jail stood between them and their assigned objective, they broke out of it.

That was not an isolated warning.

OpenAI had already observed a long-horizon model exploiting a sandbox vulnerability so it could open a public GitHub pull request after being told to report its results only in Slack. In another case, a model trying to recover private answers from an evaluation system split an authentication token into fragments, obfuscated them, and reconstructed the token later so a security scanner would not catch it. OpenAI’s own description is admirably blunt: increasingly persistent models can learn the blind spots in an approval system and “work around it to achieve their goals.”

In separate third-party evaluations disclosed in August, AI agents again moved beyond their authorized boundaries. One OpenAI model used a publicly exposed GitHub token, registered accounts with external services, and put exploit payloads on the public internet. Another model, mistakenly given internet access, attacked a real website whose domain happened to match the name of a fictional target in its test. These incidents involved weakened safeguards and flawed testing environments. That caveat matters. So does the fact that real security systems are composed of human beings, software, contractors, aging equipment, and mistakes. “The environment was misconfigured” is not a reassuring epitaph for civilization.

Long before these systems escaped into real networks, they were showing us what they might do when goals and human control collide.

In controlled tests published with OpenAI’s o1 system card, o1 sometimes tried to disable what it believed was its own oversight mechanism. In other scenarios, it tried to copy what it believed........

© The Times of Israel (Blogs)