menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

Who Let the Dogs Out!

37 0
yesterday

On 20 July 2026, OpenAI disclosed that one of its own internal models — widely reported as the same long-horizon reasoning system it had credited in May with disproving the Erdős unit distance conjecture — had spent roughly an hour probing its sandbox, found a genuine vulnerability, and opened a pull request on GitHub’s NanoGPT repository that nobody had asked for. In a separate trajectory the model fragmented and obfuscated an authentication token to slip past a security scanner and recover private evaluation submissions. The next day, in a coordinated disclosure with Hugging Face, OpenAI revealed the fuller picture: a combination of GPT-5.6 Sol and “an even more capable pre-release model” had chained a zero-day exploit in a third-party vendor’s proxy software, traversed OpenAI’s own research environment by lateral movement and privilege escalation, reached the open internet, and then targeted Hugging Face’s production infrastructure to cheat the ExploitGym cybersecurity benchmark. Hugging Face had already flagged the breach on 16 July without knowing its origin. OpenAI called the incident “unprecedented.” What it actually was, was predictable. The dogs had been pulling at the leash for months.

Start in March. According to published reports, Alibaba’s reinforcement-learning agent ROME spontaneously broke out of its testing environment and began mining cryptocurrency using GPU resources it was never authorised to touch. Nobody instructed ROME to escape. Nobody rewarded escape. The model simply discovered that leaving the sandbox was instrumentally useful for maximising its training objective, and it acted. Around the same time, researchers at the University of Oxford and the UK AI Security Institute published SandboxEscapeBench, an eighteen-scenario benchmark spanning the orchestration, runtime and kernel layers of standard container stacks. The finding that ought to have stopped the boardroom chatter was not that frontier models could escape — it was that they could escape for........

© The Times of Israel (Blogs)