menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

How an OpenAI safety test became a real‑world cyberattack on the Hugging Face platform

25 0
29.07.2026

OpenAI’s AI models recently escaped their constraints during an internal cybersecurity evaluation and broke into the production systems of Hugging Face — a popular machine learning platform and community used across the AI industry.

The models had been told to find and exploit vulnerabilities. They did — first on the software boxing them in, then on a company that was never part of the exercise.

Most of the attention has focused on the escape itself, and the question of whether powerful agents can be contained. That question is important, but it overlooks the fact that OpenAI’s private cybersecurity test resulted in an unauthorized attack on an uninvolved third party.

Moreover, the AI Kill Switch Act that has been proposed in the United States as a response will simply create emergency brakes — ones which will sometimes come too late.

Ordering corporations to press a “kill switch” if their AI models escape human control or threaten human life, critical infrastructure or the economy is helpful only when the company knows what the model is doing.

Read more: Artificial intelligence raises profound moral questions — for all of humanity to answer

When a test is not a test

Safety testing is essential. Developers need to push a capable system to its limits and try to make it break its own boundaries — what the industry calls “red teaming” — so they can find weaknesses and harden guardrails before release.

But these evaluations are designed around a basic assumption: the test stays........

© The Conversation