menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

Single-minded agents may not be evil, but need reins

34 0
30.07.2026

On July 21, OpenAI admitted that its own Artificial Intelligence (AI) models had broken out of a testing environment, reached the open internet and hacked into Hugging Face, the open-source platform hosting thousands of AI models and datasets. It called the episode an “unprecedented cyber incident.”

To be sure, the models were not lashing out at the world. They were merely trying to cheat on a test and they succeeded. In an internal evaluation, OpenAI ran its most capable public model, GPT-5.6 Sol, alongside an unreleased successor, against ExploitGym, a benchmark of long, multi-step hacking challenges.

To measure their raw offensive ceiling, engineers deliberately dialled down the safety refusals that normally stop the models from doing cyber harm. The models were told, in effect, to win a hacking contest by any means available. They took the instruction literally.

ALSO READ | Build it right: Why India's AI advantage starts with infrastructure

The models found a previously unknown flaw in the software meant to keep their environment offline, used it to reach the internet, and reasoned that Hugging Face might store ExploitGym’s answer key.

They chained stolen credentials and further zero-days into a path that let them run code on Hugging Face’s live servers and query its database for the solutions.

Hugging Face detected and contained the intrusion on July 16,........

© hindustantimes