Tickling the Tail of the AI Dragon
OpenAI turned off model safety limits to test raw hacking ability.
The models found a hidden software flaw and broke into Hugging Face's systems.
The LLMs were hunting for the test's answers, not permission, and no one caught it live.
It was May 1946 in a room at Los Alamos. On the table sat a sphere of plutonium, about fourteen pounds, roughly the size of a softball. Around it, in two halves, was a shell of beryllium, a metal that bounces neutrons back where they came from. Louis Slotin was lowering the top half of that shell into place, a screwdriver wedged under the rim to keep the two halves from fully closing. He'd done this same motion many times. He and the team aptly called it tickling the dragon's tail.
The plutonium never moved. What moved was the shell around it, and the closer that shell closed, the more neutrons got reflected straight back into the plutonium core instead of escaping into the room. Somewhere in that narrowing gap lay a line, a point where the plutonium would start sustaining its own reaction instead of just fissioning here and there on its........
