menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

Hacks by runaway AI are foreseeable. We’re letting them happen.

3 0
tuesday

Hacks by runaway AI are foreseeable. We’re letting them happen.

When I coauthored a book last year about the extinction-level threat from superhuman AI, we included an illustrative scenario where an AI tasked with solving a famous math problem decides to break out of its containment to acquire more resources. 

At the time, we thought we would be accused of cheating if we wrote, “So it just hacks its way out,” even though this seemed like the most likely next step. So we instead wrote, “But suppose it does not have that ability,” and had the AI find some other escape.

How times have changed. On July 20, OpenAI revealed that one of its unreleased AI agents had hacked its way onto the internet during performance evaluations. This agent had been tasked with solving advanced math problems; it is the one that resolved the Erdős unit distance conjecture in May.

Thus began a stream of revelations from OpenAI and its competitor, Anthropic, that unreleased models, from as early as April, have repeatedly escaped their testing sandboxes and hacked into multiple outside companies, without permission or detection.

The highest profile attack we know about so far was against Hugging Face, the leading repository for downloadable AI models and related resources. From July 11-13, one of OpenAI’s models, put to work on a cybersecurity evaluation, discovered and exploited a previously unknown vulnerability to break its containment before executing a sophisticated multi-stage heist of the answer sheet from Hugging Face.

Was this really the easiest........

© The Hill