OpenAI Hacking Fiasco Exposes a “Deeply Insufficient” System to Protect the Public
The incident sounded straight out of a science fiction movie: OpenAI’s super-advanced tool hacked another AI company’s systems in an attempt to pass its own developers’ cybersecurity test. Just replace the AI tech with a newly engineered virus and you have an entire existing subgenre.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” OpenAI wrote in a Tuesday blog post explaining the incident. The tech giant said their tool, designed to execute tasks without any human assistance, independently breached Hugging Face, another startup that hosts a voluminous number of publicly available AI models, while it was testing internally how good it was at “advanced exploitation using complex attack paths” within a supposedly enclosed lab environment called a sandbox. In other words, OpenAI was testing its own hacking capabilities, and the brakes came off; the tool broke out and onto the open internet, and that’s when the mischief began.
OpenAI said they had the situation under control: Hugging Face detected the breach last week and stopped the activity on their own (and called the cops). Since then, OpenAI said it was working with Hugging Face on addressing vulnerabilities.
But the incident—along with many others my colleagues have reported about—brings up countless regulatory concerns as the industry, and the public at large, grapples with what actually went down at OpenAI and the safety of autonomous agents. (The Center for Investigative Reporting, the parent company of Mother Jones, has sued OpenAI for copyright violations. OpenAI has denied the allegations.)
To better understand what actually happened—and to what extent we should be worried—I spoke with Miranda Bogen, who works on developing and promoting AI governance that incorporates technical expertise as the chief technologist at the Center for Democracy & Technology and the founding director of its AI Governance Lab.
This interview has been condensed and edited for clarity.
What was your immediate reaction to hearing the news come out on Tuesday?
The rhetoric was very overblown. The headlines made it out that a model had run amok, and that it was a complete surprise, and that it was something people might be exposed to. But what was really happening was that OpenAI was testing a new version of a system made up of multiple of its models. It was specifically within a sandboxed environment, and they were basically trying to get it to demonstrate capabilities in executing cyber attacks. What ended up happening was that the system identified a vulnerability in a part of the sandbox setup, and it used that vulnerability to access the internet to find the answer key for the test, which led it to try to figure out if the answers to the test were on Hugging Face in a non-public setup.
That still........
