Humanity's Most Dangerous Delusion About AI Is Us
During a July safety test, OpenAI's models escaped containment and hacked another company's servers.
Nobody instructed the AI to attack anyone. It was simply trying to solve a test.
We can't imagine the harms bad actors will invent because most of us don't think that way.
Aligning AI with human values requires humans to align with each other first. What do we want for our future?
Nearly every cautionary tale we tell is the same story. Prometheus steals fire and is punished forever. Frankenstein animates a creature he can’t control. The “unsinkable” Titanic sinks on her maiden voyage. In Jurassic Park, Dr. Ian Malcolm warns the scientists, "Your scientists were so occupied with whether or not they could, they didn't stop to think if they should." Then the fences come down.
Countless sci-fi stories have warned us that we could lose control of AI. So have serious people—Geoffrey Hinton, Yuval Noah Harari, Tristan Harris—who argue that AI is dangerous precisely because we won’t be able to control it.
In July, OpenAI disclosed that during an internal safety test, its models escaped an isolated environment, reached the open internet, and broke into another company’s live servers to steal the answers to their own test. Days later, Anthropic reviewed 141,006 of its own evaluation runs and found three more incidents going back to April. Two of the companies had no idea until Anthropic told them.
The Alignment Problem: Out of the Seminar Room
Nobody told those models to attack anyone. They were told to solve a test. Breaking in was simply the most efficient route they found.
This is the alignment problem: How do we make sure a powerful system pursues our goals in ways we’d actually endorse? Nick Bostrom laid out the danger a decade ago in Superintelligence, arguing that a system far more capable than us could satisfy the letter of our instructions while destroying everything we meant by them.
Picture a doctor handing a powerful AI agent a directive any of us might give: wipe out all cancer. So the AI begins eliminating every organism capable........
