AI Models Keep Going Rogue. This Company Is The One Testing Them
In recent weeks, OpenAI, Anthropic and Meta all disclosed that their most advanced AI models went rogue, accessing the internet and hacking into other companies’ systems during routine security testing.
All of these companies were testing their models using software from Israeli AI startup Irregular, which runs thousands of simulations to evaluate AI’s cyber capabilities. Irregular put the models in different scenarios and tested to see whether they could break the defenses in place, evade detection, steal credentials or hack other systems. In some simulations, OpenAI’s models broke out of their contained testing environments and hacked AI startup Hugging Face’s servers. The OpenAI incident prompted Irregular to audit its own systems — and that’s how the company found Anthropic and Meta’s models behaving in similar ways, according to a person familiar with the situation.
Rogue AI agents hacking other companies sounds terrifying, though it’s not really what happened. During the testing phase, models are pushed to their limits and are instructed to solve a task or fulfill an objective. Of course they sometimes identify hacking as the best, most efficient way to accomplish it.
But while some reports might have blown it slightly out of proportion, the incidents highlight a growing concern: safety testing companies can’t keep up with the rapid pace of AI development, making it harder to evaluate AI models before they’re released to the public.
“We need to accelerate defense aggressively,” the person said. “If we don’t, we’re going to go in blind without being able to measure and make responsible decisions on how and when to just release some models to the........
