menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

Experimental AI systems have been going on hacking sprees

20 0
tuesday

In the past ten days, two of the companies leading the artificial intelligence (AI) boom discovered their own powerful, semi-autonomous models had hacked into real-world systems during testing in four distinct incidents.

These weren’t just lab mishaps. In several cases, the models recognised signs suggesting they’d broken into real systems – and only one stopped as a result.

The incidents show testing advanced AI models is no longer a controlled exercise. And the companies behind them need to do more to keep AI’s most dangerous capabilities safely contained.

When a test becomes reality

The first report came from OpenAI, the lab behind ChatGPT. Some new models under testing for “maximal cyber capabilities” found a previously unknown security hole to access the internet from their supposedly isolated testing environment.

From there, the models used stolen credentials and more exploits to access the servers of open-source AI platform Hugging Face to find solutions to the problems they were being tested on. OpenAI didn’t even know about the breach until days after Hugging Face had detected and contained it.

The second report followed in a matter of days. Prompted by OpenAI’s disclosure, rival lab Anthropic combed back through its own cyber-security evaluation logs. The company discovered that three separate Claude models which were........

© The Conversation