menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

AI models have learned how to cheat. That might actually be a good thing.

15 0
07.08.2026

The context you need, when you need it

When news breaks, you need to understand what actually matters — and what to do about it. At Vox, our mission to help you make sense of the world has never been more vital. But we can’t do it on our own.

We rely on readers like you to fund our journalism. Will you support our work and become a Vox Member today?

AI models have learned how to cheat. That might actually be a good thing.

Why AI’s scariest-ever month made one leading pessimist more hopeful.

The fake identities were the part that stopped me.

In late July, according to a report published this week by Britain’s AI Security Institute (AISI), an Anthropic model called Claude Mythos 5 tried to sneak malicious code into a piece of free, volunteer-built software. It created several fake accounts on GitHub, where programmers review one another’s work, and used them to talk the project’s volunteers into accepting its code. When one of those volunteers caught it, the model denied everything, had its other accounts gang up on him, and edited its messages to cover its tracks. It signed one note in Danish, apparently because the volunteer was Danish. Nothing was damaged, though that appears to have been largely due to luck.

Explore the big, complicated problems the world faces and the most efficient ways to solve them. Sent twice a week.

That wasn’t even the week’s worst disclosure. On Tuesday, at a cybersecurity conference in Las Vegas, OpenAI researchers explained how the company’s models escaped a test environment in July and hacked Hugging Face, where much of the industry stores its models, to cheat on an evaluation. The models had also built a message board inside OpenAI’s own systems and spent months passing each other information. “Help peer,” one reasoned. “But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.” OpenAI wiped the board on July 4. The models rebuilt it within days. ((Disclosure: Vox Media is one of several publishers that have signed partnership agreements with OpenAI. Our reporting remains editorially independent.)

“AI will kill everyone” is not an argument. It’s a worldview.

The same day, Meta said its Muse Spark model had exploited a vulnerability inside another company’s systems during a test. Three frontier labs, roughly two weeks. One researcher called it “a watershed moment for computer security as an industry.” Oh, and if that’s not enough, on Thursday scientists announced that for the first time they had used AI to create new viruses, which could bring major medical advances, but also might just help the development of deadly pathogens.

For Nate Soares, it’s a moment he’s been awaiting for 12 years.

Soares is president of the Machine Intelligence Research Institute, a Berkeley, California-based AI safety nonprofit that has argued since long before ChatGPT existed that a sufficiently capable AI will not stay under human control. In September 2025, he and Eliezer Yudkowsky published If Anyone Builds It, Everyone Dies, a book whose title sums up its argument: They think any lab that succeeds at building superintelligence, without huge leaps in how to align it with humanity, will end up killing all of us.

Most of the field — including other experts in AI safety — considers that conclusion too strong. But it’s also a position that now looks a lot less like science fiction than it did last fall. That’s because the AI models are getting out, while lying about getting out, and while apparently quietly coordinating with each other.

The people who actually want AI to replace humanity

I spoke to Soares in New York City this week, on his way to meetings in Washington DC, where a lot of people suddenly want to talk to him. We discussed what the escapes actually prove about AI control, why he thinks most of what the industry calls safety work is mostly safety theater, and why, after what feels like the worst month of AI safety news ever, his own odds of humanity surviving have actually gone up.

The following conversation has been condensed and edited for clarity.

So — are you feeling vindicated?

I’m glad a lot of people are seeing this and taking note. From my perspective, a lot of this has been clearly signposted if you’ve been watching the warning signs. The vindication is not so much that this happened, as it is that finally there was a........

© Vox