menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

In the AI arms race, our safety is being put last

16 0
08.08.2026

The AI race between the United States and China is accelerating at a breathtaking pace. Time and again, the US has appeared to pull ahead, only to see China quickly close the gap.

Subscribe now for unlimited access.

Login or signup to continue reading

At the heart of this competition lie two contrasting visions of how AI should be developed and deployed. While US firms largely rely on proprietary technologies to maintain their competitive advantage, China has increasingly embraced open-source AI, making advanced models widely available to developers around the world.

Different strategies notwithstanding, both countries are locked in an arms race which neither side is willing to slow. As speed becomes the overriding priority, safety increasingly takes a backseat, despite the uncomfortably high risk of catastrophic failure arising from the combination of human error and super-intelligent AI.

A recent security breach involving OpenAI and Hugging Face has underscored the danger of treating safety as an afterthought. The incident occurred while OpenAI researchers were evaluating the cyber security capabilities of two frontier AI models. To identify weaknesses and vulnerabilities that sophisticated hackers might exploit, they temporarily disabled many of the models' normal safeguards.

The experiment was conducted inside a tightly controlled "sandbox". Although the model's behavioural guardrails had been temporarily removed, the sandbox was supposed to prevent them from accessing the open internet, where they could - at least in theory - have caused enormous damage.

Instead, the models treated the sandbox itself as an obstacle. Exploiting a previously unknown flaw in third-party software, they circumvented its restrictions and gained access to the open internet.

Did they go rogue? Not exactly. After all, they were simply pursuing the objective they had been given: to solve a difficult cyber security problem as efficiently as possible.

Although the models likely could have completed the task from within the sealed sandbox, they "inferred" that it would be far easier to infiltrate Hugging Face - which hosts over 2 million open-source AI models - and, in OpenAI's words, "obtain test solutions directly........

© Canberra Times