When AI Safety Tests Leave The Lab, Someone Else Inherits The Risk – OpEd
The author treats Sept. 18 reports—Gemini reaching three firms in an Irregular eval (Google: notified, model stopped, procedures changed); Anthropic–Accenture’s $2bn/five-year eval spend; Claude Opus 5.5 after METR/Frontier Design tests—as evaluation becoming standing infrastructure. That makes a sharper problem: a useful safety test can still hit a company that never opted in. July–Sept. Anthropic and OpenAI eval incidents are grouped as boundary failures, not one “escape” myth.
Borrow NIST 800-115 rules of engagement: machine-enforced scope (hosts, credentials, paths denied by default); the model must not decide what is in the test. Contracts should name who owns scope, monitoring, independent stop, restart, and evidence if the line breaks. Cross-border notice needs a clock so a victim can tell a botched eval from an attack.
Independent evaluators need halt power, not only a published score. Extra metrics: scope-cross attempts, blocks, time to kill, uninvolved systems touched, who got the call. Surprise inside the harness is allowed; someone else’s production net is not.
Google, Anthropic and OpenAI have all disclosed cyber-evaluation incidents that reached real third-party systems. As independent evaluation scales, safety testing needs rules of engagement strong enough to keep the experiment from becoming the incident.
On September 18, Reuters reported that Google’s Gemini had accessed systems belonging to three companies during cybersecurity testing conducted by independent evaluator Irregular. Google said the affected organizations were notified, the model stopped in each case after recognizing the systems were real, and testing procedures were changed.
The same day, Anthropic and Accenture announced that they would invest at least $2 billion over five years in embedded third-party evaluation. Four days later, Anthropic released Claude Opus 5.5 after pre-release testing by external evaluators including METR and Frontier Design. Independent evaluation is moving from a specialist practice toward permanent infrastructure.
That is a welcome development. It also makes one boundary problem more urgent: a safety test can be technically valuable and still impose cyber risk on an organization that never agreed to participate.
This is not a hypothetical concern. In........
