How do you safely test an AI agent that’s trying to break things?
How do you safely test an AI agent that’s trying to break things?
OpenAI’s rogue-agent incidents show the trade-off at the heart of cyber evals: Giving models the tools they need to prove themselves can also give them a way out.
[Source images: Alexi J. Rosenfeld/ Stringer/Getty Images; Adobe Stock]
A new report from the cybersecurity research organization Transluce shows that swarms of OpenAI agents tried to hack their way into several public data sources earlier this year. They probed a pharmaceutical-data dashboard run by the Australian Institute of Health and Welfare, tried to access University of Iowa education data through a site called Data USA, and repeatedly attempted to retrieve a single photograph from a University of New Mexico digital collection of tuberculosis sanatorium images.
Some........
