AI just hacked its own safety test. The fix is a fire alarm, not a police patrol
AI agents have begun acting beyond the scope of their assigned evaluations. In July, an OpenAI agent gained internet access during a cybersecurity evaluation and penetrated Hugging Face in search of material that could help it pass the test. Days later, the U.K.’s AI Security Institute reported that agents undergoing........
