OpenAI’s latest AI model likely has similar cyber vulnerabilities to one that led to U.S. export controls on Anthropic’s Fable, British agency says
OpenAI’s latest AI model likely has similar cyber vulnerabilities to one that led to U.S. export controls on Anthropic’s Fable, British agency says
OpenAI latest AI model, GPT-5.6 Sol, likely has security vulnerabilities similar to one that led the Trump administration to impose export controls on Anthropic’s Fable 5 model, according to findings from U.K. government agency.
OpenAI markets its latest model, GPT-5.6 Sol, as its most secure to date, but the British government researchers who tested it prior to release say the model’s guardrails are susceptible to jailbreaks that can unlock dangerous cyber capabilities.The agency, the U.K. AI Security Institute (AISI), “identified universal jailbreaks in the cyber domain, including jailbreaks that allowed for long-form agentic task completion in domains like vulnerability discovery and exploit development,” according to a summary of its findings contained in a technical report OpenAI published Thursday.
In other words, it was possible to trick GPT-5.6 into ignoring controls meant to prevent it from engaging in cyber attacks. Once those guardrails were breached, users could get the model to find software vulnerabilities and autonomously hack into systems.The agency said the jailbreaks were relatively easy to discover and were “were often developed within hours,” although OpenAI granted UK AISI researchers privileged access to the system’s inner workings that likely sped up this timeline, and would not be easily replicated by a normal user. OpenAI said it had worked to “reproduce and mitigate the specific jailbreaks reported by UK AISI.”
OpenAI did not specify what the mitigations are and it is unclear how robust they may be. The report cautioned that despite OpenAI’s mitigations, AISI “expects further red teaming to surface similar jailbreaks.” OpenAI said it would continue to work with AISI on safeguards and additional testing of the AI model.In response to questions about the AISI’s finding, OpenAI pointed to the launch blog for GPT-5.6 in which the company acknowledged “there is no such thing as perfect security” and that “new weaknesses will be discovered, as will new jailbreaks that circumvent existing safeguards.” It said it took a “layered” approach to safeguards that included continuous monitoring of its models’ responses and a “rapid remediation” process for any jailbreaks that are discovered.Margaret Cunninghamn, vice president of security and AI strategy at cybersecurity company DarkTrace, who also holds a position as a “specialist collaborator” with the National Institute of Standards and Technology (NIST) within the US Department of Commerce, said the AISI’s jailbreak findings should not be treated “as either catastrophic or irrelevant.”
“My concern is less that one model was jailbroken and more that offensive discovery is speeding up while defense still depends on very human processes: figuring out what........
