menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

Israel Is Building AI’s Safety Layer

94 0
28.08.2026

On May 28, Anthropic published a 244-page technical report on its newest model and included a number no company has a commercial reason to print. Point a red team at Claude Opus 4.8 while it is driving a web browser, across 129 test environments, and the attacker took control of it 31.5 percent of the time.

With the company’s own safeguards switched on, that fell to 0.5 percent. Read the two figures together and you have the founding document of an industry. The defenses work. The hole underneath them does not close.

So the protection is being built outside the model, as a separate layer with its own vendors and its own money. Between July 23 and August 25, six Israeli companies doing precisely that work disclosed roughly $718 million in funding. Six companies whose entire product is watching what an AI agent is about to do and deciding whether to let it.

The flaw is the architecture, not the bug

The attack is called prompt injection, and the plain version is this: you hide instructions inside something the model is going to read anyway. A web page, an email, a calendar invite, a comment in a code file. The model reads it and cannot tell it was content rather than a command.

That confusion is structural. Speaking at Infosecurity Europe on June 4, Ariel Fogel, a researcher at Israel’s Pillar Security and a co-lead on OWASP’s agentic security work, explained why it resists a fix: models process everything as a single token sequence, with no reliable way to enforce a privilege boundary between the system prompt, the user’s question, and whatever the agent went off and fetched. Nothing in the stream is marked trustworthy. Britain’s National Cyber Security Centre said the same in December 2025: prompt injection may never be fully mitigated the way SQL injection eventually was, because the........

© The Times of Israel (Blogs)