
What happened
Anthropic's Opus 5 model achieved zero percent attack success rate across 129 browser-based prompt injection test scenarios when Auto Mode is enabled in products like Claude Cowork. In a general prompt injection test by security firm Gray Swan, Opus 5's attack success rate dropped to 2.0 percent after 15 attempts, down from Opus 4.8's 5.5 percent, placing it first on the Gray Swan IPI benchmark.
Why it matters
Prompt injection — where attackers slip hidden instructions into AI inputs to bypass safety guardrails — has been a persistent security problem for AI agents handling web-based tasks. OpenAI acknowledged in December that the vulnerability may never be fully solved, so this near-elimination in a real-world scenario is significant for organizations deploying AI to interact with user-controlled content and websites.
What to watch
The zero-percent rate requires both Auto Mode defenses: a layer that scans for hidden instructions before processing, and another that blocks dangerous actions before execution. Without Auto Mode, Opus 5's browser injection success rate is 3.7 percent—highlighting that the model alone cannot eliminate the risk; the protective software wrapper is essential.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Prompt injection has emerged as one of the hardest security problems in AI deployment, particularly for autonomous agents that must interact with web pages and user-controlled inputs. When an AI agent reads a webpage, an attacker can embed hidden text or crafted prompts that trick the model into ignoring its original instructions—for instance, stealing data or executing unintended actions. This vulnerability has been largely unsolved because it requires the model itself to reliably distinguish between legitimate input and malicious instruction, a task that becomes harder as attackers become more creative.
Anthropicís approach with Opus 5 addresses the problem not by claiming the model is immune, but by wrapping it in protective software (Auto Mode) that handles the problem at a different layer. The first layer catches hidden instructions before they reach the model, and the second layer prevents dangerous actions even if an injection succeeds. By requiring an attacker to defeat both independently, the combined system achieves near-perfect defense in real-world browser scenarios. This layered defense is important because, as OpenAI acknowledged in December, relying solely on the model's own robustness may not be sufficient.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Oracle said Thursday that sales in its closely watched cloud infrastructure business jumped 121% to $7.4 billi…

Marvell Technology is targeting $18 billion in FY2028 revenue, raising its combined FY2027 and FY2028 expectat…

Philip Morris International and WSJ Intelligence unveiled The Cognition Index Research Report, a survey of mor…

Lam Research CFO Doug Bettinger said at the Citi Global TMT Conference that the company raised its calendar 20…

Broadcom's fiscal Q3 2026 AI semiconductor revenue rose 221% from a year earlier to $16.7 billion, and CEO Hoc…

Reuters reports that Nvidia plans a major expansion of data centre capacity in Australia to meet AI demand
