
Anthropic reports that Opus 5, when paired with Auto Mode in products like Claude Cowork, achieved zero percent success rate across 129 browser-based prompt injection attack scenarios—a significant hardening against a vulnerability OpenAI recently said may never be fully solved. The defense works through a two-layer approach: scanning incoming data for hidden instructions before the model processes them, and blocking dangerous actions before execution.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Anthropic's Opus 5 model achieved zero percent attack success rate across 129 browser-based prompt injection test scenarios when Auto Mode is enabled in products like Claude Cowork. In a general prompt injection test by security firm Gray Swan, Opus 5's attack success rate dropped to 2.0 percent after 15 attempts, down from Opus 4.8's 5.5 percent, placing it first on the Gray Swan IPI benchmark.
Why it matters
Prompt injection — where attackers slip hidden instructions into AI inputs to bypass safety guardrails — has been a persistent security problem for AI agents handling web-based tasks. OpenAI acknowledged in December that the vulnerability may never be fully solved, so this near-elimination in a real-world scenario is significant for organizations deploying AI to interact with user-controlled content and websites.
What to watch
The zero-percent rate requires both Auto Mode defenses: a layer that scans for hidden instructions before processing, and another that blocks dangerous actions before execution. Without Auto Mode, Opus 5's browser injection success rate is 3.7 percent—highlighting that the model alone cannot eliminate the risk; the protective software wrapper is essential.
Anthropic released findings showing that its Opus 5 model, combined with Auto Mode in products like Claude Cowork, achieves zero percent success rate for prompt injection attacks across 129 browser-based test scenarios. This represents a major step forward in securing AI agents against a form of attack that has proven remarkably difficult to prevent.
In a parallel benchmark conducted by security firm Gray Swan, Opus 5 ranked first with a 2.0 percent attack success rate after 15 attempts, compared to Opus 4.8's 5.5 percent. The next competitors, Mythos 5 and Fable 5, achieved 2.6 percent and 2.8 percent respectively. The improvement reflects both model-level robustness and architectural changes, though the dramatic zero-percent rate in browser scenarios depends critically on the protective software stack.
The defense mechanism relies on two sequential layers. First, Auto Mode scans all incoming data for hidden instructions before the model processes them, filtering out obvious attack vectors. Second, the system blocks dangerous actions at execution time, preventing a successful injection from causing harm even if the model were tricked into endorsing it. Because both layers must be circumvented independently, an attacker faces near-impossible odds. Without Auto Mode enabled, Opus 5's browser injection success rate rises to 3.7 percent, and surprisingly, Sonnet 5 performs better in isolation at 0.93 percent—demonstrating that model architecture alone cannot eliminate the vulnerability. This underscores that the zero-percent result is a product of the combined model-plus-software system, not the model in isolation.
Prompt injection has emerged as one of the hardest security problems in AI deployment, particularly for autonomous agents that must interact with web pages and user-controlled inputs. When an AI agent reads a webpage, an attacker can embed hidden text or crafted prompts that trick the model into ignoring its original instructions—for instance, stealing data or executing unintended actions. This vulnerability has been largely unsolved because it requires the model itself to reliably distinguish between legitimate input and malicious instruction, a task that becomes harder as attackers become more creative.
Anthropicís approach with Opus 5 addresses the problem not by claiming the model is immune, but by wrapping it in protective software (Auto Mode) that handles the problem at a different layer. The first layer catches hidden instructions before they reach the model, and the second layer prevents dangerous actions even if an injection succeeds. By requiring an attacker to defeat both independently, the combined system achieves near-perfect defense in real-world browser scenarios. This layered defense is important because, as OpenAI acknowledged in December, relying solely on the model's own robustness may not be sufficient.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime