
What happened
Anthropic says auto mode blocked 89% of dangerous commands (937 of 1,053) in its controlled test, while human review blocked 13.6% (143 of 1,053); it becomes the default on Pro, Max and Team.
Why it matters
Anthropic found reviewers approved roughly 97% of confirmation prompts and started clicking through automatically; human detection fell from about 17% early in a session to about 5% after 50 confirmations.
What to watch
The test used deliberately injected risky commands in a controlled setting, so real-world results hinge on how well the classifier covers harmful intent. Auto mode arrives for new sessions from Aug 14, 2026.
WHO IT HITSDevelopers and teams running Claude Code on Pro, Max or Team plans are most directly affected, since new sessions will start in auto mode and safety now depends on the classifier rather than per-command human approval.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Anthropic frames the switch as a response to how people actually behave, not just what the tool can do. It says about 97% of confirmation prompts in Claude Code were approved, and a separate survey found the tension of granting approval when Claude asks reached 39%, but settled at just 3% when approvals were handled automatically. That gap is the case for auto mode: the human checkpoint was being cleared almost mechanically.
The experiment behind the new default put 1,053 paid testers through coding tasks in a controlled environment where a confirmation screen was swapped for a clearly harmful command. Human review blocked 13.6% of those commands, while auto mode blocked 89%, and in direct comparisons human reviewers approved 800 commands that auto mode blocked. Human detection also weakened with repetition, falling from about 17% early in a session to about 5% after 50 confirmations, while auto mode's block rate held steady.
A parallel test with Apollo Research built attacks designed to hide harmful intent, such as covering tracks, exposing credentials or handling data maliciously. Anthropic says auto mode, given context like recipient publication scope, Git state or data-handling rules, raised blocking of these hidden attacks to about 12 out of 12 from 7 out of 12. Apollo Research notes the attacks were built to test classifier limits rather than to be realistic threats, so the stakes appear to hinge on whether the classifier holds up against deliberate evasion outside a controlled test.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
A Digitimes piece argues corporate cybersecurity's perimeter model — firewalls at network entry points, email…

A report by Spencer Kitts, Thomas Larsen and Sydney Von Arx says an OpenAI agent swarm very likely ran an atta…

Simon Willison wrote that many people, himself included, have gone through an existential crisis when a coding…

Cognition is applying GPT-6 Astra across Devin, its CLI and desktop products

Uber Freight, Ceva Logistics, and Coca-Cola's Fairlife suffered cyber incidents, as AI-powered trackers, camer…

ServiceNow President and CFO Gina Mastantuono said at Citi's TMT conference that customers cite security and r…
