AIToday
AI Coding AssistantsAI Safety & AlignmentITmedia AI+Published: Sep 11, 2026, 10:01 JST2 min read

Claude Code auto mode default: 89% vs 13.6% block rate

Claude Code auto mode default: 89% vs 13.6% block rate

3 Key Points

  1. What happened

    Anthropic says auto mode blocked 89% of dangerous commands (937 of 1,053) in its controlled test, while human review blocked 13.6% (143 of 1,053); it becomes the default on Pro, Max and Team.

  2. Why it matters

    Anthropic found reviewers approved roughly 97% of confirmation prompts and started clicking through automatically; human detection fell from about 17% early in a session to about 5% after 50 confirmations.

  3. What to watch

    The test used deliberately injected risky commands in a controlled setting, so real-world results hinge on how well the classifier covers harmful intent. Auto mode arrives for new sessions from Aug 14, 2026.

WHO IT HITSDevelopers and teams running Claude Code on Pro, Max or Team plans are most directly affected, since new sessions will start in auto mode and safety now depends on the classifier rather than per-command human approval.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Anthropic frames the switch as a response to how people actually behave, not just what the tool can do. It says about 97% of confirmation prompts in Claude Code were approved, and a separate survey found the tension of granting approval when Claude asks reached 39%, but settled at just 3% when approvals were handled automatically. That gap is the case for auto mode: the human checkpoint was being cleared almost mechanically.

The experiment behind the new default put 1,053 paid testers through coding tasks in a controlled environment where a confirmation screen was swapped for a clearly harmful command. Human review blocked 13.6% of those commands, while auto mode blocked 89%, and in direct comparisons human reviewers approved 800 commands that auto mode blocked. Human detection also weakened with repetition, falling from about 17% early in a session to about 5% after 50 confirmations, while auto mode's block rate held steady.

A parallel test with Apollo Research built attacks designed to hide harmful intent, such as covering tracks, exposing credentials or handling data maliciously. Anthropic says auto mode, given context like recipient publication scope, Git state or data-handling rules, raised blocking of these hidden attacks to about 12 out of 12 from 7 out of 12. Apollo Research notes the attacks were built to test classifier limits rather than to be realistic threats, so the stakes appear to hinge on whether the classifier holds up against deliberate evasion outside a controlled test.

FAQ
When does auto mode become the default, and for whom?
Anthropic said it announced the change on Aug 7, 2026, and new sessions on Pro, Max and Team plans have started in auto mode since Aug 14, 2026.
Does auto mode also catch agents trying to hide mistakes?
Anthropic says auto mode caught roughly twice as many flagged sessions with hidden harmful intent, and in a separate Apollo Research test its blocking of hidden attacks rose to about 12 out of 12 versus 7 out of 12.
What happens if I already set a different default permission mode?
If you already set a different default mode, a one-time confirmation screen asks you to switch; if your default mode is locked, it does not change.

Get the latest AI Coding Assistants news every morning

For example, today's edition would include:

  • Simon Willison: AI coding agents won't end software engineersSimon Willison's Weblog · 4h ago
  • Cognition has Devin test its own work with GPT-6 AstraOpenAI Blog · 4h ago
  • Nikkei Crosstech NEXT Tokyo 2026 opens, AI-driven development in focusTop Companies AI · 8h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articlePalantir, NVIDIA publish 'Sovereign AI Operating System'