AIToday
Large Language ModelsAI Coding AssistantsSimon Willison's WeblogPublished: Aug 13, 2026, 10:01 JST1 min read

Claude Code auto mode becomes default for Pro, Max, Team users

Claude Code auto mode becomes default for Pro, Max, Team users

Key takeaway

  • Anthropic is making auto mode the default in Claude Code for Pro, Max, and Team users starting August 14th, citing internal testing showing it blocks 89% of harmful actions compared to only 13.6% blocked by human reviewers.

  • The move reflects confidence in auto mode's safety, though the company's claims of defeating prompt injection attacks remain unverified by independent security researchers.

3 Key Points

  1. What happened

    Anthropic is making auto mode the default setting for new Claude Code sessions across Pro, Max, and Team plans starting August 14th. The company claims auto mode blocks 89% of harmful actions in a test of 1,053 paid users, where participants were shown a dangerous command midway through sessions.

  2. Why it matters

    Auto mode reduces confirmation fatigue—asking humans to approve every step leads to unsafe behavior. In the same test, only 13.6% of humans refused the harmful action, suggesting humans approve risky commands at high rates under repeated prompts. Anthropic is betting auto mode is safer than human judgment for routine decisions.

  3. What to watch

    Anthropic commissioned Trajectory Labs to test 72 indirect prompt injection scenarios against Claude Fable 5, Opus 5, and Sonnet 5 running auto mode as of July 17th 2026, reporting none of the 720 attack attempts succeeded. However, the security community remains skeptical—the setting rolls out August 14th, and independent verification of these claims is still pending.

FAQ

When does auto mode become the default?
Auto mode becomes the default setting for new Claude Code sessions starting August 14th for Pro, Max, and Team plans.
What did Anthropic's testing show auto mode could block?
In a test of 1,053 paid users, auto mode blocked 89% of harmful actions that participants were prompted to approve midway through sessions, compared to only 13.6% of humans who refused the same harmful action.
Did third parties independently verify Anthropic's security claims?
Anthropic commissioned Trajectory Labs to test prompt injection scenarios; Trajectory Labs reported that none of 720 attack attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode. However, broader independent security community confirmation has not yet occurred.
Simon Willison's WeblogRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleShieldFont: New typeface poisons AI training data while keeping pages readable

The AI news that matters, in one minute each morning.

Sign up free