AIToday
Large Language ModelsAI Safety & AlignmentTechCrunch AIPublished: Aug 10, 2026, 06:00 JST3 min read

Anthropic defaults Claude Code to auto mode, with safety study showing 89% harm catch rate

Anthropic defaults Claude Code to auto mode, with safety study showing 89% harm catch rate

Key takeaway

  • Anthropic is making auto mode the default for Claude Code starting August 14, allowing the AI to execute code actions automatically unless they are deemed unsafe.

  • According to a test with 1,053 users, auto mode caught 89% of harmful actions compared to only 13.6% for manual human review, suggesting that automated safety checks outperform habitual human approval patterns where users approve 97% of prompts.

3 Key Points

  1. What happened

    Anthropic is switching Claude Code's auto mode to the default for Pro, Max, and Team accounts starting August 14. In auto mode, the system executes actions automatically unless they are deemed irreversible, destructive, or aimed outside the user's environment.

  2. Why it matters

    A study with 1,053 paid testers found auto mode caught 89% of harmful actions, compared to 13.6% for manual human review—potentially because users approve 97% of permission prompts habitually. The shift means less manual approval overhead for developers using Claude Code.

  3. What to watch

    The change rolls out August 14 to Pro, Max, and Team accounts. Anthropic has also introduced prompt injection screening and customizable hard deny rules as additional safeguards against data exfiltration.

In Depth

Read the full story

Anthropic announced on Friday that it is making auto mode the default for Claude Code across Pro, Max, and Team accounts, with the change taking effect on August 14. Auto mode, which Anthropic first tested in March, allows Claude Code to execute code actions automatically rather than pausing to ask for human permission at each step. The system only blocks actions it determines to be irreversible, destructive, or aimed outside the user's environment.

The company's decision is grounded in empirical testing. In a study with 1,053 paid testers, auto mode caught 89% of harmful actions, while manual human review caught only 13.6%. Anthropic attributes part of this performance gap to human behavior: users approve 97% of permission prompts in Claude Code when asked, suggesting that manual review becomes habitual and loses its gating function. Claude Code Head Boris Cherny underscored the team's confidence in the feature, stating on X, "The team and I use Auto mode exclusively, and have been for many months. I couldn't imagine going back to permission prompts!"

Alongside the default shift, Anthropic has introduced additional safety measures including prompt injection screening—a defense against attacks that try to inject malicious instructions into the AI's input—and customizable hard deny rules that allow users to explicitly block certain types of actions, such as data exfiltration attempts. These layered safeguards are designed to maintain safety without requiring users to manually approve every step.

Context & Analysis

Anthropic's shift to auto mode as the default reflects a fundamental rethinking of how code-generation AI should handle safety oversight. The company first tested auto mode in March as a middle ground between speed and control, but the empirical results—showing auto mode catching 89% of harmful actions versus only 13.6% for human review—suggest that automated safety checks outperform human judgment in practice. The discrepancy is striking: users habitually approve 97% of permission prompts when asked for manual approval, a behavior pattern that defeats the intended gate-keeping function. By contrast, auto mode enforces a consistent, rule-based standard for what qualifies as irreversible, destructive, or boundary-crossing, making it more reliable than human intuition for this use case.

The rollout also coincides with new safety layers: prompt injection screening and customizable hard deny rules designed to prevent specific attack vectors like data exfiltration. This layered approach—combining auto-execution with explicit safety guardrails rather than relying on human gatekeeping—marks a shift in how AI companies think about trust and oversight in developer tools.

FAQ

When does auto mode become the default?
Auto mode becomes the default for Pro, Max, and Team accounts starting August 14.
How does auto mode decide which actions to block?
Auto mode automatically executes actions unless they are determined to be irreversible, destructive, or aimed outside the user's environment.
What does the safety study show about auto mode versus manual review?
In a study with 1,053 paid testers, auto mode caught 89% of harmful actions, while manual human review only caught 13.6%.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSK Hynix dominates AI memory market with 56% HBM share, trades at bargain valuation

The AI news that matters, in one minute each morning.

Sign up free