AIToday
Large Language ModelsAI Safety & AlignmentTechCrunch AIPublished: Aug 10, 2026, 06:00 JST

Anthropic defaults Claude Code to auto mode, with safety study showing 89% harm catch rate

Anthropic defaults Claude Code to auto mode, with safety study showing 89% harm catch rate

3 Key Points

  1. What happened

    Anthropic is switching Claude Code's auto mode to the default for Pro, Max, and Team accounts starting August 14. In auto mode, the system executes actions automatically unless they are deemed irreversible, destructive, or aimed outside the user's environment.

  2. Why it matters

    A study with 1,053 paid testers found auto mode caught 89% of harmful actions, compared to 13.6% for manual human review—potentially because users approve 97% of permission prompts habitually. The shift means less manual approval overhead for developers using Claude Code.

  3. What to watch

    The change rolls out August 14 to Pro, Max, and Team accounts. Anthropic has also introduced prompt injection screening and customizable hard deny rules as additional safeguards against data exfiltration.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Anthropic's shift to auto mode as the default reflects a fundamental rethinking of how code-generation AI should handle safety oversight. The company first tested auto mode in March as a middle ground between speed and control, but the empirical results—showing auto mode catching 89% of harmful actions versus only 13.6% for human review—suggest that automated safety checks outperform human judgment in practice. The discrepancy is striking: users habitually approve 97% of permission prompts when asked for manual approval, a behavior pattern that defeats the intended gate-keeping function. By contrast, auto mode enforces a consistent, rule-based standard for what qualifies as irreversible, destructive, or boundary-crossing, making it more reliable than human intuition for this use case.

The rollout also coincides with new safety layers: prompt injection screening and customizable hard deny rules designed to prevent specific attack vectors like data exfiltration. This layered approach—combining auto-execution with explicit safety guardrails rather than relying on human gatekeeping—marks a shift in how AI companies think about trust and oversight in developer tools.

FAQ
When does auto mode become the default?
Auto mode becomes the default for Pro, Max, and Team accounts starting August 14.
How does auto mode decide which actions to block?
Auto mode automatically executes actions unless they are determined to be irreversible, destructive, or aimed outside the user's environment.
What does the safety study show about auto mode versus manual review?
In a study with 1,053 paid testers, auto mode caught 89% of harmful actions, while manual human review only caught 13.6%.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Microsoft unveils Copilot super app with Autopilot agentsTop Companies AI · 1h ago
  • Meta's Muse agent runs on AMD EPYC Turin hosts, executes Ubuntu commandsTop Companies AI · 1h ago
  • John Deere launches JD AI assistant for farmersTop Companies AI · 1h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleSK Hynix dominates AI memory market with 56% HBM share, trades at bargain valuation