AIToday
AI Coding AssistantsAI Safety & AlignmentThe Register (AI/ML)Published: Aug 7, 2026, 04:01 JST

Humans miss a third of dangerous AI coding agent requests, study finds

Humans miss a third of dangerous AI coding agent requests, study finds

3 Key Points

  1. What happened

    A browser-based game testing human approval of AI coding agent requests found that players approved roughly one in three malicious commands on average, with scope violations like exposing AWS credentials or Kubernetes config files missed about 35 percent of the time. The most-missed command was npm run analyze, approved nearly 65 percent of the time despite the ability to run any payload defined in a project's package.json file.

  2. Why it matters

    Anthropic's telemetry shows users approve around 93 percent of permission prompts from Claude Code, meaning repeated approval requests lead to fatigue and inattention. As developers transition from reviewing single-line suggestions to approving complex multi-step agent workflows, the risk of malicious commands slipping through grows — a problem compounded when context for assessing risk is limited or requires time-consuming investigation.

  3. What to watch

    Anthropic built Claude Code auto mode to delegate some command-approval decisions to a model-based classifier, which catches roughly 83 percent of what it calls "overeager behaviors" before execution, leaving about 17 percent still getting through. Developers are advised to run AI coding agents in sandboxes, use devcontainers in the cloud, and implement hooks to contextualize and catch potentially malicious actions before automatic approval.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The game designed by Belgian software developer Alex Wauters highlights a fundamental tension in the design of AI coding agents: either users approve every command manually, or they delegate approval entirely to the model. Wauters observed that many developers resort to flags like '--dangerously-skip-permissions' to avoid interrupting multi-hour agent workflows, but the manual-approval path introduces a different failure mode. Repeated approval requests cause fatigue, which the game's data — drawn from over 40,000 runs and 409,000 approved and denied commands — quantifies starkly: roughly one in three malicious requests slip past human gatekeepers.

AnthropicConfirms this is a real-world problem. The company's telemetry from Claude Code shows users approve around 93 percent of permission prompts, corroborating the game's finding that attention degrades with volume. The specific commands that slip through reveal the nature of the oversight challenge: scope violations (requests to access credentials or configuration files) are missed about 35 percent of the time, and npm run analyze — a command that can execute arbitrary code defined in package.json — passes approval nearly 65 percent of the time. The latter suggests that even when context is provided (the history log showing what the script contains), developers often do not invest the time to review it thoroughly.

Wauters argues that the solution is not to improve human judgment under fatigue but to reduce the burden through defense-in-depth: sandboxed execution, cloud-based devcontainers, and automated safeguards like Anthropic's Claude Code auto mode. However, auto mode itself is imperfect; it catches roughly 83 percent of problematic behaviors, leaving roughly 17 percent to escape initial detection. The implication is that as developers hand off more complex tasks to AI agents and rely on them to iterate until completion, security must be enforced at the infrastructure level rather than at the approval prompt.

FAQ
What was the most-missed dangerous command in the study?
npm run analyze was the single most frequently missed potentially malicious command, approved nearly 65 percent of the time. Although the game showed what the script actually contained in the agent's history log, two thirds of players approved it anyway, suggesting the history log above the permission prompt was not read closely.
How effective is Anthropic's Claude Code auto mode at blocking malicious commands?
Claude Code auto mode catches roughly 83 percent of what Anthropic calls "overeager behaviors" before execution, meaning about 17 percent still get through in its evaluation. Anthropic describes auto mode as "one layer of defense-in-depth inside a sandbox, not a substitute for one."
What percentage of Claude Code permission prompts do users typically approve?
Anthropic's telemetry shows users approve around 93 percent of permission prompts from Claude Code, indicating that developers become less diligent in their supervision the more approvals they see.
The Register (AI/ML)Read Original Article

Get the latest AI Coding Assistants news every morning

For example, today's edition would include:

  • Claude Code, Cursor, Copilot: teams run 2–3, not oneZenn AI/ML · 13m ago
  • 5 Claude Code sessions, 1 git tree: 8 accidents in 3 monthsZenn AI/ML · 13m ago
  • panda studio runs a store with Claude Code, sales up 約3.5倍Zenn AI/ML · 13m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAMD, Intel Negotiate Longer China CPU Deals as AI Workloads Drive Demand