AIToday
AI Coding AssistantsZenn AI/MLPublished: Sep 30, 2026, 10:00 JST

Claude Code subagents silently inherit opus; echo task hits 101,418 tokens

Claude Code subagents silently inherit opus; echo task hits 101,418 tokens

3 Key Points

  1. What happened

    A freelance engineer measured his Claude Code setup and found 4 of 6 subagent definitions left the model field blank, inheriting the parent's opus model; a subagent that only echoed 'X' once used 101,418 tokens, about 90,000 of it CLAUDE.md.

  2. Why it matters

    The cost of delegating came mostly from the shared context every subagent reloads, not from the task itself — and the improvement loop that should have run every 4 hours had already missed 8 straight cycles on usage_limit.

  3. What to watch

    Whether specifying a model per role actually lowers usage limits hinges on whether the parent and subagent contexts are truly separated; the engineer measured total consumption as roughly round trips × (CLAUDE.md + foundation + work content), and three of Claude's call paths remained unspecified at the start.

WHO IT HITSEngineers running Claude Code on a flat-rate plan are the ones affected: leaving model fields blank in .claude/agents/*.md lets every subagent run on the parent's most expensive model, so the usage cap arrives before the roles that genuinely need opus get their turn.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The engineer runs a blog and an X account on Claude Code from GitHub Actions on a VPS, calling claude -p on a schedule. The audit began when a review pointed out that his model selection was not actually differentiated. He first counted every path that calls Claude — an improvement loop every 4 hours, article writing twice a week, research, chat, and a smoke test — and found three of them left the model unspecified. The subagent definitions had the same problem: of six, only two named a model.

The immediate symptom was practical rather than theoretical. Records from the improvement loop's stop-reason file showed eight consecutive runs from September 20 and 21 all hitting usage_limit, leaving the loop stuck for two days. The engineer is careful not to assert that opus alone caused this, since the flat-rate quota is shared with human sessions and other paths — but he notes that running cheap roles on an expensive model does bring the cap closer before the roles that truly need opus get their turn.

He also documents a hypothesis he nearly reported without testing: that Task was missing from --allowedTools and subagents were structurally impossible to call. A measurement showed subagents spawned and completed with zero permission denials, so the real problem was that nobody was calling them, not that they couldn't be called. The stakes in the rest of the piece flow from that distinction — the fix is a prompt change, not a permissions redesign — and from his choice to keep article writing, public post writing, and the smoke test on opus while pushing no-judgment roles down to haiku.

FAQ
What happens if I do not write a model in a Claude Code subagent definition?
It inherits the parent's model. The engineer verified this with the modelUsage field in `claude -p --output-format json`, and found in his own environment that a subagent used only for copying command output was running on opus, the most expensive model.
Does adding more subagents save tokens?
No — subagents also read CLAUDE.md in full. In the engineer's measurement, a subagent that only ran echo X once used 101,418 tokens, of which about 90,000 was the CLAUDE.md portion. Delegation can lower usage (the flat-rate cap) when it swaps an opus round trip for a cheaper model, but token count itself increases.
Why not lower the smoke test to a cheaper model?
Because a green light from a cheaper model can be false. Usage limits apply per model, so haiku answering 17×23 = 391 correctly does not prove the improvement loop's opus is still under its cap — the engineer explicitly kept the smoke test on opus, the same model production actually uses.

AI news that matters for your work, in one minute a day

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleLessWrong post argues against strong decision-theoretic realism