AIToday
Large Language ModelsZenn AI/MLPublished: Sep 28, 2026, 22:00 JST

Agent cost per successful task: a Zenn design-variable argument

Agent cost per successful task: a Zenn design-variable argument

3 Key Points

  1. What happened

    A Zenn article narrowed agent cost design to three topics: cache depends on prefix stability, routing should be judged by break-even, and budgets must be enforced in the application layer, not billing systems.

  2. Why it matters

    Designing around cost per successful task means failure costs are carried by successes, so the Evals success rate directly affects the cost metric.

  3. What to watch

    The evaluation hinges on whether prefix stability, break-even, and app-layer budget enforcement hold in practice; the body notes the 2026年9月時点 figures are official-price calculations, not measured results.

WHO IT HITSTeams building AI agents and managing LLM budgets will need to treat caching, model routing, and budget enforcement as design-time choices; the article notes figures are official-price calculations, not actual measurements.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The article positions agent cost as something decided at design time, not something fixed by shortening prompts. It focuses on three points: cache behavior determined by prefix stability (tools → system → messages), routing judged by a break-even formula, and budget enforcement in the application layer because AWS Budgets updates at most three times a day with 8–12 hour intervals. It also notes that input_tokens (inputTokens on Bedrock) only counts the portion not subject to caching, so total input must be aggregated with cache read and write. A concrete calculation for a 10-turn task on Sonnet 5 gives 0.146 for hits and $0.519 for constant misses. The thought is that once hits are working, output accounts for about one-third of cost, making the output side the next lever. One notable risk raised is that repeatedly missing with delimiters in place means paying 1.25x the write unit price continuously, which can end up costing more than no caching at all.

FAQ
What is the recommended cost management unit for agents?
The article uses cost per successful task, calculated as LLM cost during a period divided by the number of successful tasks. Because failure costs are carried by successes, the Evals success rate directly affects the cost metric.
What is the processing order for caching and what breaks it?
The order is tools → system → messages, and if an earlier layer changes, all later layers are invalidated. Non-deterministic tool order, system-leading timestamps, and mid-run changes to thinking or tool_choice are all invalidation factors.
Why can routing between models be worse than expected?
As of 2026年9月, the price gap between Sonnet 5 and Haiku 4.5 is only 2x, and Haiku 4.5 has a minimum cache length of 4,096 tokens. For short prefixes, Sonnet 5 cache reads can be cheaper, reversing the unit-price relationship.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Autoheal raises $7.9 million for self-fixing AI agentsSiliconANGLE AI · 43m ago
  • Paul Cheek: 30% of S&P 500 execs AI-literate, 78% gapFortune AI · 43m ago
  • Claude Fable 5.1 builds matrix-free Transformer site, then a 10M Japanese SLMZenn AI/ML · 43m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI Knew Pirated Books Risked Authors' Livelihoods