AIToday
Large Language ModelsAI Business & IndustryITmedia AI+Published: Oct 6, 2026, 06:01 JST

Token prices fell 8割 but AI bills rose, Postman writer says

Token prices fell 8割 but AI bills rose, Postman writer says

3 Key Points

  1. What happened

    Postman's Akihiko Kusanagi writes that LLM token prices fell roughly 8割 between 2025 and 2026, yet AI bills rose as spending on LLM APIs topped 84億ドル in 2025.

  2. Why it matters

    Per Anthropic research cited in the piece, agentic workflows consume about 4x the tokens of chat and multi-agent setups about 15x, so cheaper units do not guarantee smaller invoices.

  3. What to watch

    Kusanagi says the fix is governance and API design, not price comparison — and points to a next installment with five ways to curb token consumption.

WHO IT HITSThis lands hardest on platform and API teams that expose internal services to AI agents, plus the finance and IT managers who now track AI spend as its own budget line. Teams wrapping legacy APIs into agent tools may already be paying for context they never intended to send.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The article builds on the two earlier installments in this series, which examined API keys and access tokens — the credentials agents use to reach outside services — and the risks of handing agents a human's keys. Kusanagi frames credential management as only the entry point to running agents, with cost as the equally overlooked second problem that shows up later.

The mismatch between unit price and total bill has drawn investor attention too. Andreessen Horowitz has analyzed that the token price for the same capability keeps falling at roughly a tenth per year, and the piece notes that GPT-3-class performance dropped from about $60 per million tokens in 2021 to about $0.06 in late 2024 — a nearly thousandfold decline in three years.

The piece also connects tool design directly to cost. Wrapping existing REST APIs one-to-one into MCP tools is easy and works at first, but mechanically generating a server from an OpenAPI spec turns 200 endpoints into 200 tools, inflating context and causing selection errors. A cited study found that narrowing tools to only what is needed raised tool-selection accuracy more than threefold and cut prompt tokens by more than half. Whether enterprises can bend their spending curve downward therefore appears to hinge less on model vendors' price lists than on how carefully they curate the tools and context they hand to agents.

FAQ
Why are AI bills rising if token prices keep falling?
Because usage grows faster than price falls. Each agent task can call a model dozens of times or nearly 200 times, and Anthropic research cited in the article says agentic workflows use about 4x the tokens of ordinary chat, or 15x for multi-agent setups.
How much context do tool definitions eat up?
GitHub's official MCP server consumes about 4万2000 tokens just listing tool definitions. Linking four or five such servers can exceed 6万トークン, eating 3 to 5割 of a frontier model's 12万8000〜20万トークン context window before any work is done.
Is anyone treating AI costs as a separate budget item?
Yes. A FinOps Foundation survey found organizations separating AI spend as an independent managed category rose from 31% in 2024 to 63% in 2025, and reached 98% in 2026.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAnthropic, Microsoft leaders drive AI consciousness debate