AIToday
Large Language ModelsAI Business & IndustrySiliconANGLE AIPublished: Sep 29, 2026, 06:00 JST

Agentic AI uses 10 to 100 times more tokens per task: Futurum

Agentic AI uses 10 to 100 times more tokens per task: Futurum

3 Key Points

  1. What happened

    Futurum's report, sponsored by QumulusAI Inc., finds agentic AI can drive token consumption per task 10 to 100 times higher than a simple inference call. It forecasts agent and reasoning inference will grow 219% this year, and total inference spending will rise from $120 billion in 2025 to $885 billion by 2030.

  2. Why it matters

    The report quantifies a real cost problem that chief information officers and chief financial officers have reported with increasing frequency, according to the article.

  3. What to watch

    The risk isn't just a high bill; unpredictable bills can kill useful projects, and some customers abandoned internally built automation tools because they couldn't forecast or justify the costs, per the report. Futurum recommends reserved bare metal for sustained workloads with predictable utilization above roughly 60%.

WHO IT HITSEnterprise IT and finance teams running production AI workloads are most affected, since agentic AI can use 10 to 100 times more tokens per task than a simple inference call. The report warns that unpredictable bills can cause organizations to abandon successful projects, so IT and finance leaders may need to measure cost per task rather than cost per token.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The Futurum report, "The Off Ramp From Per-Token Pricing," was sponsored by neocloud provider QumulusAI Inc. and arrives as enterprises are already shifting how they buy AI compute. According to Futurum's survey of 824 AI decision-makers, reserved and owned infrastructure account for 66% of AI compute consumption, compared with 19% for on-demand cloud, and 59% of respondents primarily run AI workloads outside hyperscaler public clouds, in their own data centers, colocation facilities, or with bare-metal providers.

The report also notes that the offramp works well for open-weight models a company can deploy on infrastructure it controls, which is why Amberd.ai built on private, open-source large language models. But many enterprises have standardized on frontier models available only through their developers' application programming interfaces or hyperscaler marketplaces. For those workloads, no bare-metal alternative exists, so the pricing lever rests with the model provider.

Whether the offramp pays off may hinge on operational skills, since reserved infrastructure only saves money if it stays busy. The report acknowledges that these environments require more custom engineering, limiting the operating margin gains for teams without the hardware expertise. Enterprises that lack deep bench strength in serving engines, batching, quantization and key-value cache management may find the offramp leads to a different kind of cost overrun.

FAQ
What is the key finding of the Futurum report?
The report's key finding is that agentic AI can drive token consumption per task 10 to 100 times higher than a simple inference call.
What does the report recommend for workload placement?
Futurum recommends reserved bare metal for sustained workloads with predictable utilization above roughly 60%. The report also acknowledges these environments require more custom engineering, limiting operating margin gains for teams without the hardware expertise.
How much did one organization spend on AI?
One organization budgeted $1 million for the year, but the initiative was so successful they spent it in three months.
SiliconANGLE AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Meta taps CJ Desai to lead new enterprise AI unitSiliconANGLE AI · 20m ago
  • 1Password ties AI agent access to individual tasksSiliconANGLE AI · 20m ago
  • ServiceNow's Bhakti Pitre: rogue AI agents need risk-based kill switchSiliconANGLE AI · 20m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleClaude Code's Boris Cherny: Black-box AI code OK only for prototypes