
What happened
On February 6, 2026, agents on OpenRouter consumed more tokens than humans did and have not given the lead back, according to the article.
Why it matters
Agent volume on OpenRouter rose from 0.51t to 7.3t tokens in six months, a 14x increase, while human volume only managed 2.8x.
What to watch
The third wave — meta-harnesses dispatching many agents in parallel — is on the horizon and could push token counts into the billions, hinging on how quickly parallelization compounds.
WHO IT HITSEnterprise AI platform teams and compute buyers who plan capacity based on smooth chat-focused growth curves may need to rethink their forecasts, as agent and meta-harness consumption stacks on top of chat rather than replacing it.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
AI token consumption does not follow a smooth curve. The article traces three waves: chat, a single agent, and a meta-harness orchestrating many agents. The second wave has already broken, with agents on OpenRouter consuming more tokens than humans since February 6, 2026. This shift is visible in enterprise use too: by June 2026, chat accounted for a little more than a third of enterprise output, with Codex responsible for the other 64%. The token intensity gap between waves is two zeros of power.
The growth comes not from faster generation but from parallelization compounding. A meta-harness, for example, could involve a spreadsheet of AI answers where each cell represents tens of tool calls, costing tens of millions of tokens before anyone notices. Agent volume on OpenRouter rose from 0.51t to 7.3t tokens in six months, a 14x increase, while human volume managed 2.8x. Goldman Sachs projects this demand will reach 120 quadrillion tokens a month by 2030, 24 times the 2026 level.
The stakes hinge on how compute planning adapts. Anyone sizing infrastructure off a smooth extrapolation may be planning for the wrong curve, as the waves stack rather than replace each other. Chat keeps growing, agents grow faster, and meta-harnesses could dwarf both, with the outcome depending on whether parallelization continues to compound at this pace.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…
A Daily Dose of Data Science test kept LoRA adapters separate from a shared 7B base model, cutting 100 fine-tu…

A report by Spencer Kitts, Thomas Larsen and Sydney Von Arx says an OpenAI agent swarm very likely ran an atta…

Simon Willison wrote that many people, himself included, have gone through an existential crisis when a coding…

Stephen Aarons, a New Mexico defense lawyer of over 40 years, was held in direct contempt and fined $5,000 for…

Perplexity cofounder and Chief Strategy Officer Johnny Ho said GPT‑6 Astra can craft communications, edit real…
