
As AI agents consume vastly more tokens than simple chatbots—up to 200,000 to a million per complex task—enterprises face unexpected scaling costs. Goldman Sachs forecasts token consumption will multiply roughly 24 times by 2030 to 120 quadrillion a month. Dell's new Deskside Agentic AI system, launched in May, lets companies run agents on their own workstations instead of cloud APIs, cutting token costs by up to 87% over two years and reaching break-even in as little as three months.
Summaries like this, in your inbox every morning.
Sign up free →What happened
A simple AI agent consumes up to 15,000 tokens per task; complex multi-agent systems use 200,000 to over a million. Goldman Sachs projects total token consumption will multiply roughly 24 times by 2030, to 120 quadrillion a month. Dell launched Deskside Agentic AI in May, a system that runs production-ready agents on company workstations using open-source models, with governance built in from the start.
Why it matters
Between mid-2023 and early 2026, token prices fell 80%, but enterprise AI spending jumped 320% because companies deployed far more agents consuming vastly more tokens. The total bill climbed despite lower per-token prices. For IT teams, the shift means infrastructure, security, budgeting, and governance all have to change — and token strategy is now a board-level question, not just an IT decision.
What to watch
Analysis by Signal65 and Futurum shows running agents on-premises saves up to 87% on token spend over two years compared to public-cloud APIs, with break-even in as little as three months. The system handles workflows for coding, research, and private assistants on models from 30 billion to trillion parameters and can migrate to data center servers without redesign.
Every interaction with an AI model consumes tokens—small units of text, each representing roughly three-quarters of a word. The scale of consumption has become staggering as enterprises deploy agentic systems. While a chatbot exchange burns a few hundred tokens, a simple agent consumes up to 15,000 tokens per task, and a complex, multi-agent system uses from 200,000 to more than a million. Stanford researchers discovered that agentic coding is uniquely token-dense, consuming around 1,000 times more tokens than ordinary code chat. Goldman Sachs expects this trend to accelerate, projecting total token consumption will multiply roughly 24 times by 2030, reaching 120 quadrillion a month.
The paradox driving this change is straightforward: token prices fell by around 80% between mid-2023 and early 2026, yet companies saw their AI spending increase by roughly 320% in the same period. Rather than banking savings from cheaper tokens, enterprises deployed agents consuming far more tokens and ran more of them across more workflows. The result was that total token costs climbed despite the lower per-token price. As Bharat Patel, a solution architect at Dell Technologies Customer Solution Center, explains, "Where you run AI matters as much as which AI you run." This shift demands changes across the entire supporting system—hardware, security, budgeting, and management tools—and turns token strategy into a board-level question rather than an IT-only concern.
Dell Deskside Agentic AI, launched in May, addresses this challenge by enabling work groups to run production-ready agents on Dell workstations using the Nvidia NemoClaw open-source stack. The system handles validated workflows for coding, research, and private assistants across models ranging from 30 billion to trillion parameters. Agents operate inside an OpenShell environment that enforces policy-based privacy and security rules and logs all actions, establishing governance from the start. A key advantage is portability: when a prototype scales beyond desktop capacity, it can migrate to Dell PowerEdge servers in the data center without architectural redesign. This approach is particularly valuable for organizations that cannot send sensitive work to public cloud APIs—such as engineers protecting proprietary source code or researchers handling pre-publication and patient data subject to privacy regulations.
Analysis by Signal65 and Futurum, commissioned by Dell, shows the cost impact is substantial: running agentic workflows on-premises delivers savings of up to 87% on token spend over two years compared to public-cloud APIs, with break-even possible in as little as three months. Patel frames his guidance as six words: "Start local, govern early, scale smart." The implication is that token costs are a design decision, not an afterthought—decided before the first agent ships, with clear governance around where each workload runs. When that discipline is in place, the bill becomes something teams manage rather than something that happens to them.
The economics of agentic AI have shifted fundamentally. When token prices dropped 80% between mid-2023 and early 2026, enterprises did not pocket those savings; instead, they deployed far more agents across more workflows, driving enterprise AI spending up 320% in the same period. This disconnect reveals a deeper truth: the marginal cost of individual tokens matters less than the total volume of tokens an agentic system consumes. Goldman Sachs' projection of a roughly 24-fold increase in token consumption by 2030 to 120 quadrillion a month underscores that the bill is driven by the scale and complexity of agent workflows, not token pricing alone.
The implications reshape how enterprises should think about infrastructure. Running agents on public cloud APIs—treating frontier models as the default workhorse—locks companies into variable consumption costs that compound quickly as agent deployments grow. Dell's approach, articulated by Bharat Patel, inverts this logic: run everyday tasks on hardware you already own, where the marginal cost of a token approaches the price of electricity, and reserve expensive cloud APIs for the largest frontier models that require specialist capabilities. This design decision, made early and with governance in place, transforms the token bill from an uncontrolled variable into a managed cost.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime