AIToday

Why cheaper AI tokens won't solve enterprise cost crisis

Top Companies AI — US (2/2)13h ago
Why cheaper AI tokens won't solve enterprise cost crisis

Key takeaway

Enterprise spending on AI has shifted from maximizing use cases to controlling costs, but cheaper tokens alone won't solve the problem. Most companies waste tokens through inefficient architecture—processing thousands of daily agent interactions that consume 5–10x more tokens than necessary. The real advantage goes to companies that invest in data governance and master data management to extract more value from fewer tokens, not those that simply buy cheaper compute.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Enterprise AI spending has shifted from maximizing use cases to controlling costs. Companies like Uber have tightened AI spending, while AI providers have cut token prices and rolled out caching. However, a Salesforce VP argues that lower costs alone miss the real problem: most companies waste tokens through poor architecture.

  • Why it matters

    Enterprises process thousands of daily AI agent interactions, each consuming 5–10x more tokens than necessary due to unoptimized data pipelines. At $10–15 per million tokens, this waste compounds quickly. Companies that invest in data governance and master data management gain a structural cost advantage—not by buying cheaper tokens, but by extracting more value from fewer of them.

  • What to watch

    Token optimization now requires five architectural fixes: eliminating context bloat through master data management, establishing data governance as a routing signal, cost-routing work to the right model rather than always using frontier models, adding persistent memory to reduce per-interaction reprocessing, and caching repeated requests. The next competitive edge in enterprise AI belongs to companies that turn AI consumption into governed, measurable business value.

In Depth

Over the past two years, enterprises embraced a maximalist approach to AI adoption, launching multiple use cases, onboarding many employees, and deploying numerous agents. This strategy was rational for a new technology: companies did not yet know which workflows would be transformed, which models were reliable enough, or which employees would actually adopt the tools. Rapid deployment and pilots were the fastest way to find out what worked. But the conversation has changed dramatically this year. In nearly every executive review, the question is no longer "How do we get more AI?" or "Which model is smartest?" It has become "How do I generate impactful outcomes while keeping AI costs under control?" Different companies are taking different paths. Uber tightened internal AI spending after consumption ran far ahead of plan. Anthropic, OpenAI, and Google have aggressively cut input token prices and rolled out prompt caching. Cursor, with Composer 2.5, now considers cost an important factor in model selection, not just performance.

However, the article argues that cutting token and model costs addresses only a symptom. The deeper issue is that most companies are not architected to use tokens efficiently. The modern AI pipeline leaks tokens and dollars at every phase of execution. According to a Salesforce VP, pouring cheaper tokens through a leaky foundation is not sustainable; companies need an architectural foundation that closes those leaks. The leaky pipeline has five predictable failure points. Context bloat occurs when too much information is stuffed into a prompt, causing higher costs and model drift. The fix comes through master data management—ensuring agents have access to the right customer, account, policy, transaction, and recent activity built into infrastructure rather than passed through the prompt. Ungoverned data access forces agents to roam aimlessly through data warehouses searching for credible information, wasting time and tokens. Establishing a data catalog, lineage, permissions, and quality signals turns governance into a routing signal rather than a checkpoint. Wrong model selection means sending every task to a frontier (most advanced) model instead of cost-routing work to the right model for each task—treating models as a portfolio where small, efficient models handle lookups and classification, tuned mid-tier models handle domain workflows, and frontier models are reserved for novel reasoning and sensitive judgment. Lack of persistent memory forces stateless agents to start each interaction from scratch, reloading context and reprocessing history. Structured memory lets agents keep and retrieve relevant state—facts, decisions, preferences, and open tasks—reducing token consumption and cost per interaction. Single-use semantics means enterprise agents process recurring requests as if seeing them for the first time, re-reading context, re-running retrieval, and re-generating answers that could have been cached, embedded, or pre-computed.

The business impact is substantial. A large enterprise running AI agents for sales or service operations might process thousands of agent interactions daily. Each interaction that passes raw, uncurated data into a language model consumes 5–10x more tokens than necessary. At $10–15 per million tokens, this waste compounds quickly across a fleet of agents. Enterprises that invest in master data management and data governance create a natural advantage by drawing on pre-curated, semantically enriched, quality-scored data rather than raw noise. The Salesforce VP states from firsthand work with enterprise customers that when Informatica's master data management integrates with Data 360, Salesforce's Customer Data Platform, every AI agent is grounded in trusted customer and business context, turning ungoverned AI consumption into governed, measurable business value. The conclusion is that every company can buy more tokens, but very few know how to extract more value from fewer of them. The next phase of enterprise AI will not be defined by who consumes the most tokens, runs the largest models, or fills the biggest context windows, but by who can turn AI consumption into governed, measurable business value.

Context & Analysis

The article captures a fundamental shift in enterprise AI strategy. Two years of aggressive AI rollouts—maximizing agents, pilots, and token consumption—have given way to cost discipline. The trigger is clear: spending has run ahead of measurable business value. Vendors have responded with price cuts (OpenAI, Anthropic, Google) and efficiency features (prompt caching), but the Salesforce VP's argument cuts deeper. Cheaper tokens are a surface fix; the real leak is architectural. Enterprises that lack master data management, data governance, and memory structures force every AI interaction to start from scratch, regenerate the same queries, and process raw noise at the model level. This is not a vendor problem—it's an organizational design problem that only the enterprise can fix by investing in foundational data infrastructure before building more agents. The implication is that token costs will remain a competitive advantage for companies that have already invested in data governance, while those without it face a structurally higher cost per interaction, regardless of list price.

FAQ

How much extra do inefficient AI pipelines cost?
Each agent interaction that passes raw, uncurated data into an AI model consumes 5–10x more tokens than necessary. At $10–15 per million tokens, this multiplies quickly across a fleet of agents in large enterprises processing thousands of interactions daily.
What are the five main sources of wasted tokens?
Context bloat (stuffing too much information into prompts), ungoverned data access (agents searching aimlessly through warehouses), wrong model selection (sending every task to expensive frontier models), lack of persistent memory (reprocessing history on every interaction), and single-use semantics (re-running the same retrieval and regenerating answers that could be cached).
How do companies fix token waste?
The Salesforce VP recommends master data management to provide curated, single-source-of-truth data before the prompt stage, establishing data governance as a routing signal, treating models as a portfolio (small models for lookups, frontier only for novel reasoning), adding structured memory so agents retrieve relevant prior context, and caching repeated requests through prompt caching and embedding reuse.

Get the latest Top Companies' AI Moves news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →