
Enterprise AI budgets are shifting from "how do we deploy more AI" to "how do we control costs while driving outcomes." The real problem is not token price but pipeline inefficiency—companies leak 5–10× more tokens than necessary by sending raw data into prompts instead of building governed data foundations. Enterprises that invest in master data management and data governance will extract more value from fewer tokens, becoming the competitive winners of the next phase of enterprise AI.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Enterprises' AI spending has shifted from maximalist deployment toward cost control. Uber tightened spending after consumption outpaced plans; Anthropic, OpenAI, and Google cut token prices and added caching; Cursor's Composer 2.5 now weights cost alongside performance in model selection.
Why it matters
Cheaper tokens alone do not solve the problem—most companies waste 5–10× more tokens than necessary by feeding raw, uncurated data into AI prompts instead of building governed data foundations. A large enterprise running thousands of daily AI agent interactions compounds this inefficiency at $10–15 per million tokens, making architecture and data governance the real lever for cost-effective AI.
What to watch
The competitive advantage will go to enterprises that invest in master data management and data governance (integrating systems like Informatica and Salesforce's Customer Data Platform) to deliver curated, trusted data to AI agents, rather than those that simply consume more tokens or use larger models.
Over the past two years, enterprises adopted AI aggressively and without strict cost discipline. The logic was sound: with a new technology, the fastest way to determine which workflows would transform and which employees would adopt the tool was to deploy as many agents and run as many pilots as possible. But the conversation shifted this year. In executive reviews across industries, the question is no longer "How do we get more AI?" or "Which model is smartest?" It has become "How do I generate impactful outcomes, while keeping AI costs under control?" Different companies have responded differently. Uber tightened internal AI spending after consumption ran far ahead of plan. Anthropic, OpenAI, and Google have aggressively cut input token prices and rolled out prompt caching. Cursor, with Composer 2.5, now considers cost an important factor in model selection—not just performance. Yet according to the analysis, cutting token and model costs alone misses the deeper issue: most companies are not architected to use tokens efficiently. The modern AI pipeline functions like a sieve, leaking tokens and dollars at every phase of execution. The article identifies five predictable failure points. First, context bloat: stuffing too much raw information into a prompt inflates costs and degrades model reasoning. The fix is a data foundation that provides agents with accurate, unified data before the prompt stage—through master data management using deterministic matching, merging, and deduplication to create a single golden record. Second, ungoverned data access: without clear governance, agents search aimlessly through data warehouses and check results after the fact, wasting tokens and generating inconsistent actions. Establishing a data catalog, lineage, permissions, and quality signals turns governance into a routing signal on the first hop, not a checkpoint. Third, wrong model selection: sending every task to a frontier model is expensive. Cost-routing work through the right model—small and efficient for lookups and classification, tuned mid-tier models for domain workflows, and frontier models reserved for novel reasoning—increases efficiency dramatically. Fourth, lack of persistent memory: stateless agents start each interaction from scratch, reloading context and reprocessing history. Structured memory lets agents keep and retrieve relevant facts, prior decisions, user preferences, and open tasks, reducing token consumption and lowering cost per interaction. Fifth, single-use semantics: enterprise agents often reprocess recurring requests as if seeing them for the first time, re-reading context, re-running retrieval, and regenerating cached answers. Pairing prompt caching, embedding reuse, and pre-computed outputs with semantic structure already present in data flows makes repeat requests progressively cheaper and more reliable. The economic impact is substantial. A large enterprise running AI agents for sales or service operations might process thousands of agent interactions daily; each interaction that passes raw, uncurated data into a prompt consumes 5–10× more tokens than necessary. At $10–15 per million tokens, this inefficiency compounds fast across a fleet of agents. Enterprises that invest in master data management and data governance create a natural advantage by drawing on pre-curated, semantically enriched, quality-scored data. When systems like Informatica's master data management integrate with Salesforce's Customer Data Platform, for example, every AI agent is grounded in trusted customer and business context, turning ungoverned consumption into governed, measurable business value. The conclusion is that token optimization will define the next competitive phase of enterprise AI. Every company can buy more tokens; very few know how to extract more value from fewer of them. The winner will not be defined by who consumes the most tokens, runs the largest models, or fills the biggest context windows, but by who can turn AI consumption into governed, measurable business value.
The enterprise AI conversation has matured rapidly in two years. What began as a maximalist experimentation phase—deploying as many pilots and agents as possible to find which workflows would transform—has collided with fiscal reality. Uber's spending tightening, aggressive token-price cuts by Anthropic, OpenAI, and Google, and Cursor's decision to factor cost into model selection all signal that the industry has moved from "How do we adopt AI?" to "How do we derive ROI from AI while controlling burn?" Yet the article argues that cheaper tokens misdiagnose the real problem. The core issue is architectural: enterprises lack the data governance and master data management infrastructure to feed clean, curated context to AI agents. Instead, they send raw, noisy data through prompts—a practice that multiplies token consumption by 5–10×. This inefficiency is baked into the pipeline itself: context windows bloat, agents wander through ungoverned data, large models handle simple lookups, stateless interactions require reprocessing, and semantic duplicates regenerate answers unnecessarily. The implication is that the next phase of competitive advantage in enterprise AI will not belong to those who negotiate cheaper tokens or deploy the largest models, but to those who invest in the unsexy but essential work of master data management, data governance, and semantic architecture—turning raw AI consumption into measurable, governed business value.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No discussion yet for this article
Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack