
AI pricing is no longer a mystery—the issue is that companies keep pricing based on token costs rather than customer value. Three models now work: bundling AI into existing subscription tiers, charging for specific outcomes the AI delivers, and gating usage to prevent surprise bills. The key is metering consumption accurately per customer and per task, which most companies still struggle with; those that do see much healthier margins and retention.
Summaries like this, in your inbox every morning.
Sign up free →What happened
A pricing strategist argues that AI pricing has become solvable, but the industry keeps confusing the unit cost (tokens) with the actual price. Three working models are now clear: bundling tokens into subscription plans (Notion, Cursor, Perplexity), pricing based on outcomes rather than tokens (Fin at $0.99 per resolution), and gating consumption to prevent runaway costs.
Why it matters
Most AI applications built on top of models treat tokens as a pass-through cost to customers, which caps margins at the vendor's own COGS curve—exactly when customers are paying for an outcome worth many times that cost. The piece shows AI add-ons launched much more intensely in 2025 than 2024, with 1/5th of companies adding AI features in Q2 2025, often as separate charges. Companies that switch to value-based or bundled billing can protect margin while avoiding customer churn from surprise overages.
What to watch
The five KPIs that actually matter for AI pricing: token cost as a percentage of revenue (AI gross margin rose from 41% in 2024 to 52% in 2026, with a durable floor around 60–65%), gross margin per customer (the spread matters more than the average), bundle utilization rate, consumption/overage as a percentage of revenue (rising is usually healthy), and net revenue retention (usage/hybrid models should hit 115–130%+ versus 95–105% for flat-rate models).
The article opens by challenging a widespread claim: that AI pricing remains unsolved. The author argues it is not—rather, the confusion stems from a single repeated substitution: treating tokens (the unit of inference cost) as if they were a price. Tokens are the vendor's cost to run inputs and outputs; outcomes are what customers actually want. Once that distinction is made clear, three working pricing models become visible.
The first model bundles tokens into plans or seats. Notion folded AI into higher-priced subscription tiers instead of selling it as a separate add-on. Cursor sells seats at $20 per month (Pro) or $200 per month (Ultra) with far more usage headroom. Perplexity spans $20 to $325 tiers. These companies absorb token costs and bet that adoption, retention, and expansion pay it back. The second model prices on the value tokens produce: Fin charges $0.99 per resolved customer-service ticket on top of a $49 base that includes the first 50 resolutions. Zendesk, AgentForce, Justt, and Chargeflow follow the same pattern—tokens are the vendor's invisible cost, never shown to the buyer. The third model gates usage: Cursor's tiers are gating in practice, with each seat having a usage threshold and higher tiers priced for accounts that outgrow it. The author notes this third move is often overlooked but makes the other two safe, because unbundled plans with no ceiling invite overuse.
The author cites data from partner PricingSaaS showing that AI add-on launches were far more intense in 2025 than in 2024, with 1/5th of companies adding AI in Q2 2025 as separate charges. Solvimon's customers bill on average about 5 different items now, up from roughly 2.5 before. When done correctly, add-ons with gating protect both vendor and buyer: the vendor holds margin tight, and the customer gets a voluntary credit balance to watch and budget against. Done badly, they become invisible traps that drive churn.
The article then introduces five KPIs for AI pricing. Token cost as a percentage of revenue maps inversely to gross margin; AI-native gross margins rose from about 41% in 2024 to 52% in 2026, with a durable floor around 60–65%. Gross margin per customer is more important than the blended average—top 5% of users are either best case studies or margin losers. Bundle utilization rate tracks credit breakage and overrun. Consumption or overage as a percentage of revenue often rises healthily with expansion. Net revenue retention shows whether the model works longer-term: usage or hybrid models should hit 115–130%+ versus 95–105% for flat-rate models. The author emphasizes that per-customer metering—attributing token and infrastructure cost to each account—is essential; blended averages hide loss-making whales, often the accounts companies are proudest of.
The article reframes a persistent industry confusion: tokens are a unit of cost to run a model, not a price to charge customers. This distinction matters because output tokens run 3–8x the cost of input tokens, and the actual outcome a customer pays for is unmeasured in tokens at all. For foundation-model APIs selling access directly, token pricing works fine; but for applications layered on top—customer service bots, AI writing tools, research assistants—passing through token costs at a markup leaves margin on the table and frustrates buyers who don't understand their bills.
Three strategies have emerged as working: bundling (Notion, Cursor, Perplexity absorb token costs into higher-priced subscription tiers and bet on retention and expansion), outcome-based pricing (companies like Fin charge per resolved ticket, making the invoice match the customer's actual value), and gating (usage thresholds and tiers prevent surprise overages that cause churn). The author notes that 1/5th of companies added AI features as separate add-ons in Q2 2025, up sharply from 2024, suggesting the market is still testing pricing shapes—and that add-ons done badly become invisible cost traps. The real unsolved problem isn't strategy; it's the plumbing: accurate per-customer metering and clear gating rules that let companies charge for bundled or outcome-based plans without leaving margin on the table or losing customers to surprise bills.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime