AIToday
Large Language ModelsAI Business & IndustryTomasz Tunguz (Theory Ventures)Published: Aug 5, 2026, 06:01 JST5 min read

AI labs segment pricing to sustain demand amid capacity shortage

AI labs segment pricing to sustain demand amid capacity shortage

Key takeaway

  • Hyperscalers announced in Q2 2026 earnings that AI capacity will remain tight through 2027 and beyond, with demand already booked for 2028 outpacing supply.

  • Rather than cut prices uniformly, AI labs are introducing tiered models—premium, mid-market, and value—at sharply different price points to absorb workloads and sustain overall consumption growth.

  • Router software that directs queries to the appropriate model tier is emerging as the key competitive battleground, forcing major labs to maintain offerings across all price segments or risk losing customers to more specialized competitors.

3 Key Points

  1. What happened

    Hyperscalers including Alphabet, Amazon, and NVIDIA stated on Q2 2026 earnings calls that AI capacity remains constrained through 2026 and into 2027, with demand already booked for 2028 outpacing supply. Simultaneously, AI model pricing is fragmenting into three tiers—premium, mid-market, and value—rather than falling uniformly. Anthropic launched Fable 5 at $50 per million output tokens (doubling Opus 5), Google's Gemini flagship rose from $1.50 to $12 across generations, yet OpenAI cut GPT-5.6 Luna prices by 80% five days after Fable 5's launch.

  2. Why it matters

    Rather than compete on price alone, AI labs are betting that tiered offerings—the mid-market tier delivering 96% of frontier intelligence at 40% of the cost, and the value tier offering 84% of frontier capability at 1 to 5% of premium cost—will absorb workloads priced out of premium services and keep total GPU consumption growing. This segmentation strategy is meant to sustain Jevons' Paradox (where lower cost drives higher consumption) even as headline prices rise, preventing demand from collapsing under supply constraints.

  3. What to watch

    Routers—software that directs queries to the right-priced model—are becoming the strategic chokepoint between buyers and models. Each major lab needs entries across all three pricing tiers or risks losing router auctions. Startups can win by owning a single frontier point (as DeepSeek V4 Flash does at $0.03 per million tokens) that large labs cannot economically match.

In Depth

Read the full story

In Q2 2026 earnings calls, major cloud and AI companies acknowledged that AI compute capacity remains the binding constraint on growth. Sundar Pichai of Alphabet stated simply, "We continue to be supply constrained, a sign of momentum & rapid adoption." Andy Jassy of Amazon was more explicit: "We will still not have enough capacity to meet all the demand we have in 2026, & I believe this dynamic will also be true in 2027 too. The demand we already have for 2028 is striking." Jensen Huang at NVIDIA added that "Blackwell sales are off the charts, & cloud GPUs are sold out." These declarations set up a paradox: persistent scarcity typically triggers price increases, yet AI labs—facing capacity constraints—are pursuing a mixed pricing strategy rather than raising prices uniformly across the board. Anthropic moved first, launching Fable 5 on July 24 at $50 per million output tokens, double the price of its prior flagship Opus 5 and a new ceiling for model pricing. Google's trajectory was similar: Gemini's flagship climbed from $1.50 to $12 per million tokens across four generations. But the picture reversed when OpenAI cut prices sharply, reducing GPT-5.6 Luna by 80% just five days after Fable 5's launch, an aggressive move attributed either to market capture or a fundamental breakthrough in cost. The labs' apparent solution is segmentation into three tiers, each serving a different customer segment and use case. The premium tier—Anthropic's Fable 5, Anthropic's Opus 5, and others—commands the highest price. The mid-market tier is newly ascendant: GPT-5.6 Sol (OpenAI) and Kimi K3 deliver 96% of frontier intelligence at 40% of premium cost, making them attractive for workloads where maximum capability is less critical. The value tier ranges from GLM-5.2 down to DeepSeek V4 Flash at $0.03 per million tokens, offering 84% of frontier intelligence at 1 to 5% of premium cost. The bet is that this segmentation sustains Jevons' Paradox—the economic principle that as the cost of a resource falls, total consumption rises—by absorbing workloads that would otherwise drop out of the market at higher prices. As this segmentation becomes standard, the router—software that intelligently routes a query to the most appropriate model for a given task and budget—emerges as the strategic battleground. Routers will live in three places: embedded inside models themselves, built into customer applications, and in middleware harnesses. The major labs each need competitive entries across all three tiers, or they lose the router auction to competitors who can undercut them on price. Startups have an advantage: by owning a single point on the frontier—such as DeepSeek V4 Flash's existence proof of extreme cost efficiency at $0.03 per million tokens—they can capture router selection even if they cannot compete across all tiers.

Context & Analysis

Hyperscalers' Q2 2026 earnings calls painted a picture of sustained scarcity: Sundar Pichai (Alphabet), Andy Jassy (Amazon), and Jensen Huang (NVIDIA) all affirmed that AI compute supply will lag demand not just through 2026 but into 2027, with 2028 demand already heavily booked. This creates a classic pricing dilemma: capacity shortage typically forces prices up, but higher prices destroy the consumption growth that justifies further infrastructure investment. AI labs appear to be solving this via segmentation rather than uniform pricing. Instead of raising all prices or cutting all prices, they are launching distinct model tiers—premium (Anthropic's Fable 5 at $50/M output tokens), mid-market (GPT-5.6 Sol, Kimi K3), and value (GLM-5.2 down to DeepSeek V4 Flash at $0.03/M tokens)—each capturing a different cost-sensitivity cohort. The mid-market tier is the novel entry, positioned to convert price-sensitive buyers who would otherwise forgo consumption entirely. By keeping the value tier cheap and the mid-market tier performant at a fraction of premium cost, labs aim to sustain Jevons' Paradox—the dynamic that lower cost drives higher total consumption—even while raising premium prices. This hinges on a new layer of competition: routers, the software that directs customer queries to the economically optimal model for each workload. Routers become the gatekeeper between buyer and model, and labs that lack competitive offerings in all three tiers risk losing router selection battles to rivals or startups that can undercut them on price or capability at a specific tier.

FAQ

Are AI prices going up or down?
Both. Anthropic raised Fable 5 to $50 per million output tokens (doubling Opus 5), and Google's Gemini flagship climbed from $1.50 to $12 across generations. However, OpenAI cut GPT-5.6 Luna prices by 80% five days after Fable 5's launch. Labs are using tiered pricing rather than uniform price moves: premium tiers cost 13× the value tier, while mid-market models deliver 96% of frontier intelligence at 40% of the cost.
What is the mid-market tier and why does it matter?
The mid-market tier is a new offering from labs like OpenAI (GPT-5.6 Sol) and Kimi (K3), delivering 96% of frontier intelligence at 40% of premium cost. AI labs believe that if mid-market and value tiers absorb workloads that cannot afford premium pricing, total GPU-hours consumed will keep growing despite capacity constraints and rising headline prices.
How is DeepSeek V4 Flash competing?
DeepSeek V4 Flash is priced at $0.03 per million tokens, in the value tier offering 84% of frontier intelligence at 1 to 5% of premium cost. The article cites it as the existence proof that startups can win by owning a single point on the frontier that large labs cannot economically match.
Tomasz Tunguz (Theory Ventures)Read Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleBest uncensored AI image tools: local FLUX, Krea 2 Turbo, Wan models

The AI news that matters, in one minute each morning.

Sign up free