AIToday
Tomasz Tunguz (Theory Ventures)Published: Aug 25, 2026, 01:00 JST3 min read

AI infrastructure bottleneck shifts: GPU to memory to CPU to storage

AI infrastructure bottleneck shifts: GPU to memory to CPU to storage

Key takeaway

  • AI infrastructure shortages have moved through GPUs, memory, CPUs, and storage. Each bottleneck takes years to resolve.

  • Data center costs now exceed $20b per gigawatt. Suppliers have sold out production through 2029.

  • Overcapacity risk looms if software revenues lag.

3 Key Points

  1. What happened

    The AI infrastructure bottleneck has shifted over time, starting with GPU shortages in early 2023, then memory, then server CPUs by late 2025, and reaching bulk storage by 2026. Nvidia H100 rental rates surpassed $9 per hour, server shipments fell 22% in 2023, enterprise SSD prices rose 80% in one quarter, Intel server CPU ASPs rose 27% year over year, and Western Digital & Seagate confirmed their entire 2026 nearline production is sold out.

  2. Why it matters

    Each bottleneck freezes the next component's supply chain, with lag times running in years and every wave locking in a higher baseline cost. The fastest-rising cost is now the data center itself, with facilities costing upwards of $20b per gigawatt and construction costs tripled to $1,033 per square foot. GE Vernova & Siemens Energy have sold out turbine production through 2029.

  3. What to watch

    Over $2b in domestic transformer expansions, next-generation 300-layer NAND fabs, and new turbine production lines will deliver in 2027 & 2028. If end-user software revenues do not keep pace with $20b per gigawatt facilities, capital expenditure will face a classic crack of the whip—a potential overcapacity risk in long-lead-time capital goods.

Ask the AI about this article →

Context & Analysis

The article traces a clear sequence of supply chain shocks, each triggered by the previous one. In early 2023, GPU scarcity dominated as ChatGPT launch drove buyers to concentrate capital on procuring GPUs, starving conventional computing and causing server shipments to fall 22%. Memory makers, already reeling from a post-pandemic glut costing over $20b, lost server demand that would have absorbed inventory—leading them to convert production to HBM, which then drove up memory prices.

Eighteen months later, the pressure migrated to memory, with enterprise SSD prices rising 80% in a single quarter. By late 2025, agents squeezed server CPUs as agentic workflows pushed the CPU-to-GPU ratio toward 1:1. By 2026, the shortage reached bulk storage, with cloud architects retreating to traditional HDDs as flash hit $150 per terabyte.

Beyond the server chassis, data center construction costs have tripled to $1,033 per square foot, with electrical systems consuming half the budget. The quote from Siemens' Barry Powell captures the dilemma: build too little and lose market share, build too much and face fixed costs. The article suggests that if software revenues do not keep pace with the massive capital expenditures, a bubble is possible—with the wave finally breaking on long-lead-time capital goods.

FAQ

Why did memory prices surge after the GPU shortage?
Manufacturers converted cleanrooms and lithography tools toward High Bandwidth Memory (HBM) to chase AI margins. HBM consumes roughly three times the wafer capacity per gigabyte of standard DDR5, so every bit of HBM output removes about three bits of conventional supply, driving up prices.
What caused the server CPU shortage in late 2025?
Training clusters ran one CPU to eight GPUs, but agentic workflows invert that ratio toward 1:1. Autonomous systems spend their cycles compiling code, calling tools, and managing state, pushing demand up. Intel reported server CPU average selling prices rose 27% year over year against falling unit volumes.
What is the Bullwhip Effect in this context?
When a value chain suffers from multi-year manufacturing latency, sudden demand shocks downstream amplify into massive, lagged overreactions upstream. Relieving pressure at one bottleneck pushes it into the next component with a predictable delay, leaving long-lead-time capital goods exposed to overcapacity when the wave breaks.
Tomasz Tunguz (Theory Ventures)Read Original Article

Get AI news like this every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Next articleAbacus.AI全产品指南:从ChatLLM到Claw与Hermes