
What happened
GPU rental prices doubled in six months, from $4.40 to $8.08 per GPU-hour, even as Claude Opus 5.5 costs 40% less to run and OpenAI cut Luna 80% then 50%.
Why it matters
Rising GPU costs are likely to persist while data center buildouts face higher costs for everything from concrete to electricity, yet efficiency gains may be keeping gross profit per GPU-hour in balance.
What to watch
The metric that will ultimately matter is gross profit dollars per GPU-hour — watch whether Microsoft's 90% year-over-year gain in tokens per GPU continues, keeping the two curves neck and neck.
WHO IT HITSThis directly affects AI infrastructure investors and data-center operators who are bidding on limited GPU-hours, as well as the finance teams at AI model companies whose margins depend on whether efficiency gains outpace rising compute costs.
Summaries like this, in your inbox every morning.
The article frames a puzzle: GPU rental prices have doubled in six months, to $8.08 per GPU-hour from $4.40, yet AI inference keeps getting cheaper. Tom Tunguz points to the buildout's rising input costs — concrete, copper, credit — with electricity as the limiting factor. Oracle's force majeure on its New Mexico campus, after the natural-gas pipeline feeding it slipped by six months, illustrates how physical delays ripple through plans.
At the same time, demand from inference and model companies is posting record-setting growth, emboldening investors to value them more richly and bid up limited GPU-hours. Efficiency is improving just as fast: the same benchmark threshold that cost $0.55 eighteen months ago now clears for $0.0015, a 377x reduction. Capital itself is more expensive too, and the historical relationship between the 10-year Treasury and tech valuations has flipped — the correlation went from −0.50 over the full history to +0.39 over the last two years, while rates climbed from 3.63% to 5.18% and the NASDAQ rose roughly 78%.
Tunguz's reading is that the industry is holding both forces in balance for now, with gross profit dollars per GPU-hour as the ultimate test. Whether that balance holds seems to hinge on whether efficiency gains — like Microsoft's 90% more tokens per GPU year over year, concentrated in smaller models — continue to run neck and neck with rising GPU costs. If they do not, the economics for the companies paying those GPU-hour rates could shift.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Anthropic launched Claude Sonnet 5.5, a mid-tier model for everyday tasks, priced at $2 per million input toke…
NEAR, the token of the NEAR Protocol, has more than doubled in value over the past two weeks, helped by surgin…

Michael Burry wrote in a Substack chat that Trump's team knows the AI buildout is 'the only thing keeping this…

Anthropic released Claude Sonnet 5.5, which it says runs output more than 30 percent faster and costs up to 30…

The University of Florida found only 10-15% of its data was semantically defined; using NaviGator AI and Snowf…

In a Nikkei Money discussion, private investor Okeido said he sets his investing goal years out and basically…
