AIToday
AI Business & IndustrySiliconANGLE AIPublished: Oct 2, 2026, 04:00 JST

Nvidia's Ian Buck ties AI factory economics to tokens per watt

Nvidia's Ian Buck ties AI factory economics to tokens per watt

3 Key Points

  1. What happened

    Nvidia's Ian Buck said at the Fully Connected event that a data center's natural cap is its power, and that Blackwell delivered a 30x improvement in tokens per watt.

  2. Why it matters

    That makes tokens per watt a central measure of AI factory economics, so each hardware generation's efficiency gain is what determines how much useful output a site can produce.

  3. What to watch

    The test is whether time-sensitive workloads like fintech adopt Nvidia's Groq 3 LPX inference accelerator with the Vera Rubin platform, which Buck said is drawing interest.

WHO IT HITSOperators of AI data centers and the customers buying capacity from them, such as CoreWeave users, face buying decisions based on tokens per watt rather than raw chip count.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Buck's argument is that the unit of account for an AI data center has changed. Rather than counting servers, cars or devices, he told theCUBE Research's Dave Vellante and John Furrier that the output is tokens, describing them as appreciating, revenue-generating, fungible, durable and productive parts of an economy. That framing matters because it shifts attention away from the chips themselves and toward the networking, storage, processors and software that must work together at scale.

He also pushed back on the idea that inference replaces training. As companies use models, he said, they refine, align and add data to them, which amounts to a little bit of training, visible in reinforcement learning and online alignment work. A second economic tier is emerging for low-latency workloads where faster reasoning commands a higher value, which is where the Groq 3 LPX and Vera Rubin combination fits.

The stakes appear to hinge on execution rather than demand. Power capacity caps how much infrastructure a site can deploy, so the pressure is on vendors to keep raising tokens per watt with each generation. Buck pointed to Blackwell as evidence of that pace and to CoreWeave's menu of configurations and higher-level inference services as a way to spare customers from overwhelming choices. Whether that partner-led model holds as agentic systems pull on multiple models, databases and tools is likely to shape how quickly AI factories convert power into revenue.

FAQ
What does Nvidia say is the limiting factor for AI data centers?
Nvidia VP Ian Buck said data centers have a natural cap, and that cap is their power. That makes tokens per watt the central measure of AI factory economics.
How much did Blackwell improve tokens per watt?
Buck said Nvidia targets upwards of 10 times more efficiency in tokens per watt each GPU generation, and that with Blackwell it got a 30x improvement in tokens per watt.
What is the Groq 3 LPX used for?
It is Nvidia's inference accelerator that works with the Vera Rubin platform to increase per-user token rates for time-sensitive workloads. Buck said Nvidia is seeing interest in areas like fintech.
SiliconANGLE AIRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleCardinal debuts OmegaHealth at Scaling Up 2026