
What happened
Nvidia revealed performance benchmarks for its Vera Rubin chip system, a CPU–GPU combo that pairs 36 Vera CPUs with 72 Rubin GPUs in a single NVL72 superchip. The system processes 10 times as many tokens per watt as Nvidia's previous Grace Blackwell chip, and OpenAI already has one Vera Rubin rack in use.
Why it matters
Nvidia is no longer just a GPU supplier—it's positioning itself as a full AI infrastructure provider. As AI workloads shift toward more complex, agentic systems, demand for CPUs to orchestrate data and networking has grown. Vera Rubin's monolithic chip design offers nearly three times as much memory bandwidth as Blackwell, addressing a persistent high-bandwidth memory shortage that constrains AI deployments.
What to watch
Nvidia says Vera Rubin will ship in the second half of this year, with early customers including Microsoft, OpenAI, and Oracle. The company claims significantly reduced installation time—from a couple of hours to a few minutes—thanks to "cable-free compute" and hot-swappable design. AMD is revealing its competing Helios AI chip rack this week, intensifying the race for multiyear contracts with AI hyperscalers and labs.
Summaries like this, in your inbox every morning.
Nvidia's Vera Rubin push marks a strategic pivot that the company has been signaling for months. As AI workloads have matured beyond pure training toward inference and agentic systems—where AI models orchestrate multiple tasks, call tools, and manage data flows—the bottleneck has shifted. GPUs remain dominant for raw computation, but CPUs now handle coordination, networking, and other orchestration duties. By packaging both into a single, optimized superchip with unified memory access, Nvidia is trying to lock customers into its entire stack rather than allow them to mix suppliers.
The timing is deliberate and competitive. AMD has been making inroads in data center CPUs over the past two years, leveraging its long expertise with x86 chiplet architectures—the industry standard. Nvidia is betting that its ARM-based, monolithic design will prove superior for the specific demands of AI (high memory bandwidth, low latency across data movement). The benchmarks Nvidia disclosed suggest it believes the architectural choice is a winner: a single integrated circuit avoids the "heavy tax on memory bandwidth" that stitched-together chiplets impose, a claim Nvidia executives made directly to journalists.
Nvidia's marketing blitz also comes at a sensitive moment. The company's Blackwell chips faced overheating issues when packed together in Nvidia's custom server racks, forcing design changes and delaying shipments—a credibility hit in a market where hyperscalers and AI labs demand reliability. By pre-announcing Vera Rubin benchmarks and revealing that OpenAI, Microsoft, and Oracle are already piloting the hardware, Nvidia is trying to rebuild confidence that Vera Rubin will ship on schedule and deliver the promised efficiency gains.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic released Claude Opus 5.5 today and cut its price 20%, with input at $4 per million tokens and output…
Firecrawl announced it has raised $75 million in a Series B round led by Smash Ventures, with participation fr…
ASRock is shifting its business focus toward AI

Anthropic and OpenAI, which spent early September warning that model capabilities are outrunning the safeguard…

Jessica Wachter of Wharton and co-authors estimate hyperscaler spending will reach nearly 1.1兆ドル by 2027, and…

OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, trained the same way as its top-tier GPT-6 Astra, an…
