AIToday

Nvidia pushes CPU push with Vera Rubin chip ahead of AMD showdown

WIRED AI3h ago
Nvidia pushes CPU push with Vera Rubin chip ahead of AMD showdown

Key takeaway

Nvidia unveiled its Vera Rubin chip system, which bundles CPUs and GPUs to power AI data centers and delivers 10 times as many tokens per watt as its previous generation. The move reflects a broader industry shift: as AI workloads become more complex and agentic, demand for CPUs—not just GPUs—has surged, and Nvidia wants to own that entire stack. Early adopters including OpenAI, Microsoft, and Oracle are already testing the hardware, which Nvidia says will ship in the second half of 2025.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Nvidia revealed performance benchmarks for its Vera Rubin chip system, a CPU–GPU combo that pairs 36 Vera CPUs with 72 Rubin GPUs in a single NVL72 superchip. The system processes 10 times as many tokens per watt as Nvidia's previous Grace Blackwell chip, and OpenAI already has one Vera Rubin rack in use.

  • Why it matters

    Nvidia is no longer just a GPU supplier—it's positioning itself as a full AI infrastructure provider. As AI workloads shift toward more complex, agentic systems, demand for CPUs to orchestrate data and networking has grown. Vera Rubin's monolithic chip design offers nearly three times as much memory bandwidth as Blackwell, addressing a persistent high-bandwidth memory shortage that constrains AI deployments.

  • What to watch

    Nvidia says Vera Rubin will ship in the second half of this year, with early customers including Microsoft, OpenAI, and Oracle. The company claims significantly reduced installation time—from a couple of hours to a few minutes—thanks to "cable-free compute" and hot-swappable design. AMD is revealing its competing Helios AI chip rack this week, intensifying the race for multiyear contracts with AI hyperscalers and labs.

In Depth

Nvidia held a technical workshop at its Santa Clara headquarters last week to showcase its Vera Rubin chip system, a hybrid CPU–GPU superchip designed to power the next generation of AI data centers. Ian Buck, Nvidia's vice president of accelerated computing and architect of the company's CUDA software, led the briefings. "We're on a road map to crank out new architectures, not just GPUs but CPUs," Buck told reporters. "We're going to keep innovating, because it's do this or die in Silicon Valley."

Vera Rubin is the successor to Nvidia's Grace Blackwell hybrid system and represents the company's strategy to supply not just the compute engines for AI but the entire infrastructure stack. In a single Vera Rubin NVL72 superchip, there are 36 Vera CPUs paired with 72 Rubin GPUs—designed to handle the orchestration, networking, and data-flow management that increasingly complex AI agents require. Nvidia is also selling the Vera CPU as a stand-alone product and has told Chinese customers these could be ready by August.

The performance gains are substantial. Nvidia claims the Vera Rubin NVL72 system will process 10 times as many tokens per watt as Grace Blackwell, the measure of energy efficiency in AI workloads. The new Vera CPU is faster at processing agentic AI tasks than rival CPUs from AMD and Intel, though Nvidia's own testing used slightly older generations of competitors' hardware. Localized memory subsystems will offer nearly three times as much memory bandwidth as Blackwell—a critical feature given the persistent shortage of high-bandwidth memory that has constrained AI deployments.

Nvidia is also highlighting operational improvements. The Vera Rubin NVL72 racks—stacks of chips in a single liquid-cooled platform—are touted as "cable-free compute" and "hot-swappable," meaning customers can install each rack in a few minutes rather than a couple of hours. Andrew Bell, Nvidia's senior vice president of hardware engineering, emphasized this alongside Buck. Full liquid cooling reduces the energy needed to cool the chips compared to air-cooling, adding another efficiency angle.

The timing of the announcement is strategic. Nvidia unveiled Vera Rubin in spring 2025 and has been slowly releasing details while insisting the system will ship on schedule. CEO Jensen Huang has repeatedly said Vera Rubin is ramping to "full production" and will ship in the second half of 2025, with early customers including Microsoft, OpenAI, and Oracle. During a tour of a Nvidia data center lab in Silicon Valley, executives confirmed that OpenAI already has one Vera Rubin rack in use.

Nvidia is particularly sensitive to any appearance of delay. Its previous-generation Blackwell chips reportedly overheated when connected in Nvidia's customized server racks, forcing design changes and pushing back shipments—a credibility wound in a market where AI hyperscalers and labs demand absolute reliability. The Vera Rubin rollout is therefore being presented as engineered and proven.

The push comes directly ahead of rival AMD's annual conference, where the company is expected to showcase its next-generation AI and data center chips. On Sunday, AMD revealed more details about its Helios AI chip rack, designed to compete with Nvidia's wares. Both companies are vying for large-scale, multiyear contracts with AI hyperscalers like Meta and Amazon and AI labs like OpenAI, Anthropic, and SpaceX. AMD has significantly grown its share of the data center CPU market over the past two years, leveraging its recognized expertise in chiplet architecture for x86 processors, which still dominate data center CPU revenue. Nvidia, by contrast, builds its data center CPUs on ARM, an alternative architecture known for power efficiency. Nvidia executives Buck and Hannah Coutand, who runs product marketing for Nvidia DGX Cloud, argued that Vera Rubin's monolithic chip design—abandoning the chiplet approach—eliminates the "heavy tax on memory bandwidth and data movement" that stitching multiple chiplets together imposes, allowing data to move more quickly across a single integrated circuit.

Context & Analysis

Nvidia's Vera Rubin push marks a strategic pivot that the company has been signaling for months. As AI workloads have matured beyond pure training toward inference and agentic systems—where AI models orchestrate multiple tasks, call tools, and manage data flows—the bottleneck has shifted. GPUs remain dominant for raw computation, but CPUs now handle coordination, networking, and other orchestration duties. By packaging both into a single, optimized superchip with unified memory access, Nvidia is trying to lock customers into its entire stack rather than allow them to mix suppliers.

The timing is deliberate and competitive. AMD has been making inroads in data center CPUs over the past two years, leveraging its long expertise with x86 chiplet architectures—the industry standard. Nvidia is betting that its ARM-based, monolithic design will prove superior for the specific demands of AI (high memory bandwidth, low latency across data movement). The benchmarks Nvidia disclosed suggest it believes the architectural choice is a winner: a single integrated circuit avoids the "heavy tax on memory bandwidth" that stitched-together chiplets impose, a claim Nvidia executives made directly to journalists.

Nvidia's marketing blitz also comes at a sensitive moment. The company's Blackwell chips faced overheating issues when packed together in Nvidia's custom server racks, forcing design changes and delaying shipments—a credibility hit in a market where hyperscalers and AI labs demand reliability. By pre-announcing Vera Rubin benchmarks and revealing that OpenAI, Microsoft, and Oracle are already piloting the hardware, Nvidia is trying to rebuild confidence that Vera Rubin will ship on schedule and deliver the promised efficiency gains.

FAQ

What is Vera Rubin and how does it differ from Nvidia's earlier chips?
Vera Rubin is Nvidia's successor to Grace Blackwell and pairs CPUs with GPUs in a single liquid-cooled superchip system. In the NVL72 configuration, there are 36 Vera CPUs for every 72 Rubin GPUs. It uses a monolithic chip design (one single integrated circuit) rather than the chiplet architecture rival AMD favors, allowing faster data movement across the chip.
When will Vera Rubin be available?
Nvidia says Vera Rubin will ship in the second half of 2025. The company has told Chinese customers that Vera CPUs sold as a stand-alone product could be ready as soon as August. Early customers already using the hardware include OpenAI, Microsoft, and Oracle.
How does Vera Rubin improve on Grace Blackwell?
Vera Rubin processes 10 times as many tokens per watt as Grace Blackwell and offers nearly three times as much memory bandwidth. Installation time is reduced from a couple of hours to a few minutes thanks to "cable-free compute" and hot-swappable design, and the system is 100 percent liquid-cooled to reduce cooling energy.

Get the latest AI Coding Assistants news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →