
Nvidia has revived its Rubin CPX AI chip with a major redesign.
Production is expected to begin in the first quarter of 2027.
It will work alongside the Vera Rubin systems, not replace them.
What happened
Analyst Ming-Chi Kuo says Nvidia has revived the Rubin CPX AI accelerator with a substantially redesigned architecture, with production expected to begin in the first quarter of 2027. The chip had appeared to disappear from Nvidia's roadmap earlier this year.
Why it matters
The revived CPX is designed to accelerate the prefill stage of AI inference — the process of reading and processing a model's input before it generates a response. It will feature 168GB of HBM4 memory per GPU, compared with 288GB in Nvidia's Rubin GPUs and 128GB of GDDR7 in the earlier design, and will move into a standalone MGX ETL rack, with customers able to configure systems with 64, 128, 192, or 256 CPX GPUs.
What to watch
Rather than replacing Rubin, CPX is expected to work alongside Nvidia's Vera Rubin NVL72 systems. Kuo says Nvidia recommends a 1:1 ratio of CPX to Rubin GPUs, with CPX handling prefill and generating the KV cache before transferring that information to Rubin over Ethernet RDMA for the decode stage.
Ask the AI about this article →
The revival of the Rubin CPX marks a strategic pivot for Nvidia, coming after the company appeared to remove the chip from its roadmap at GTC 2026, where it instead highlighted Groq 3 LPUs and LPX racks. Kuo's industry checks suggest that Nvidia has not abandoned the prefill acceleration niche, but is repositioning it with a stronger architecture. The shift to standalone MGX ETL racks and the larger memory footprint indicate a focus on handling the growing demands of AI inference, where prefill performance is critical. By recommending a 1:1 ratio of CPX to Rubin GPUs, Nvidia is positioning the two chips as complementary rather than competing, with CPX handling the compute-heavy prefill stage and Rubin handling the decode stage. This division of labor could optimize overall system efficiency for customers deploying large-scale AI models. The fact that Nvidia did not immediately respond to requests for comment leaves room for further clarification, but Kuo's track record as a supply-chain analyst lends weight to his claims.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Chinese large-model developer Z.ai says it can now support large-scale inference using roughly 100,000 domesti…

AI system scaling has pushed interconnect requirements inside data centers from chips and boards up to racks…

Palantir Technologies stock has posted multi-year gains, including an 11x return over 3 years

Apple has escalated its legal battle against OpenAI, claiming in a new court filing that OpenAI is actively de…

Samsung Electronics has locked up as much as 70% of its memory production capacity under long-term supply agre…

Recent controversies include Ajinomoto's official X account posting an AI-edited image and a restaurant menu s…
