
Nvidia's Groq 3 LPX is now in full production.
It set a record of 3,400 tokens per second on a standard test.
Nvidia claims it's four times faster than Cerebras, but experts say the comparison favors Nvidia.
What happened
Nvidia has moved its Groq 3 LPX inference accelerator into full production. An independent benchmark from Artificial Analysis shows it hitting 3,400 tokens per second on the open model Gemma 4 31B, the highest figure ever recorded for that model.
Why it matters
Nvidia says this makes the chip four times faster than the Cerebras chip, which runs at 882 tokens per second. Speed is critical for agentic AI systems, which generate huge numbers of tokens across many inference steps, so faster token generation means more reasoning steps and tool calls within a user-acceptable wait time.
What to watch
The benchmark may not reflect real-world performance. The comparison leaves out chip counts—Cerebras uses one or two accelerators while Nvidia needs at least 64—and doesn't factor in Cerebras' newest CS-4 generation. Nebius plans to be the first cloud provider to offer the chip through its Token Factory; Groq itself is an early user.
Ask the AI about this article →
Nvidia's move to production for the Groq 3 LPX comes after a late-December acquisition of the Groq license for about $20 billion, bringing in founder Jonathan Ross and president Sunny Madra. Groq's processors are specialized for inference, not training, which fits Nvidia's push into agentic AI where rapid token generation is key. The benchmark result is impressive, but as The Register notes, the architecture's memory limitations—each LPU has only 500 MB—mean larger models like DeepSeek V3 would need over five racks of accelerators, and the Cerebras comparison omits chip counts and the newer CS-4 generation. This suggests the real-world advantage may be narrower than the headline number implies, especially for complex, large-scale deployments. The arrival of Nebius as a first cloud provider could offer a practical test of the chip's performance outside Nvidia's own ecosystem.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
WalkMe's third annual AI at Work Pulse Survey, of 2,037 U.S

Hugging Face revived Papers with Code with a hybrid search system combining keyword and vector search, using i…

OpenAI's ChatGPT Work, a platform for white-collar workers to use AI agents, has reached 20 million users, acc…

An unknown AI model called Ox Alpha appeared on OpenRouter on August 20 and reached #1 in weekly token consump…

The Vanguard Utilities ETF (VPU) is highlighted as a potential way to profit from the AI boom, as AI data cent…

AWS hit $169 billion in annual recurring revenue last quarter, with sales up 37% year over year
