AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Aug 25, 2026, 22:01 JST2 min read

Nvidia Groq 3 LPX hits 3400 tokens/s, 4x faster than Cerebras

Nvidia Groq 3 LPX hits 3400 tokens/s, 4x faster than Cerebras

Key takeaway

  • Nvidia's Groq 3 LPX is now in full production.

  • It set a record of 3,400 tokens per second on a standard test.

  • Nvidia claims it's four times faster than Cerebras, but experts say the comparison favors Nvidia.

3 Key Points

  1. What happened

    Nvidia has moved its Groq 3 LPX inference accelerator into full production. An independent benchmark from Artificial Analysis shows it hitting 3,400 tokens per second on the open model Gemma 4 31B, the highest figure ever recorded for that model.

  2. Why it matters

    Nvidia says this makes the chip four times faster than the Cerebras chip, which runs at 882 tokens per second. Speed is critical for agentic AI systems, which generate huge numbers of tokens across many inference steps, so faster token generation means more reasoning steps and tool calls within a user-acceptable wait time.

  3. What to watch

    The benchmark may not reflect real-world performance. The comparison leaves out chip counts—Cerebras uses one or two accelerators while Nvidia needs at least 64—and doesn't factor in Cerebras' newest CS-4 generation. Nebius plans to be the first cloud provider to offer the chip through its Token Factory; Groq itself is an early user.

Ask the AI about this article →

Context & Analysis

Nvidia's move to production for the Groq 3 LPX comes after a late-December acquisition of the Groq license for about $20 billion, bringing in founder Jonathan Ross and president Sunny Madra. Groq's processors are specialized for inference, not training, which fits Nvidia's push into agentic AI where rapid token generation is key. The benchmark result is impressive, but as The Register notes, the architecture's memory limitations—each LPU has only 500 MB—mean larger models like DeepSeek V3 would need over five racks of accelerators, and the Cerebras comparison omits chip counts and the newer CS-4 generation. This suggests the real-world advantage may be narrower than the headline number implies, especially for complex, large-scale deployments. The arrival of Nebius as a first cloud provider could offer a practical test of the chip's performance outside Nvidia's own ecosystem.

FAQ

How much memory does the Groq 3 LPX have?
Each LPU has just 500 MB of memory, which is 576 times less than a Rubin GPU's 288 GB.
When will the Groq 3 LPX be available?
Nvidia says it will go live later this year. Nebius plans to be the first cloud provider to offer it.
Why is the benchmark not conclusive?
The benchmark uses a model that fits entirely in one rack, and the comparison doesn't account for chip counts or Cerebras' newest CS-4 generation. Larger models might require many more accelerators.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAI Companion Robot OlloNi SS1 Targets Loneliness with 'Gentle Intelligence'