
Cerebras introduced the CS-4, an AI accelerator it claims is the fastest in the industry.
It doubles CS-3 performance on the same chip while keeping memory at 44 GB per wafer.
The product is used by OpenAI for Codex Spark, among others.
What happened
Cerebras has introduced its CS-4 AI accelerator, a rack-scale product that CEO Andrew Feldman calls the fastest system in the industry. It runs on the same 5nm WSE-3 chip as the CS-3 but delivers double the performance by boosting clock speed through more power and better cooling.
Why it matters
A single rack now holds three wafers instead of two and delivers up to 4,400 tokens per second per user, which Cerebras says is up to 30 times faster than setups running on Nvidia GPUs. Memory capacity stays the same at 44 GB per wafer.
What to watch
The CS-4 uses a new modular 'Backpack' design for faster assembly and disaggregated inference through partners like AMD and AWS Trainium. Analysts at SemiAnalysis see the networking gains as fairly small, and more details are coming at the Hot Chips conference.
Ask the AI about this article →
Cerebras's CS-4 represents a significant performance leap achieved not through a new chip design but through engineering around the existing 5nm WSE-3. By increasing power and improving cooling, the company has doubled performance, allowing a single rack to hold three wafers and deliver up to 4,400 tokens per second per user. This positions the system as a faster alternative to Nvidia GPU setups, though memory capacity remains unchanged.
The new 'Backpack' modular design aims to speed up assembly, and the move toward disaggregated inference with partners like AMD and AWS Trainium suggests a strategy to integrate into broader data center ecosystems. However, analysts at SemiAnalysis downplay the significance of the networking gains, indicating that while raw performance is impressive, the system's overall value may depend on other factors. With more details expected at the Hot Chips conference, the industry will be watching to see how this hardware performs in real-world deployments, especially given its use by OpenAI for Codex Spark.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Thomson Reuters Corp. today launched Thomson, its first proprietary large language model, combining its legal…
Amazon told investors it now expects to spend $220 billion in 2026, which is $20 billion more than its prior c…

BMO Capital started coverage of AMD with an Outperform rating and a $550 price target

More than 500 seed- or venture-backed private companies have sold to other private, venture-backed companies s…

Thomson Reuters launched its first in-house language model, built on Alibaba's Qwen, after spending about $40…

Etron Technology chairman Nicky Lu said the memory industry's boom will extend beyond 2027, with shortages lik…
