
Cerebras, an AI chipmaker competing with Nvidia, has unveiled the CS-4, a new system that the company claims delivers 30x the tokens per second per user compared with graphics processing units.
The system uses faster memory (SRAM) and single large processors rather than multiple smaller chips, allowing it to run large AI models at high speed for inference tasks.
Despite the announcement, Cerebras shares have fallen more than 35% since the company's May IPO, weighed down by Q2 losses and investor skepticism.
What happened
AI chipmaker Cerebras unveiled the CS-4, a rack-scale system with three wafer-scale WSE-3 Turbo processors containing 4 trillion transistors. The company claims the system delivers 30x the tokens per second per user compared with graphics processing units and is designed for AI inference rather than training.
Why it matters
Cerebras uses static random-access memory (SRAM) instead of the standard dynamic random-access memory (DRAM) found in Nvidia and AMD systems, making data transfer faster and shorter since the processors are a single product rather than multiple chips that must move data between them. CEO Andrew Feldman stated the CS-4 'delivers industry-leading speeds on the largest frontier models, fundamentally changing the paradigm' — meaning companies can now run full-scale AI models at high speed, not just smaller, less capable ones.
What to watch
Cerebras went public in May at $185 per share and began trading at $350, but shares have fallen more than 35% to $218 as of midday Tuesday. The company reported a Q2 loss per share of -$2.98 versus a profit of $1.91 in the same quarter last year, and while Q3 guidance beat expectations, it failed to impress investors.
Ask the AI about this article →
Cerebras is directly challenging Nvidia's dominance in the AI acceleration market with a fundamentally different architectural approach. While Nvidia's systems rely on multiple discrete GPUs connected together, Cerebras has built dinner-plate-sized wafers that can fit faster SRAM memory — a technology that is normally impractical at smaller scales because it requires too much physical space. By consolidating processing into a single monolithic chip, Cerebras reduces the distance data must travel, a key bottleneck in multi-chip systems.
The timing of this announcement is significant given the intense competition for inference workloads. As AI models grow larger and companies deploy them in production, the speed and efficiency of inference — the step where an AI model produces an answer to a user query — has become a critical competitive factor. Cerebras' claim of 30x throughput advantage, if validated, could reshape how companies choose hardware for serving large models at scale.
However, Cerebras faces an uphill battle in investor confidence. The company's stock has tumbled since its May IPO, driven by a swing from profitability in Q2 of last year to a $2.98 loss per share in Q2 of this year. Even better-than-expected Q3 guidance has not restored investor appetite, suggesting the market is skeptical of the company's path to sustainable returns despite its technological claims.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Visko raised $10 million in pre-seed funding from Llama Ventures and opened public access to its first foundat…
U.S. markets ended August higher, with the S&P 500 up 2.6% and the Nasdaq up 3.9%

Neurovia AI, an Abu Dhabi-based company, is pitching Saudi security agencies software that it says can compres…

AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." I…

John Deere introduced JD, a conversational AI tool that lets farmers ask open-ended questions about their hist…

Nvidia CEO Jensen Huang said on Fox Business that AI is creating 'hundreds of thousands' of jobs, including in…
