
OpenAI's custom chip Jalapeño beats Nvidia's Blackwell and Rubin in inference benchmarks. It delivers 1.5x to 1.9x more work per watt.
The chip is still in samples, not yet shipping.
SemiAnalysis says the CUDA moat may be dead.
What happened
OpenAI unveiled benchmarks for its custom inference chip "Jalapeño" at the Hot Chips conference. It reportedly outperforms Nvidia's Blackwell and Rubin in throughput per watt and token latency.
Why it matters
Jalapeño delivers 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower latency than the best commercial systems. This could challenge Nvidia's dominance in AI chips, with SemiAnalysis suggesting the "CUDA moat" may be "potentially dead."
What to watch
Jalapeño is still in engineering samples, while Nvidia's Rubin systems are shipping. OpenAI claims it built the chip in nine months using its own AI models, and SemiAnalysis notes it beats Vera Rubin in output tokens per megawatt even without multi-token prediction.
Ask the AI about this article →
OpenAI's Jalapeño chip, developed with Broadcom, took only nine months from first design to manufacturing blueprint, a speed that SemiAnalysis says signals the end of Nvidia's CUDA moat. The chip delivers up to 104x token throughput per kilowatt at matched decoding speed, and it does so without relying on optimizations like multi-token prediction, leaving room for further gains. However, caveats remain: competitors like Nvidia and AMD have tested larger models on their chips, and Jalapeño is still an engineering sample, not yet in production. OpenAI's CFO Sarah Friar frames the chip as part of an integrated compute strategy, complementing existing partnerships with Nvidia, AMD, and others, rather than replacing them. This dual dynamic of cooperation and competition underscores the broader industry trend where major players are developing custom silicon, even as they continue to buy from each other.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic PBC announced it is changing how Claude, its flagship AI product, uses memory, allowing users to see…
Banco BS2, a Brasil-based digital bank, is pursuing a foundation-first approach to scaling enterprise AI
McKinsey's 2026 State of AI survey of 1,719 professionals found that 37% of respondents attribute at least som…

Chinese automakers are moving aggressively into humanoid robotics, drawing on existing technology, supply chai…

Nvidia has guaranteed up to $105 billion of OpenAI's data center leases at the PORTS-Pike campus in Ohio

A design principle for AI agents is proposed: daily assistants should have a 24-hour session that resets at mi…
