AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Aug 26, 2026, 04:01 JST2 min read

OpenAI's custom chip beats Nvidia in inference benchmarks

OpenAI's custom chip beats Nvidia in inference benchmarks

Key takeaway

  • OpenAI's custom chip Jalapeño beats Nvidia's Blackwell and Rubin in inference benchmarks. It delivers 1.5x to 1.9x more work per watt.

  • The chip is still in samples, not yet shipping.

  • SemiAnalysis says the CUDA moat may be dead.

3 Key Points

  1. What happened

    OpenAI unveiled benchmarks for its custom inference chip "Jalapeño" at the Hot Chips conference. It reportedly outperforms Nvidia's Blackwell and Rubin in throughput per watt and token latency.

  2. Why it matters

    Jalapeño delivers 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower latency than the best commercial systems. This could challenge Nvidia's dominance in AI chips, with SemiAnalysis suggesting the "CUDA moat" may be "potentially dead."

  3. What to watch

    Jalapeño is still in engineering samples, while Nvidia's Rubin systems are shipping. OpenAI claims it built the chip in nine months using its own AI models, and SemiAnalysis notes it beats Vera Rubin in output tokens per megawatt even without multi-token prediction.

Ask the AI about this article →

Context & Analysis

OpenAI's Jalapeño chip, developed with Broadcom, took only nine months from first design to manufacturing blueprint, a speed that SemiAnalysis says signals the end of Nvidia's CUDA moat. The chip delivers up to 104x token throughput per kilowatt at matched decoding speed, and it does so without relying on optimizations like multi-token prediction, leaving room for further gains. However, caveats remain: competitors like Nvidia and AMD have tested larger models on their chips, and Jalapeño is still an engineering sample, not yet in production. OpenAI's CFO Sarah Friar frames the chip as part of an integrated compute strategy, complementing existing partnerships with Nvidia, AMD, and others, rather than replacing them. This dual dynamic of cooperation and competition underscores the broader industry trend where major players are developing custom silicon, even as they continue to buy from each other.

FAQ

What is Jalapeño?
Jalapeño is OpenAI's first custom inference chip, developed with Broadcom. It runs AI models but doesn't train them, and it's a general-purpose LLM inference accelerator.
How does Jalapeño compare to Nvidia's chips?
Jalapeño outperforms Nvidia's Blackwell and Rubin in throughput per watt and token latency. Even against Vera Rubin with HBM4, it produces more output tokens per megawatt, though total cost per token is roughly even.
Is Jalapeño available now?
No, Jalapeño hasn't moved beyond engineering samples. In contrast, Nvidia's Rubin systems are already shipping to customers.

Also reported by TechCrunch AI

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleLinkedIn 'AI slop' button hits 1M clicks