AIToday
Large Language ModelsAI Stocks & MarketsAI Business & IndustryLatent SpacePublished: Aug 27, 2026, 13:00 JST2 min read

OpenAI custom chip Jalapeño beats Nvidia in tests

OpenAI custom chip Jalapeño beats Nvidia in tests

Key takeaway

  • OpenAI unveiled benchmark results for its custom inference chip Jalapeño. It claims better efficiency and latency than NVIDIA GB200/GB300 systems.

  • Deployment into OpenAI's own infrastructure begins by year-end.

  • The chip also benefited from AI-assisted kernel optimization.

3 Key Points

  1. What happened

    OpenAI revealed benchmark results for its custom inference chip Jalapeño, claiming 1.5–1.9× more work per watt at peak throughput and 1.7–3.6× lower end-to-end latency than NVIDIA GB200/GB300 systems. The chip is rated at 700W but reportedly stayed at or below 550W on tested runs.

  2. Why it matters

    This suggests frontier labs may no longer be strictly downstream of NVIDIA for inference economics, even if packaging and foundry capacity remain a hard bottleneck. OpenAI says deployment into its own infrastructure begins by year-end, with Gen 2 already deep in development.

  3. What to watch

    OpenAI's post also said GPT-Astra + Codex helped write and optimize low-level kernels, bringing three previously unplanned open-weight models to high performance on Jalapeño in about two months. Selected attention and MoE blocks reportedly ran 1.5–1.8× faster than existing human-expert-written code.

Ask the AI about this article →

Context & Analysis

The Jalapeño announcement dominated the 37th Hot Chips conference, less than a year after OpenAI's Broadcom announcement. The chip is not an ASIC but a full alternative that reportedly beats Blackwell-class systems. SemiAnalysis framed it as unusually strong for a first-generation chip, comparing it directly against Blackwell and Rubin-class systems.

A second-order story is model-assisted systems optimization. GPT-Astra + Codex helped write and optimize low-level kernels, bringing three previously unplanned open-weight models to high performance in about two months. This suggests compiler and kernel work is increasingly being folded into the model improvement loop, not just application-layer coding.

Multiple technical reactions noted Jalapeño performed well even without tricks like aggressive prefill/decode disaggregation or speculative decoding in some setups, while beating systems that did use them. Several posts tie this to a larger industry transition where frontier labs may no longer be strictly downstream of NVIDIA for inference economics, though packaging and foundry capacity remain a hard bottleneck.

FAQ

What are the specific performance claims for Jalapeño?
OpenAI claims 1.5–1.9× more work per watt at peak throughput and 1.7–3.6× lower end-to-end latency than NVIDIA GB200/GB300 systems. For highly interactive workloads, it delivered 2.1–4.1× higher performance.
When will Jalapeño be deployed?
OpenAI says deployment into its own infrastructure begins by year-end. Gen 2 is already deep in development, and Gen 3 is underway.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNvidia forecasts 70% revenue growth for fiscal 2028