
OpenAI unveiled benchmark results for its custom inference chip Jalapeño. It claims better efficiency and latency than NVIDIA GB200/GB300 systems.
Deployment into OpenAI's own infrastructure begins by year-end.
The chip also benefited from AI-assisted kernel optimization.
What happened
OpenAI revealed benchmark results for its custom inference chip Jalapeño, claiming 1.5–1.9× more work per watt at peak throughput and 1.7–3.6× lower end-to-end latency than NVIDIA GB200/GB300 systems. The chip is rated at 700W but reportedly stayed at or below 550W on tested runs.
Why it matters
This suggests frontier labs may no longer be strictly downstream of NVIDIA for inference economics, even if packaging and foundry capacity remain a hard bottleneck. OpenAI says deployment into its own infrastructure begins by year-end, with Gen 2 already deep in development.
What to watch
OpenAI's post also said GPT-Astra + Codex helped write and optimize low-level kernels, bringing three previously unplanned open-weight models to high performance on Jalapeño in about two months. Selected attention and MoE blocks reportedly ran 1.5–1.8× faster than existing human-expert-written code.
Ask the AI about this article →
The Jalapeño announcement dominated the 37th Hot Chips conference, less than a year after OpenAI's Broadcom announcement. The chip is not an ASIC but a full alternative that reportedly beats Blackwell-class systems. SemiAnalysis framed it as unusually strong for a first-generation chip, comparing it directly against Blackwell and Rubin-class systems.
A second-order story is model-assisted systems optimization. GPT-Astra + Codex helped write and optimize low-level kernels, bringing three previously unplanned open-weight models to high performance in about two months. This suggests compiler and kernel work is increasingly being folded into the model improvement loop, not just application-layer coding.
Multiple technical reactions noted Jalapeño performed well even without tricks like aggressive prefill/decode disaggregation or speculative decoding in some setups, while beating systems that did use them. Several posts tie this to a larger industry transition where frontier labs may no longer be strictly downstream of NVIDIA for inference economics, though packaging and foundry capacity remain a hard bottleneck.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Microsoft co-founder Bill Gates published an almost-6,000-word essay Tuesday warning that the transition to th…
SpaceX disclosed a partnership with Nvidia to deploy Vera CPUs as part of the architecture behind SpaceXAI's S…

The Agentic AI Foundation (AAIF) under the Linux Foundation has published a roadmap outlining five priority ar…

A Semafor analysis found that over the past month, 10 out of 310 guest submissions to The Wall Street Journal…

Nvidia projected its revenue would rise 70% next fiscal year, its first-ever year-ahead forecast

NVIDIA reported record profits, driven by strong demand for its AI chips, and is stepping up efforts to create…
