AIToday
Large Language ModelsAI Business & IndustryTechCrunch AIPublished: Aug 26, 2026, 01:01 JST2 min read

OpenAI chip beats Nvidia in inference test

OpenAI chip beats Nvidia in inference test

Key takeaway

  • OpenAI unveiled benchmark results for its Jalapeño chip at Hot Chips.

  • It beat Nvidia's Blackwell on tokens per user and throughput per kilowatt.

  • The chip is slated for small deployment by end of 2026.

3 Key Points

  1. What happened

    At the Hot Chips conference on Tuesday, OpenAI shared the first benchmark results for its Jalapeño chip. Tested on SemiAnalysis' InferenceX benchmark, it registered more tokens per user and more throughput per kilowatt than currently available state-of-the-art inference processors, which in the test was an Nvidia Blackwell system.

  2. Why it matters

    OpenAI's head of hardware, Richard Ho, said the results show "a very, very significant performance advance over state of the art," meaning the chip can serve more AI work per unit of power while returning responses more quickly. This is notable because the comparison is against an Nvidia system, though Ho noted the competition may have advanced by the time Jalapeño is fully deployed.

  3. What to watch

    Ho estimated Jalapeño will deploy at the end of 2026 "in very small volumes," with more significant deployment coming in 2027. It is designed to minimize delays during the prefill and communication phases of inference, which OpenAI says often act as bottlenecks.

Ask the AI about this article →

Context & Analysis

OpenAI's first benchmark results for Jalapeño are a significant data point in its effort to build its own hardware. The chip's performance on the InferenceX benchmark, showing an advantage over Nvidia's current Blackwell system, comes from a full-stack design approach. By developing the AI models, chip, and memory in concert, OpenAI targeted specific bottlenecks that slow down the serving of AI responses.

The chip is part of a multigenerational platform strategy, indicating a move toward tighter integration between OpenAI's software and its hardware. The deployment timeline is gradual, with very small volumes by the end of 2026 and more significant volumes the following year. This suggests that while the company is making progress, it will be some time before Jalapeño becomes a major factor in its data centers. The comparison against Nvidia was noted to be a snapshot in time, since Nvidia is likely to have advanced its own technology by the time Jalapeño reaches full deployment.

FAQ

When will the Jalapeño chip be available?
Jalapeño is estimated to deploy at the end of 2026 in "very small volumes," with more significant deployment coming in 2027.
Who developed the Jalapeño chip?
Jalapeño was developed by OpenAI in close collaboration with Broadcom, with OpenAI's own models assisting in the development process.
What makes Jalapeño different from other inference processors?
Jalapeño is designed to minimize delays during the prefill and communication phases of processing, which OpenAI says often act as bottlenecks. It minimizes data movement and communication delays by explicitly placing model state, including the KV cache, locally.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleIBM Releases Granite 4.2 Reasoning LLMs