
OpenAI's new AI chip, Jalapeño, promises faster responses and more efficient processing.
It outperforms Nvidia chips in tests, offering up to 1.9 times more work per watt.
Deployment starts in small volumes this year.
What happened
OpenAI announced its new AI chip, Jalapeño, which delivers faster responses and higher throughput compared to other AI systems. In tests, it performed 1.5–1.9 times more work per watt and had 1.7–3.6 times lower latency than Nvidia's GB200 or GB300 superchips.
Why it matters
The chip is designed for AI inference, the process of running a trained AI model to complete tasks. This could mean faster and more reliable AI responses for users, as noted by OpenAI hardware VP Richard Ho.
What to watch
OpenAI plans to deploy Jalapeño in small volumes by the end of this year, with volume ramp-up beginning in 2027. The company says it will not replace all its chips, continuing to work with partners like Nvidia.
Ask the AI about this article →
OpenAI's introduction of Jalapeño marks a significant step in its effort to optimize AI hardware. The chip's design as an ASIC, made with Broadcom, focuses on AI inference, which is the execution phase of AI models. The reported performance gains—1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower latency compared to Nvidia's GB200/GB300—are notable because they address the typical trade-off between speed and efficiency in AI systems. This could lead to faster and more reliable AI responses for end users, as well as more responsive agents, without necessarily escalating costs or energy use. However, OpenAI's announcement that it will continue relying on partners like Nvidia suggests that Jalapeño is not intended to replace existing infrastructure entirely but to complement it. The phased deployment—small volumes by end of this year, with a ramp-up into 2027—indicates a measured approach to integrating the chip. The benchmarking results, while promising, come from OpenAI's own tests using the InferenceX platform, so independent validation remains a future step. As demand for AI services grows, having a chip that can deliver faster responses at lower latency could prove crucial, but the company's broader compute strategy will involve balancing multiple hardware options.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Banco BS2, a Brasil-based digital bank, is pursuing a foundation-first approach to scaling enterprise AI
Google is launching Gemini Enterprise for Legal, a product that uses MCP connectors to link Google's AI model…

OpenAI unveiled benchmarks for its custom inference chip "Jalapeño" at the Hot Chips conference

Anthropic announced on Tuesday that it is merging the memory system used by chat and Claude Cowork, so Claude…

Cisco Systems Inc. and Nvidia Corp
Cisco announced an expansion of its Secure AI Factory to include full rack-scale compute, partnering with Supe…