
OpenAI and Cerebras have launched GPT-5.6 Sol Ultrafast, a new inference tier delivering 750 output tokens per second without quality degradation.
In head-to-head benchmarks on the Humanity's Last Exam (2,500 PhD-level questions), Ultrafast completed all questions in 11 hours and 11 minutes—nearly 7× faster than Claude Fable 5—making it practical for time-critical applications like production incident response and security threat detection.
The service is available now in limited preview.
What happened
OpenAI and Cerebras announced Ultrafast Mode, a new service tier delivering up to 750 output tokens per second for GPT-5.6 Sol, powered by Cerebras' Wafer-Scale Engine. The service launched in limited preview to select customers, with access expanding over time.
Why it matters
Ultrafast resolves the traditional tradeoff between response speed and model quality. On the Humanity's Last Exam benchmark (2,500 PhD-level questions), GPT-5.6 Sol Ultrafast answered all questions in 11 hours and 11 minutes—nearly 7× faster than Claude Fable 5 (78 hours and 27 minutes)—with comparable accuracy. For organizations, this enables real-time problem-solving in time-critical domains: production outages, security incidents, and agent-driven workflows where latency directly impacts business outcomes.
What to watch
GPT-5.6 Sol Ultrafast is available now in limited preview; sign up for updates to gain access as capacity grows. The service runs on Cerebras' architecture, which packs 44 GB of SRAM on a single wafer-sized chip to eliminate memory bandwidth bottlenecks that slow inference on GPUs.
OpenAI and Cerebras announced Ultrafast Mode on the OpenAI API, a new inference tier powered by Cerebras' Wafer-Scale Engine architecture. The service delivers up to 750 output tokens per second for GPT-5.6 Sol while maintaining output quality. Initially available to a select group of customers, access will expand as capacity grows.
The core problem Ultrafast addresses is the longstanding tradeoff between model intelligence and response speed. As models scale in capability, they incur higher computational and data movement costs, slowing inference times. Users traditionally faced a choice: wait for high-quality results or accept inferior outputs delivered quickly. Ultrafast resolves this by bringing frontier-grade intelligence to applications where latency matters.
Benchmark results demonstrate the speed advantage concretely. On the Humanity's Last Exam—a challenging benchmark of 2,500 questions typically answerable only by PhD-holders in fields such as chemistry, economics, and literature—GPT-5.6 Sol Ultrafast answered all questions in 11 hours and 11 minutes. Claude Fable 5 required 78 hours and 27 minutes for the same task, meaning Ultrafast worked through the benchmark nearly 7× faster with comparable accuracy. Cerebras performed its testing on July 10 using GPT-5.6 Sol Ultrafast with Codex on xhigh reasoning; Claude Fable 5 was tested on July 13–15 with Claude Code on xhigh reasoning. On GDP-Val, a benchmark for economically valuable knowledge work tasks, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation, tested by Cerebras on July 31, 2026.
Rohan Varma, Product at OpenAI, stated: "With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate. We're excited to see how workflows and applications are transformed by Ultrafast inference." Jeffrey Wang, an OpenAI researcher, noted the practical impact: "Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive."
The speed improvement is enabled by Cerebras' Wafer-Scale Engine architecture, designed specifically for frontier AI workloads. Traditional GPU-based inference is bottlenecked by memory bandwidth: model weights must be repeatedly transferred between on-chip memory and off-chip storage to generate successive tokens. Cerebras takes a contrarian approach by packing 44 GB of SRAM on each wafer-sized chip. Weights stay on-chip, and tokens flow uninterrupted through model layers pipelined across wafers. This design scales smoothly with model size, potentially sustaining the speed advantage on future frontier models.
Concretely, GPT-5.6 Sol Ultrafast enables new modes of work. In production incident response, teams can leverage Ultrafast to root-cause and address outages quickly, preserving customer trust and preventing lost revenue. Security teams can use it to detect and respond to cyberattacks in real time. More broadly, Ultrafast enables agents to be deployed on the critical path of problems where every second counts, allowing researchers and engineers to focus their attention on deep, high-value problems while offloading parallelizable commodity tasks to Standard processing. GPT-5.6 Sol is positioned as OpenAI's best model for legal briefs, financial models, and engineering reports.
The announcement resolves a long-standing constraint in AI deployment: the need to choose between response speed and model quality. As language models have grown more capable, they have also become computationally heavier, forcing organizations to either wait for high-quality outputs or accept degraded results within tight time windows. GPT-5.6 Sol Ultrafast changes this calculus by delivering frontier-grade intelligence—comparable accuracy to standard GPT-5.6 Sol—at speeds (750 tokens per second) that make real-time applications feasible.
The performance gains are substantive and grounded in benchmark evidence. On the Humanity's Last Exam—a dataset of 2,500 PhD-level questions spanning chemistry, economics, and literature—Ultrafast completed the full set in 11 hours and 11 minutes versus 78 hours and 27 minutes for Claude Fable 5. This speed advantage compounds in practical scenarios: organizations can now deploy AI agents on the critical path of production incident response, security threat detection, and financial modeling, where latency translates directly to operational and financial risk. The architecture's ability to scale smoothly with model size suggests the speed advantage may persist across future frontier models.
The limited preview model indicates Cerebras and OpenAI are being cautious about capacity constraints, likely due to the specialized hardware required. As infrastructure expands, broader adoption could reshape workflows in law, finance, engineering, and security—domains where both intelligence and response time are mission-critical.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Sam Altman is demanding OpenAI's IPO valuation reach $1 trillion, but SoftBank must repay a $40 billion bridge…

Bank of America raised its 2030 server CPU market estimate to more than $210 billion from about $170 billion…

DeepSeek moved its V4-Pro flagship model to production (build V4-Pro-0813), released its proprietary agent har…

Google released Gemini 3.7 Flash, its successor to Gemini 3.6 Flash (released three weeks earlier), available…

Microsoft has begun integrating its consumer and enterprise Copilot applications into a unified platform, star…

A functional programming team built LLM agent systems in Clojure and Elixir, comparing them directly with Pyth…

The AI news that matters, in one minute each morning.
Sign up free