
OpenAI has launched a preview of Ultrafast mode for its flagship GPT-5.6 Sol model, powered by Cerebras infrastructure from a ten-billion-dollar partnership, delivering up to 750 output tokens per second.
The service is available through the OpenAI API to select customers initially, with gradual expansion planned.
OpenAI positions the speed as enabling real-time use cases—such as analyzing logs during outages, evaluating financial transactions as they occur, and resolving customer inquiries instantly—and monetizes it as a premium tier above its existing Fast Mode offering.
What happened
OpenAI is previewing Ultrafast mode, powered by Cerebras infrastructure from a ten-billion-dollar partnership signed earlier this year. The service delivers up to 750 output tokens per second from GPT-5.6 Sol and will initially be available only through the OpenAI API to select customers, with gradual expansion as capacity grows.
Why it matters
The speed boost—described by OpenAI as combining smaller models' speed with large models' full capabilities—unlocks use cases that require real-time analysis. OpenAI highlights incident response (analyzing logs and code during outages), finance (evaluating market signals live), customer support (resolving complex inquiries instantly), and research (turning overnight batch jobs into interactive sessions). The company is already using the model internally for incident response.
What to watch
Ultrafast is the third tier in OpenAI's speed-based pricing ladder, following Fast Mode (up to 2.5x speed at roughly double the standard price). Companies interested in early access can sign up through a form on OpenAI's site.
OpenAI announced the preview launch of Ultrafast mode, a new tier of inference acceleration for its GPT-5.6 Sol model that delivers up to 750 output tokens per second. The capability is powered by Cerebras, the AI infrastructure company with which OpenAI signed a ten-billion-dollar partnership earlier in the year. The service will roll out first through the OpenAI API, restricted to select customers, with plans to expand access gradually as capacity permits. Interested companies can request early access via a signup form.
The speed breakthrough opens several classes of applications OpenAI believes were previously impractical at real-time scales. In incident response, engineers can have logs, code changes, and reports analyzed while an outage is still unfolding, enabling faster diagnosis and repair. OpenAI reports already using the model internally for this purpose. In finance, traders and risk teams could evaluate market signals and flag suspicious activity as conditions shift. Customer support could resolve complex multi-step inquiries without delay, and e-commerce platforms could answer product questions, check inventory, and personalize recommendations before a customer abandons their session. Research teams can convert overnight batch experiments into interactive work—testing an idea, reviewing results, adjusting approach, and rerunning without interrupting the workflow.
OpenAI's pricing strategy treats speed as a performance tier analogous to cloud infrastructure pricing. The company already offers Fast Mode for GPT-5.6 Sol at roughly double the standard price for up to 2.5x speed and lower latency. Ultrafast becomes a third, faster tier at an implied further premium. This model mirrors how AWS and other cloud providers charge more for the same computational resource delivered at higher performance, applying the same logic to AI inference. If inference speed becomes a bottleneck across industries—particularly in finance, customer support, and operations—this tiered model gives OpenAI direct revenue capture from the economic gains that faster inference unlocks.
OpenAI's Ultrafast mode marks a strategic shift in how the company monetizes its AI output. Rather than charging solely on token volume or model capability, OpenAI is now selling speed itself as a tiered service—a model long established in cloud infrastructure (AWS charges premium rates for higher performance) but novel in the AI inference market. The ten-billion-dollar Cerebras partnership, announced earlier in the year, now delivers tangible product value: the ability to serve 750 tokens per second, positioning speed as a direct revenue lever. This tiered approach (standard, Fast Mode at 2.5x with roughly double pricing, and now Ultrafast at an implied premium) assumes that speed will become a binding constraint for many use cases, particularly where real-time decisions matter—trading, incident response, or customer-facing interactions. OpenAI's internal use of the model for incident response during outages suggests the company views this not just as a commercial offering but as architecturally sound for high-stakes workflows.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Comfort Systems USA, an engineering and construction contractor specializing in plumbing, HVAC, and electrical…

NVIDIA unveiled a roughly $500 billion compute financing plan to support AI infrastructure buildout, according…

Snowflake's Observe announced general availability of a redesigned MCP server and new CLI tool that give AI ag…

Tim O'Reilly, the publisher and tech pioneer, is promoting open-source AI—not just open-weight models, but the…

Kog, a French startup founded by Gaël Delalleau, is using deep GPU-level software optimization to accelerate A…

Amazon, Google, Meta, and Microsoft have committed to building large natural gas power plants (ranging from gi…

The AI news that matters, in one minute each morning.
Sign up free