AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Aug 15, 2026, 01:00 JST3 min read

OpenAI launches Ultrafast mode for GPT-5.6 Sol, delivering 750 tokens/second

OpenAI launches Ultrafast mode for GPT-5.6 Sol, delivering 750 tokens/second

Key takeaway

  • OpenAI has launched a preview of Ultrafast mode for its flagship GPT-5.6 Sol model, powered by Cerebras infrastructure from a ten-billion-dollar partnership, delivering up to 750 output tokens per second.

  • The service is available through the OpenAI API to select customers initially, with gradual expansion planned.

  • OpenAI positions the speed as enabling real-time use cases—such as analyzing logs during outages, evaluating financial transactions as they occur, and resolving customer inquiries instantly—and monetizes it as a premium tier above its existing Fast Mode offering.

3 Key Points

  1. What happened

    OpenAI is previewing Ultrafast mode, powered by Cerebras infrastructure from a ten-billion-dollar partnership signed earlier this year. The service delivers up to 750 output tokens per second from GPT-5.6 Sol and will initially be available only through the OpenAI API to select customers, with gradual expansion as capacity grows.

  2. Why it matters

    The speed boost—described by OpenAI as combining smaller models' speed with large models' full capabilities—unlocks use cases that require real-time analysis. OpenAI highlights incident response (analyzing logs and code during outages), finance (evaluating market signals live), customer support (resolving complex inquiries instantly), and research (turning overnight batch jobs into interactive sessions). The company is already using the model internally for incident response.

  3. What to watch

    Ultrafast is the third tier in OpenAI's speed-based pricing ladder, following Fast Mode (up to 2.5x speed at roughly double the standard price). Companies interested in early access can sign up through a form on OpenAI's site.

In Depth

Read the full story

OpenAI announced the preview launch of Ultrafast mode, a new tier of inference acceleration for its GPT-5.6 Sol model that delivers up to 750 output tokens per second. The capability is powered by Cerebras, the AI infrastructure company with which OpenAI signed a ten-billion-dollar partnership earlier in the year. The service will roll out first through the OpenAI API, restricted to select customers, with plans to expand access gradually as capacity permits. Interested companies can request early access via a signup form.

The speed breakthrough opens several classes of applications OpenAI believes were previously impractical at real-time scales. In incident response, engineers can have logs, code changes, and reports analyzed while an outage is still unfolding, enabling faster diagnosis and repair. OpenAI reports already using the model internally for this purpose. In finance, traders and risk teams could evaluate market signals and flag suspicious activity as conditions shift. Customer support could resolve complex multi-step inquiries without delay, and e-commerce platforms could answer product questions, check inventory, and personalize recommendations before a customer abandons their session. Research teams can convert overnight batch experiments into interactive work—testing an idea, reviewing results, adjusting approach, and rerunning without interrupting the workflow.

OpenAI's pricing strategy treats speed as a performance tier analogous to cloud infrastructure pricing. The company already offers Fast Mode for GPT-5.6 Sol at roughly double the standard price for up to 2.5x speed and lower latency. Ultrafast becomes a third, faster tier at an implied further premium. This model mirrors how AWS and other cloud providers charge more for the same computational resource delivered at higher performance, applying the same logic to AI inference. If inference speed becomes a bottleneck across industries—particularly in finance, customer support, and operations—this tiered model gives OpenAI direct revenue capture from the economic gains that faster inference unlocks.

Context & Analysis

OpenAI's Ultrafast mode marks a strategic shift in how the company monetizes its AI output. Rather than charging solely on token volume or model capability, OpenAI is now selling speed itself as a tiered service—a model long established in cloud infrastructure (AWS charges premium rates for higher performance) but novel in the AI inference market. The ten-billion-dollar Cerebras partnership, announced earlier in the year, now delivers tangible product value: the ability to serve 750 tokens per second, positioning speed as a direct revenue lever. This tiered approach (standard, Fast Mode at 2.5x with roughly double pricing, and now Ultrafast at an implied premium) assumes that speed will become a binding constraint for many use cases, particularly where real-time decisions matter—trading, incident response, or customer-facing interactions. OpenAI's internal use of the model for incident response during outages suggests the company views this not just as a commercial offering but as architecturally sound for high-stakes workflows.

FAQ

When will Ultrafast mode be available and who can use it?
Ultrafast is currently in preview and available only through the OpenAI API to select customers. OpenAI plans to expand access gradually as capacity grows; companies can sign up for updates through a form.
How much faster is Ultrafast compared to other OpenAI modes?
Ultrafast delivers up to 750 output tokens per second. OpenAI's Fast Mode offers up to 2.5x speed; Ultrafast is positioned as a faster, third tier in the company's tiered pricing model.
What is powering this speed improvement?
The inference acceleration comes from Cerebras, which signed a ten-billion-dollar partnership with OpenAI earlier this year.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleObserve launches MCP server and CLI for AI agents to access telemetry

The AI news that matters, in one minute each morning.

Sign up free