AIToday
Large Language ModelsAI Business & IndustryTechCrunch AIPublished: Aug 14, 2026, 06:01 JST3 min read

OpenAI launches Ultrafast mode for GPT-5.6 Sol, hitting 14x standard speed

OpenAI launches Ultrafast mode for GPT-5.6 Sol, hitting 14x standard speed

Key takeaway

  • OpenAI has introduced Ultrafast, a new mode for GPT-5.6 Sol that runs at 14x the speed of standard processing and outputs up to 750 tokens per second.

  • The company says this achieves "more useful work per second" without sacrificing the model's full capability, addressing a long-standing tradeoff between speed and power that typically forced users to pick a smaller model for faster responses.

  • Ultrafast is rolling out in preview to a limited set of customers, with broader access to follow as infrastructure capacity expands.

3 Key Points

  1. What happened

    OpenAI rolled out Ultrafast, a new mode for GPT-5.6 Sol that processes at 14x the speed of standard mode, delivering up to 750 output tokens per second. The feature is currently in preview and available to a small group of customers; expansion depends on capacity growth.

  2. Why it matters

    Real-time responsiveness has typically required choosing a smaller or more specialized model, forcing businesses to trade capability for speed. Ultrafast suggests it is now possible to maintain GPT-5.6 Sol's full power while achieving rapid response, potentially unlocking new use cases in incident response, customer service, financial analysis, and e-commerce where latency matters.

  3. What to watch

    Ultrafast is powered by OpenAI's partnership with chipmaker Cerebras and is currently in preview for a limited group; OpenAI says access will expand as capacity grows. Competitors like Anthropic offer Claude's fast mode, though it does not match Ultrafast's speed.

In Depth

Read the full story

OpenAI has introduced Ultrafast, a new operational mode for its flagship GPT-5.6 Sol model, designed to dramatically reduce latency without sacrificing model capability. The mode operates at 14x the speed of standard processing and can produce up to 750 output tokens (discrete units of text generated by an AI when responding to a prompt) per second. This represents a shift in the AI performance design space: as the company explained in a Thursday blog post, achieving real-time speed has historically required developers to choose a smaller or more specialized model. "Ultrafast," OpenAI wrote, "points to progress in a new direction: more useful work per second."

The technology is powered by OpenAI's partnership with chipmaker Cerebras, indicating that the speed gains rely on specialized hardware acceleration rather than algorithm improvements alone. Ultrafast is currently being rolled out in preview form to a limited number of customers; OpenAI has not committed to a fixed expansion date but says access will broaden as capacity grows. This phased release strategy allows the company to validate performance in production workloads before wider deployment.

OpenAI has highlighted several business use cases where Ultrafast's latency profile matters: incident response (where response delay directly affects operational impact), customer service and support (where wait time affects customer satisfaction), financial market analysis (where time-sensitive data feeds demand rapid processing), and e-commerce (where real-time personalization requires fast inference). Competitors have begun offering similar capabilities; Anthropic's Claude includes a fast mode, though OpenAI's benchmark suggests it does not yet match Ultrafast's speed envelope. The announcement underscores the intensifying competition around production-grade AI performance, with companies moving beyond raw accuracy metrics toward latency-sensitive use cases where speed directly translates to business value.

Context & Analysis

OpenAI's Ultrafast mode addresses a fundamental engineering constraint that has long shaped AI product design: the speed-versus-capability tradeoff. Historically, achieving real-time latency required deploying smaller models or ones optimized for narrow tasks, sacrificing the richness and general-purpose capability of flagship models. By achieving 14x speedup on GPT-5.6 Sol—outputting 750 tokens per second—Ultrafast suggests that infrastructure and algorithmic advances (backed by Cerebras chipmaking hardware) may be loosening this constraint.

The announcement also reflects competitive pressure from Anthropic, which has offered fast-mode variants of Claude, though at lower speeds. Notably, OpenAI frames Ultrafast not as a product for individual users seeking quicker chat responses, but as an enabling technology for high-stakes business workflows: incident response (where latency can affect mean time to resolution), customer service (where response time drives satisfaction), financial market analysis (where microseconds matter), and e-commerce (where real-time personalization drives conversion). These are domains where speed directly translates to revenue or risk mitigation, likely justifying premium pricing and early-stage limited availability.

FAQ

When will Ultrafast be available to the general public?
Ultrafast is currently in preview and available only to a small group of customers. OpenAI says it will expand access as capacity grows, but has not announced a specific timeline.
How does Ultrafast compare to Anthropic's Claude fast mode?
Ultrafast delivers 750 output tokens per second, whereas Claude's fast mode does not deliver the same level of speed.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleBloom Energy stock surges 1,920% as AI power demand drives green energy play

The AI news that matters, in one minute each morning.

Sign up free