AIToday
Large Language ModelsAI Business & IndustryLatent SpacePublished: Jul 31, 2026, 16:00 JST4 min read

GPT 5.6 prices cut 20–80%; intelligence cost drops 13× in 4 months

GPT 5.6 prices cut 20–80%; intelligence cost drops 13× in 4 months

Key takeaway

  • OpenAI slashed GPT-5.6 model prices by 20–80% and introduced a 2.5× faster Sol mode, driven by backend efficiency improvements including self-optimization, speculative decoding, and prompt caching.

  • The cuts mean GPT-5.4's flagship intelligence (scored 51 on benchmarks) now costs roughly one-thirteenth what it did four months ago, undercutting open competitors on cost per task and signaling continued steep deflation in AI inference pricing.

3 Key Points

  1. What happened

    OpenAI cut GPT-5.6 Luna prices by 80% (now $0.20 per million input tokens and $1.20 per million output tokens) and Terra by 20% (to $2/$12), while launching Sol Fast mode at 2.5× speed for 2× the price. The cuts follow systemic efficiency gains: self-optimization reduced serving costs by 20%, speculative decoding improved token efficiency by over 15%, and optimizations to KV caching, context bloat avoidance, and prompt caching lowered infrastructure demands.

  2. Why it matters

    GPT-5.4 full intelligence (scored 51 on Artificial Analysis Cost per Task benchmark) cost $2.50/$15 four months ago; that same intelligence level is now available in Luna at $0.20/$1.20—roughly one-thirteenth the token price. For workflows like auto-review in ChatGPT and Codex CLI moving from GPT-5.4 to Luna, OpenAI projects roughly 10× lower cost. This undercuts open models like DeepSeek and specialized cheap models like Gemini Flash-Lite on cost-per-task.

  3. What to watch

    OpenAI's pricing gains reflect sustained deflationary pressure on commodity AI inference. The body notes an annualized rate of roughly 2000× cost reduction per year (though benchmarks like Artificial Analysis may be partially trained-to), and the article frames this as an acceleration beyond the 1000× drop over 18 months observed a year prior. Sol Fast mode is available in the API; Luna and Terra cuts are effective immediately.

In Depth

Read the full story

On July 30, 2026, OpenAI announced steep price cuts for GPT-5.6 models and a faster inference mode, the culmination of months of backend optimization. Sam Altman posted the headline numbers: GPT-5.6 Luna dropped 80% to $0.20 per million input tokens and $1.20 per million output tokens; Terra fell 20% to $2/$12; and Sol gained a Fast mode offering up to 2.5× speed for 2× the standard price, with "no change in intelligence."

The cuts emerged from a detailed systems optimization effort spanning three domains. For self-optimization, OpenAI used GPT-5.6 Sol itself to analyze production traffic, tune load balancing, and autonomously rewrite production kernels in Triton and Gluon, reducing end-to-end serving costs by 20%. On speculative decoding—a technique where a smaller model drafts tokens and a larger model validates them—GPT-5.6 designed and ran hundreds of experiments, testing changes in size, structure, and features while monitoring training and autonomously intervening during hardware failures and training instability, increasing token-generation efficiency by over 15%. For KV caching (a memory optimization for inference), the team optimized batching, sharding, and cache management for different workloads, especially Sol in Codex. To streamline multi-step agentic tasks, OpenAI optimized the Rust orchestration layer, implementing deferred tool discovery (tools surface only when needed) and capping tool outputs at 10,000 tokens to prevent context bloat. Prompt caching treats all model-visible history as append-only, preserving the prompt prefix and allowing the system to reuse previously computed data at high cache-hit rates.

The market impact was sharp. An observer noted that GPT-5.4 full at "xhigh" quality scored 51 on Artificial Analysis Cost per Task, exactly where Luna max now sits. GPT-5.4 cost $2.50/$15; Luna now costs $0.20/$1.20—roughly one-thirteenth the token price, achieved in approximately four months. The article frames this as an annualized rate of roughly 2000× cost reduction per year, with the caveat that public benchmarks may be partially trained-to. The downstream effect is immediate: auto-review in ChatGPT and Codex CLI is migrating from GPT-5.4 to Luna, with OpenAI expecting roughly 10× lower cost. The article notes that on cost-per-task measures, OpenAI has surpassed open models like DeepSeek, GLM, and MiniMax, as well as specialized cheap-but-good models like Gemini Flash-Lite, making proprietary non-finetuned inference the lowest-cost option available.

Context & Analysis

The article traces a striking pattern of sustained cost reduction in AI inference. A year ago, holding model quality constant (measured by LMSys Elo), GPT-4-level intelligence fell 1000× in cost over 18 months; the new data suggest this deflation has not only continued but accelerated. The specific claim—that GPT-5.4 intelligence now costs roughly one-thirteenth what it did four months earlier—translates to an annualized rate the article frames as "roughly 2000× a year," though it hedges this with the caveat that public benchmarks like Artificial Analysis may be partially trained-to, whereas Elos resist direct optimization.

The price cuts are tied directly to systems-level improvements across three layers: the model itself (self-optimization where GPT-5.6 analyzed production traffic and autonomously rewrote serving kernels), the inference stack (speculative decoding and KV caching tuning), and the agentic harness (context bloat avoidance and prompt caching). This multi-layer optimization appears to reflect a shift from low-hanging fruit (moving from completions to reasoning, from dense to mixture-of-experts) to harder architectural and infrastructure wins. The practical impact is immediate: teams using Codex and ChatGPT are moving auto-review from GPT-5.4 to Luna with an expected 10× cost reduction.

The article also notes that on cost-per-task benchmarks, OpenAI has "beat even open models like DeepSeek, GLM, and MiniMax, and dedicated cheap-but-good models like Gemini Flash-Lite," positioning proprietary inference as the cheapest option for non-finetuned intelligence—a reversal from prior periods when open models offered better unit economics.

FAQ

Which GPT-5.6 models got price cuts and by how much?
GPT-5.6 Luna was cut 80% to $0.20 per million input tokens and $1.20 per million output tokens; GPT-5.6 Terra dropped 20% to $2/$12. GPT-5.6 Sol also got a new Fast mode offering up to 2.5× speed for 2× the standard price with no change in intelligence.
How much cheaper is GPT-5.4 intelligence now compared to four months ago?
GPT-5.4 full at high quality scored 51, the same benchmark level Luna now reaches. GPT-5.4 cost $2.50/$15; Luna now costs $0.20/$1.20—roughly one-thirteenth the token price.
What technical changes enabled the price cuts?
OpenAI cited self-optimization (reducing serving costs by 20%), speculative decoding (improving token-generation efficiency by over 15%), optimized KV caching and batching, avoiding context bloat through deferred tool discovery and 10,000-token output caps, and prompt caching to reuse previously computed data.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Related Articles

Next articleSemco expands SoulBrain partnership for AI chip packaging materials

The AI news that matters, in one minute each morning.

Sign up free