
OpenAI cut prices for two of its three GPT-5.6 API tiers on July 30, 2026: Luna dropped 80% to $0.20/$1.20 per million tokens, and Terra fell 20% to $2/$12 per million tokens, while Sol pricing stayed the same.
The company also introduced Fast mode for Sol, delivering up to 2.5× faster responses at 2× the cost.
These moves reflect efficiency gains in the models, inference serving, and agentic workflows, and make it practical for developers to run high-volume AI agents and multi-step tasks at scale without prohibitive token costs.
What happened
On July 30, 2026, OpenAI reduced GPT-5.6 Luna pricing by 80% to $0.20 per million input tokens and $1.20 per million output tokens, and cut GPT-5.6 Terra by 20% to $2 per million input tokens and $12 per million output tokens. The flagship Sol model kept its existing price. OpenAI also replaced Priority Processing with Fast mode for Sol, which delivers responses up to 2.5 times faster at twice the Standard rate.
Why it matters
The price drops make high-volume agentic work—multi-step workflows, agents using tools—economically viable at scale for the first time. Luna now costs roughly 6 cents on the dollar per task compared to models that were frontier-class a year ago, and on Agents' Last Exam (a professional agentic benchmark), Luna outperforms Fable 5 at an estimated cost per task nearly 99% lower. For developers, the wider price-performance range lets them match the right model tier to each step of a workflow instead of defaulting to the flagship.
What to watch
Fast mode is available now in the API and via the /fast command in Codex; it is backward compatible, so existing integrations using Priority Processing automatically route to Fast mode without code changes. Pricing takes effect July 30, 2026, with AWS rollout shortly after. Lower Terra and Luna costs also reduce credit consumption in Codex and ChatGPT Work subscriptions, stretching the same plan further without changing subscription prices or quota budgets.
On July 30, 2026, OpenAI announced significant pricing reductions and a new processing mode for its GPT-5.6 model family. Luna, the fastest and most affordable tier, now costs $0.20 per million input tokens and $1.20 per million output tokens—an 80% reduction that makes it one of the cheapest capable models on the market. Terra, the balanced everyday model, drops 20% to $2 per million input tokens and $12 per million output tokens. Sol, the flagship, retains its existing pricing.
The price reductions flow through OpenAI's subscription products as well. In Codex and ChatGPT Work, Terra and Luna now consume fewer credits against a user's quota, allowing the same subscription plan to stretch further without any change to subscription prices or quota budgets themselves.
OpenAI also introduced Fast mode for Sol, replacing the previous Priority Processing option. Fast mode delivers responses up to 2.5 times faster than Standard processing at twice the base price, with no change in model intelligence. The feature is backward compatible—existing API requests tagged as priority automatically route to Fast mode without requiring code rewrites. This aligns with a pattern Anthropic has already established with Claude Opus 4.8, which also runs roughly 2.5 times faster at twice the base rate.
The efficiency gains behind the price cuts stem from improvements across three layers: the models themselves, the inference systems serving them, and the agentic harness. Better routing keeps hardware busy, optimized serving software generates tokens more cheaply, and smarter context management prevents agents from repeating completed work. A particularly notable development is that GPT-5.6 Sol is now helping optimize itself. Within a human-led process, Sol autonomously rewrote and optimized production kernels and ran hundreds of experiments to improve token generation, reducing the end-to-end cost of serving the model by about 20% and increasing token-generation efficiency by more than 15%.
For developers, the broader implication is strategic. Rather than defaulting every call to the flagship Sol or assuming the cheapest tier is always right, teams can now match intelligence, speed, and cost to each workflow step. Luna is positioned for high-volume, cost-sensitive work; Terra for balanced quality and cost; and Sol for the most demanding reasoning tasks. On Agents' Last Exam, a benchmark for professional agentic work, OpenAI reports that Luna outperforms Fable 5 at an estimated cost per task nearly 99% lower. According to Artificial Analysis, Luna also sits well ahead of similarly priced models on its Intelligence Index. The practical pattern emerging is that cheaper Luna and Terra pricing makes it realistic to run agentic features at scale—where a coding agent, for example, might use Sol to resolve uncertainty and write a plan, then hand well-specified changes to Luna to implement, test, and evaluate.
The price cuts on Luna and Terra are not a margin decision but a result of efficiency gains across the entire stack. OpenAI attributes the improvements to optimizations in the models themselves, the inference systems serving them, and the agentic harness that connects them to tools and context. A notable feedback loop is now at work: GPT-5.6 Sol autonomously rewrote and optimized production kernels and ran hundreds of experiments, reducing the end-to-end cost of serving the model by about 20% and increasing token-generation efficiency by more than 15%. This creates a virtuous cycle in which better autonomous capability accelerates the discovery of the next round of cost savings, which then cascade to customers as lower prices.
The introduction of Fast mode mirrors a pattern Anthropic established with Claude Opus 4.8—allowing developers to pay a premium only when latency is the constraint, rather than blanket speed upgrades. For developers building agentic systems, the new pricing unlocks a multi-tier strategy: use Sol for high-stakes reasoning steps that require maximum intelligence, hand well-specified work to Luna for implementation and testing, and reserve Terra for balanced everyday tasks where quality is important but speed matters too. The real leverage comes not from a single model becoming cheaper, but from having a wider range of price-performance points to match against each workflow step.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Supermicro reported a sharp recovery in gross margins and issued fiscal 2027 revenue guidance that exceeded Wa…

IBM and Together AI announced a collaboration to deploy a dedicated Nvidia-powered inference cluster on IBM Cl…

Siliconware Precision Industries (SPIL), a subsidiary of ASE Technology Holding, held a groundbreaking ceremon…

CoreWeave reported second-quarter fiscal 2026 revenue of US$2.6 billion, up 112% year-over-year and 24% sequen…

CoreWeave, an AI cloud provider, more than doubled its second-quarter revenue while managing a $104 billion ba…

Palantir Technologies reported second-quarter revenue up 93% year over year to $1.94 billion on Aug

The AI news that matters, in one minute each morning.
Sign up free