AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Jul 31, 2026, 04:02 JST4 min read

OpenAI cuts Luna model price 80%, matches year-old rivals at 6 cents per task

OpenAI cuts Luna model price 80%, matches year-old rivals at 6 cents per task

Key takeaway

  • OpenAI has slashed prices for its GPT-5.6 Luna model by 80 percent, bringing the cost of running a task down to about 6 cents—nearly nine times cheaper than leading models from a year ago—effective July 30.

  • The company says the cuts are enabled by efficiency gains in its own Sol model, which optimized GPU software and improved token generation.

  • The aggressive pricing reflects growing competition from Chinese providers and Microsoft's push for lower-cost alternatives, though sustained price wars could strain the finances of frontier AI labs with heavy infrastructure costs.

3 Key Points

  1. What happened

    OpenAI is cutting GPT-5.6 Luna prices by 80 percent effective July 30, dropping to $0.20 per million input tokens and $1.20 per million output tokens. Terra pricing falls 20 percent to $2 and $12 per million tokens respectively, while Sol stays the same. Luna matches the performance of leading models from a year ago, but a task that cost a dollar with those models now runs about 6 cents on Luna, nearly nine times faster.

  2. Why it matters

    OpenAI says the cuts are possible because GPT-5.6 Sol optimized GPU software on its own, cutting deployment costs by 20 percent and improving token generation by more than 15 percent through speculative decoding. Growing price pressure from low-cost Chinese providers is likely a factor—Microsoft is now openly promoting its own MAI models as cheaper alternatives to OpenAI. A prolonged price war could hurt the broader market if it slows revenue growth at frontier labs whose balance sheets are tied to massive infrastructure investments.

  3. What to watch

    All three models (Luna, Terra, and Sol) are available through ChatGPT Work, Codex, and the OpenAI API starting July 30. The move may signal whether OpenAI can sustain profitability while competing on cost rather than capability alone.

In Depth

Read the full story

OpenAI announced significant price reductions across its GPT-5.6 model lineup, effective July 30. Luna, the company's smallest and most affordable model, will see an 80 percent price cut, dropping to $0.20 per million input tokens and $1.20 per million output tokens. Terra will drop 20 percent to $2 and $12 per million tokens respectively, while Sol pricing remains flat. All three models remain accessible via ChatGPT Work, Codex, and the OpenAI API.

The centerpiece of the announcement is Luna's price-to-performance proposition. OpenAI positions Luna as a model whose capabilities match those of leading competitors from a year ago, yet at dramatically lower cost. The company quantifies this gap: a task that cost one dollar using those legacy models now costs about 6 cents on Luna—a 94 percent savings and nearly nine times faster execution. This framing allows OpenAI to offer low-cost inference without claiming to have built a new frontier model.

OpenAI attributes the price cuts to efficiency gains in its own infrastructure, particularly driven by GPT-5.6 Sol. According to the company, Sol optimized GPU software on its own, reducing deployment costs by 20 percent. Additionally, Sol improved token generation speed by more than 15 percent through speculative decoding—a technique that generates multiple candidate tokens in parallel and discards incorrect ones, thereby reducing the number of inference steps required. These improvements in model efficiency translate directly into lower per-token costs that OpenAI can pass to customers.

The broader context is intensifying competition in the AI inference market. Chinese providers have established themselves as low-cost competitors, and Microsoft is now actively marketing its own MAI models as cheaper alternatives to OpenAI's offerings. The article notes that this price war carries systemic risk: if frontier AI labs—whose research and development are funded by high-margin inference revenue—see their margins compressed, they may struggle to sustain the massive infrastructure investments required to train next-generation models. OpenAI's move suggests the company believes it can maintain profitability at lower prices through operational efficiency, but whether that holds across the entire sector remains uncertain.

Context & Analysis

OpenAI's aggressive 80 percent price cut on Luna reflects a fundamental shift in how the company is competing in the AI market. Rather than racing purely on capability, OpenAI is now emphasizing price-to-performance—positioning Luna as a model that delivers year-ago performance levels at a fraction of the cost. The company attributes this move to internal efficiency gains: Sol's ability to optimize its own GPU software (reducing deployment costs by 20 percent) and improve token generation by more than 15 percent through speculative decoding. This suggests that frontier labs are finding ways to cut costs not just by using cheaper hardware, but by refining the software layer itself.

The timing and scale of the cut signal heightened competition. Microsoft's recent push to market MAI models as cheaper alternatives to OpenAI, combined with pressure from low-cost Chinese providers, has forced OpenAI's hand. Luna's new price point—6 cents per task versus a dollar on legacy models—is designed to capture price-sensitive customers and prevent defection to rivals. However, the article hints at a longer-term risk: if this price war spreads across the frontier AI lab sector, it could erode the revenue growth that underpins massive infrastructure investments. Companies like OpenAI rely on high-margin inference revenue to fund their next-generation training runs, so sustained margin compression could slow the pace of AI development industry-wide.

FAQ

When do the price cuts take effect?
The price cuts are effective July 30. Luna drops to $0.20 per million input tokens and $1.20 per million output tokens, while Terra falls to $2 and $12 per million tokens. Sol pricing remains unchanged.
How did OpenAI make these price cuts possible?
OpenAI says GPT-5.6 Sol optimized GPU software on its own, cutting deployment costs by 20 percent, and improved token generation by more than 15 percent through speculative decoding.
How does Luna's performance compare to older models?
Luna matches the performance of leading models from a year ago, but a task that cost a dollar with those models now runs about 6 cents on Luna, nearly nine times faster.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Related Articles

Next articleMicrosoft ships Agent Framework–Squad integration for multi-agent teams

The AI news that matters, in one minute each morning.

Sign up free