
What happened
OpenAI published a guide to its GPT‑6 family, splitting it into GPT‑6 Astra for hardest reasoning, GPT‑6.1 Sol for complex coding and research, and GPT‑6 Luna for focused, repeated tasks.
Why it matters
The guide treats model choice and reasoning effort as an intelligence/price tradeoff, so teams are being steered by cost and latency rather than raw capability alone.
What to watch
OpenAI says cached input tokens cost up to 95% less than uncached ones, depending on the model — but only if prompt caching and compaction actually fit the workflow.
WHO IT HITSStartup engineering and product teams choosing which GPT-6 model to deploy, and the finance and platform leads who own their per-task API cost and latency budgets, are the ones this guide is written for.
Summaries like this, in your inbox every morning.
The guide is framed as a decision aid rather than a launch note: OpenAI presents GPT‑6 as a suite with three named models, and asks teams to think of model choice and reasoning level as an intelligence/price tradeoff. That framing runs through the whole document — from the model list (Astra for the hardest reasoning, Sol for complex coding and research, Luna for repeated tasks like extracting invoice fields) to the reasoning levels (low, medium, high, and extra high/max) and speed options (Fast mode and Ultrafast, the latter available for GPT‑6 Astra).
Much of the operational advice concerns cost and context rather than raw capability. OpenAI points to prompt caching, saying cached input tokens cost up to 95% less than uncached ones depending on the model, and to compaction for reducing context size in longer conversations. It also tells teams to measure task success, latency, and cost per successful task before deploying, and to plan for monitoring and data controls.
For longer work, the guide leans on features that keep a run going rather than restarting it: mid-turn steering through the Responses WebSocket API, asynchronous tool calling, and multi-agent workflows in the Responses API via GPT‑6.1 Sol, which OpenAI says is currently in beta. Whether teams actually capture the caching savings and the long-running-task gains is likely to depend on how closely their prompts, skills, and repository instructions follow the guide — a discipline question as much as a model one.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Talkdesk's survey of 252 director-level-and-above leaders found 98% use some form of AI in customer experience…
The Allen Institute for AI announced Olmo-core 3 on Thursday, a training framework it says scaled mixture-of-e…
Ben Thompson wrote that Meta's new Meta Enterprise Platform "won't work" and is "a distraction from the bigges…

Broadcom is reportedly raising $60 billion to fund chips for Anthropic, on top of a loan of up to $42 billion…

Meta is letting go of employees it hired from AI safety startup Virtue AI, four months after they joined in Ju…

Bain & Company found the industry's compute demand will require $6 trillion in annual revenue by 2031, includi…
