
What happened
OpenAI's three GPT-5.6 models—Sol (for reasoning tasks), Terra (for general production), and Luna (for fast inference)—are now generally available on Amazon Bedrock. All three support a 272K-token context window, text and image input, and reasoning effort levels from none to max, and are accessible through the OpenAI Responses API on the bedrock-mantle endpoint.
Why it matters
Developers can now use frontier OpenAI models through AWS infrastructure without managing separate model infrastructure, while staying within their VPC, AWS IAM policies, and regional data residency requirements. Pricing matches OpenAI's first-party rates, and usage counts toward existing AWS commitments.
What to watch
Sol and Terra are available in US East (N. Virginia) and US East (Ohio); Luna is also available in US West (Oregon). GPT-5.6 supports prompt caching (implicit by default, explicit for fine control) to reduce costs in repeated or multi-step workloads. Developers can authenticate using auto-refreshing short-term keys from AWS credentials or environment variables, and the new Amazon Bedrock console offers side-by-side model comparison before writing code.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
OpenAI's release of its GPT-5.6 family on Amazon Bedrock represents a direct integration of frontier models into AWS's managed infrastructure, eliminating the need for developers to operate separate model endpoints. The three-tier approach—Sol for reasoning-heavy tasks, Terra for balanced production work, and Luna for high-throughput, latency-critical inference—allows teams to right-size both capability and cost for different workloads within a single API ecosystem.
The integration preserves AWS's security model: all requests operate within customers' IAM policies, VPCs, and CloudTrail logging, with optional regional residency. This addresses a key concern for enterprise workloads subject to data-governance requirements. Pricing parity with OpenAI's first-party offering, combined with credit toward existing AWS commitments, removes a common barrier to adoption. The prompt caching feature (available in implicit and explicit modes) directly targets the economics of agentic and multi-turn reasoning workloads, which repeat context between calls—a pattern that will likely dominate as developers scale autonomous systems.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Copado extended its Agentia AI DevOps platform with Headless, which uses Model Context Protocol and command-li…
The United Arab Emirates said it plans to invest EUR40 billion, or about US$46 billion, in Germany across arti…

After Anthropic CEO Dario Amodei's Sept 12 essay urged AI companies to slow capability improvements, AI-relate…

Investor Michael Burry challenged recent calls by OpenAI, Anthropic and other large AI companies to slow AI de…

Fast Company named Architech to its eighth annual Best Workplaces for Innovators list in AI, marking a second…

AI researcher Jacob Coxon resigned from Anthropic and posted on X that AI could kill us all by the end of the…
