
Amazon Bedrock now enables cross-Region inference for OpenAI GPT-5.6 models across more than 25 AWS Regions, allowing requests to route to whichever Region has available capacity.
The launch includes geographic profiles (restricted to specific regions like the US) and global profiles (spanning all supported AWS commercial Regions), helping organizations balance performance, capacity, and data residency requirements.
Developers can call the models through the OpenAI SDK or Amazon Bedrock's native APIs.
What happened
Amazon Bedrock now supports OpenAI GPT-5.6 models (Sol, Terra, and Luna variants) across more than 25 AWS Regions with cross-Region inference capabilities. The service offers both geographic inference profiles (keeping data within a specific region like the US) and global profiles (routing requests across all supported AWS commercial Regions based on real-time capacity).
Why it matters
Cross-Region inference works as a capacity mechanism—by routing requests to a broader pool of compute rather than limiting them to a single Region, it improves throughput and maintains consistent performance under load. Organizations with data residency requirements can use geographic profiles to scale while staying within specific geographic boundaries, while others can use global profiles for maximum capacity access.
What to watch
Developers can access the models via the Amazon Bedrock console playground (requiring no coding), the OpenAI Responses API, the OpenAI Chat Completions API, or the Amazon Bedrock Converse API. All three GPT-5.6 variants support a 1 million token context window, reasoning mode, server-side tool calling, and prompt caching.
Ask the AI about this article →
Amazon's launch of cross-Region inference for GPT-5.6 addresses a core operational challenge for enterprises: balancing model capacity with data governance. Cross-Region inference profiles act as a routing layer, allowing requests to draw from a shared pool of compute across multiple AWS Regions rather than exhausting capacity in a single Region. This is particularly valuable for organizations operating at scale or during traffic spikes, as it improves throughput without requiring pre-provisioning of capacity in every Region.
The dual-profile approach reflects different customer priorities. Geographic profiles serve organizations subject to data residency regulations (such as keeping data within the US or within Europe), while global profiles unlock the full capacity footprint for those without such constraints. Billing and quota remain tied to a single account view regardless of which backend Region processes the request, simplifying cost tracking and rate limits. The integration with OpenAI's API formats (Responses API and Chat Completions API) lowers the barrier to adoption for teams already using OpenAI models—developers can point their existing SDK clients at Amazon Bedrock's OpenAI-compatible endpoint and swap in the inference profile ID without rewriting application logic.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic is privately hoping to file for its initial public offering by the end of this month, targeting a ra…
Broadcom is reportedly seeking to borrow up to $100 billion in debt financing to support growth efforts at Ant…
As AI technology matures, the bottleneck in the industry is moving beyond semiconductor constraints like GPUs…

Elice Group, a South Korean AI infrastructure provider, announced the launch of the country's first AI data ce…

On August 12, AT&T's Chief Data and AI Officer said OpenAI models power about 25% of the telecom's total AI us…

On August 11, IBM announced a multi-year $240 million agreement with Together AI to deploy NVIDIA HGX B300 sys…
