AIToday
Large Language ModelsAI Business & IndustryAmazon AI BlogPublished: Sep 30, 2026, 13:00 JST

Bedrock adds in-region inference for Claude Opus 5, Sonnet 5

Bedrock adds in-region inference for Claude Opus 5, Sonnet 5

3 Key Points

  1. What happened

    Amazon Bedrock added in-region inference for Claude Opus 5 and Claude Sonnet 5 in Seoul (ap-northeast-2) and Claude Sonnet 5 in Singapore (ap-southeast-1). Requests are served by that single Region alone.

  2. Why it matters

    Under strict single-Region processing, prompts and outputs stay inside the Region a company calls for the full request lifecycle.

  3. What to watch

    The trade-off hinges on whether a single Region's capacity and quotas can support an application at scale, since there is no routing layer to fall back on. Billing follows standard on-demand pricing for the Region called.

WHO IT HITSEnterprise teams with strict single-Region data processing rules, such as in financial services, healthcare, and the public sector, can now run these models where processing stays inside the Region they call.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The update puts a specific kind of guarantee in front of customers: a request sent to the Seoul (ap-northeast-2) or Singapore (ap-southeast-1) Region is served by that Region alone, with no routing layer like the one used by cross-Region inference profiles. Input prompts and output results stay inside it for the full lifecycle of the request. That matters most to organizations whose data residency rules are strict enough that multi-Region routing is not an option, and the post names financial services, healthcare, and the public sector as examples.

The trade-off is stated just as plainly. Because the request cannot be routed elsewhere, throughput is bounded by that Region's capacity, and requests are subject to per-Region service quotas. Billing follows standard on-demand pricing for the Region you call, and the operational footprint is deliberately narrow: quota consumption, CloudWatch metrics, and CloudTrail log entries are all scoped to that same Region, so there is no source-versus-destination distinction to track in monitoring. For teams already using Amazon Bedrock Guardrails and intelligent prompt routing, the post says those features remain available alongside in-region inference.

Whether this is the right fit comes down to how a team weighs residency against headroom. The guarantee that data never leaves the Region is what some regulated workloads need to proceed, but it also means the ceiling moves with that Region's capacity rather than being smoothed across a wider footprint. Teams with spiky or very large demand may find the per-Region quotas to be the deciding constraint, while those whose main obstacle was data residency get a path to run these models at scale.

FAQ
Which models and Regions does this cover?
Claude Opus 5 and Claude Sonnet 5 in Seoul (ap-northeast-2), and Claude Sonnet 5 in Singapore (ap-southeast-1), available on the bedrock-runtime endpoint.
How is this different from cross-Region inference profiles?
There is no routing layer; the request is served by the single Region alone, and prompts and outputs stay within it for the full lifecycle of the request.
How is usage billed and monitored?
Billing follows standard on-demand pricing for the Region you call, and quota consumption, CloudWatch metrics, and CloudTrail logs are all scoped to that same Region.
Amazon AI BlogRead Original Article

AI news that matters for your work, in one minute a day

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleEliseAI raises $350 million, valued at $4 billion