
Amazon Bedrock now offers OpenAI GPT-5.6 models in India, keeping data within the country.
Financial services, healthcare, and the public sector can use these models at scale.
Both models support a 1-million-token context window and text-image input.
What happened
Amazon Bedrock now supports OpenAI GPT-5.6 models Terra and Luna in India, with India geographic cross-Region inference. Requests route only between the Mumbai and Hyderabad Regions, so processing stays within the country.
Why it matters
This lets organizations with local data processing requirements, such as financial services, healthcare, and the public sector, use these OpenAI models at scale without data leaving India. Both models offer a 1-million-token context window and accept text and image input, enabling processing of long documents and large code bases in a single request.
What to watch
India geographic inference profiles (in.openai.gpt-5.6-terra and in.openai.gpt-5.6-luna) keep inference within India. Prompt caching with these profiles offers a 90 percent discount on cached reads compared to uncached input tokens, and cached prefixes stay warm for at least 30 minutes.
Ask the AI about this article →
This announcement extends Amazon Bedrock's OpenAI GPT-5.6 model support to India with a geographic data-residency boundary. Previously, cross-Region inference for these models could route requests globally; now, organizations with strict local processing requirements can use the India profiles to keep inference within the country, routing only between the Mumbai and Hyderabad Regions. The move is particularly relevant for regulated industries like financial services and healthcare, where data sovereignty is a compliance concern.
Technically, the India geographic profiles work through inference profiles, which Amazon Bedrock uses to define routing constraints. Calls to these profiles track billing and quota against the source Region, and CloudWatch and CloudTrail logs also stay in the source Region, simplifying monitoring. Notably, the models support a 1-million-token context window and multimodal input, enabling complex workloads like invoice extraction or long-document processing without leaving the country.
For developers, Amazon Bedrock supports both OpenAI-compatible APIs (Responses and Chat Completions) and the Bedrock-native Converse API. Prompt caching, available under the India profiles, offers a 90 percent discount on cached reads and keeps cached prefixes warm for at least 30 minutes, which could reduce costs for RAG and agent workloads that repeat context. This combination of data residency, API flexibility, and cost-saving features may encourage more enterprise adoption of these models within India, though the actual business impact will depend on how organizations weigh performance against compliance requirements.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
DataDirect Networks (DDN), with Super Micro Computer and Solidigm, introduced DDN Enterprise AI HyperPOD, buil…
Zep built Konig, a graph storage system that keeps hot graphs in RAM, recently used graphs on local NVMe, and…

Nvidia is set to acquire AI model repository Hugging Face for $13 billion, as reported by Ars Technica

Runable Inc., a platform using AI agents to help businesses build, run and grow, announced Wednesday it raised…
AI shopping agents tested by Wharton School researchers changed product picks by up to 99 percentage points wh…

Google updated Gemini Omni Flash to version 1.1, improving scene extension to analyze up to ten seconds of vid…
