AIToday
Large Language ModelsOpen-Source AIAI Business & IndustryAmazon AI BlogPublished: Sep 19, 2026, 04:00 JST

Kimi K3 hits Amazon Bedrock with 2.8T params

Kimi K3 hits Amazon Bedrock with 2.8T params

3 Key Points

  1. What happened

    Moonshot AI's Kimi K3 is now on Amazon Bedrock, with a 1-million-token context window, native vision, and an approximate 2.5x gain in scaling efficiency over Kimi K2.

  2. Why it matters

    Kimi K3 is the first open-weight model on Bedrock with explicit prompt caching, which is designed to cut latency and input costs when context is reused across calls.

  3. What to watch

    The value hinges on whether reused-prompt workloads actually hit the cache — cached tokens are billed higher upfront and kept for at least 30 minutes.

WHO IT HITSEnterprise AI platform teams running long-context coding and knowledge workflows on Amazon Bedrock are the first to be affected, since caching and cross-Region profiles change latency and input-token billing for those workloads.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

Moonshot AI's Kimi K3 arrives on Amazon Bedrock as AWS continues to widen its catalog of open-weight models. The post notes that since 2025, Bedrock has added dozens of open-weight models from providers including DeepSeek, Google, MiniMax, Mistral AI, Moonshot AI, NVIDIA, OpenAI, and Qwen. In 2026, Bedrock also added platform-level capabilities such as tool calling, structured output, reasoning, response streaming, and the Responses and Chat Completions APIs. Because those are platform features rather than per-model integrations, newer open-weight models can pick them up as they land on Bedrock.

On the model itself, the post positions Kimi K3 for long-running coding and knowledge workflows that need sustained context across large repositories, documents, and images. It pairs native vision with a 1-million-token context window and an approximate 2.5x improvement in scaling efficiency over Kimi K2, and it is described as the first open-weight model on Bedrock to support explicit prompt caching — a mechanism that lets requests reuse a marked prompt prefix instead of resending it. For teams weighing cost and latency against data control, the post also notes that data is processed within the AWS data boundary, is not shared with the model provider, and is not used for training, with zero data retention always enabled for inference requests.

What that amounts to in practice is likely to depend less on the headline capability than on how well prompt-caching fits a given workload. Explicit caching charges more for tokens written to cache and keeps them for at least 30 minutes, so the payoff hinges on whether a team reuses the same long, stable context often enough to hit the cache before it expires. The choice between the global and US profiles adds a second trade-off — AWS says the global profile costs about 10% less than a geographic one, while the US profile keeps processing within US geography for data-residency requirements. Buyers with strict residency or cost rules may find those two dials, rather than raw model specs, determine whether Kimi K3 earns a place in production.

FAQ
How much cheaper is the global cross-Region inference profile?
Amazon says global cross-Region inference costs approximately 10% less than a geographic profile.
Does my data get shared with Moonshot AI or used for training?
AWS says your data is processed within the AWS data boundary, is not shared with the model provider, and is not used to train the underlying model.
What is the minimum prompt size for caching?
You can mark the end of a reusable prompt prefix after at least 1,024 tokens using a prompt_cache_breakpoint.
Amazon AI BlogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Anthropic opens AI-powered wet lab in San Francisco Bay AreaSiliconANGLE AI · 28m ago
  • Meta's Muse hits nearly 600,000 downloads in first 5 daysYahoo Finance AI · 28m ago
  • Federal Register site drops Alibaba Qwen search toolArs Technica AI · 28m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSam Altman, Elon Musk back Amodei's AI slowdown call