
What happened
Moonshot AI's Kimi K3 is now on Amazon Bedrock, with a 1-million-token context window, native vision, and an approximate 2.5x gain in scaling efficiency over Kimi K2.
Why it matters
Kimi K3 is the first open-weight model on Bedrock with explicit prompt caching, which is designed to cut latency and input costs when context is reused across calls.
What to watch
The value hinges on whether reused-prompt workloads actually hit the cache — cached tokens are billed higher upfront and kept for at least 30 minutes.
WHO IT HITSEnterprise AI platform teams running long-context coding and knowledge workflows on Amazon Bedrock are the first to be affected, since caching and cross-Region profiles change latency and input-token billing for those workloads.
Summaries like this, in your inbox every morning.
Moonshot AI's Kimi K3 arrives on Amazon Bedrock as AWS continues to widen its catalog of open-weight models. The post notes that since 2025, Bedrock has added dozens of open-weight models from providers including DeepSeek, Google, MiniMax, Mistral AI, Moonshot AI, NVIDIA, OpenAI, and Qwen. In 2026, Bedrock also added platform-level capabilities such as tool calling, structured output, reasoning, response streaming, and the Responses and Chat Completions APIs. Because those are platform features rather than per-model integrations, newer open-weight models can pick them up as they land on Bedrock.
On the model itself, the post positions Kimi K3 for long-running coding and knowledge workflows that need sustained context across large repositories, documents, and images. It pairs native vision with a 1-million-token context window and an approximate 2.5x improvement in scaling efficiency over Kimi K2, and it is described as the first open-weight model on Bedrock to support explicit prompt caching — a mechanism that lets requests reuse a marked prompt prefix instead of resending it. For teams weighing cost and latency against data control, the post also notes that data is processed within the AWS data boundary, is not shared with the model provider, and is not used for training, with zero data retention always enabled for inference requests.
What that amounts to in practice is likely to depend less on the headline capability than on how well prompt-caching fits a given workload. Explicit caching charges more for tokens written to cache and keeps them for at least 30 minutes, so the payoff hinges on whether a team reuses the same long, stable context often enough to hit the cache before it expires. The choice between the global and US profiles adds a second trade-off — AWS says the global profile costs about 10% less than a geographic one, while the US profile keeps processing within US geography for data-residency requirements. Buyers with strict residency or cost rules may find those two dials, rather than raw model specs, determine whether Kimi K3 earns a place in production.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Reuters reported that Anthropic PBC opened a wet lab in the San Francisco Bay Area, where robots powered by it…
Micron Technology executives said memory has become a core bottleneck setting the upper limits of AI performan…

Meta's personal AI agent Muse reached nearly 600,000 downloads in its first five days, according to SensorTowe…

The Zacks Research Daily featured new reports on AMD, Linde and Amgen, whose shares have outperformed their in…

The FAA awarded an $875 million, 12-year contract to Air Space Intelligence in June for SMART, an AI tool to h…

On Wednesday, US government officials removed a Chinese Alibaba Qwen AI search tool from the Federal Register…
