AIToday
Large Language ModelsAI Coding AssistantsAmazon AI BlogPublished: Oct 6, 2026, 01:00 JST

Amazon Bedrock Managed Knowledge Bases: agentic hits 6 of 6

Amazon Bedrock Managed Knowledge Bases: agentic hits 6 of 6

3 Key Points

  1. What happened

    AWS built a LangChain RAG app on Amazon Bedrock Managed Knowledge Bases and ran one six-intent question through both paths; standard retrieval at 5 results covered 4 of 6 sub-intents, while the other path covered 6 of 6.

  2. Why it matters

    This shows that when a single query hides six intents, one embedding can miss sub-intents, so teams may need to weigh higher cost and latency against fuller coverage. Hedged: the trade-off hinges on query shape.

  3. What to watch

    The result hinges on whether production traffic is mostly single-intent or multi-part, since the cheaper path is recommended for short, well-scoped questions; watch maxAgentIteration — below 4 the planner stops decomposing.

WHO IT HITSThis lands on enterprise RAG teams and application developers running self-hosted or managed retrieval on Amazon Bedrock who must choose between one hybrid search and a planning loop per query.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Amazon Bedrock Managed Knowledge Bases removes the self-managed vector store, embeddings, and re-ranking models from the RAG architecture; the walkthrough uses an S3 bucket as the data source, and the API takes no storageConfiguration, which the post calls the clearest signal that Amazon handles the storage layer. From there the same complex query can go through two paths: a standard LangChain retriever wrapping the Retrieve API, or a planning loop exposed as agentic_retrieve and the lower-level agentic_retrieve_stream call. The post notes friction in the current integration — agentic retrieval is a function rather than a LangChain retriever, so it needs a RunnableLambda to sit in a chain, and trace events that show the query plan require a direct boto3 call. For context, AWS evaluated agentic retrieval on MuSiQue, a public multi-hop benchmark, and reported improved recall over single-shot retrieval, with the largest gains on the hardest questions and single-hop questions seeing gains under five points. The post also notes that agentic retrieval registers up to five knowledge bases in one request and routes sub-queries using a natural-language description attached to each, which the other API cannot do. The outcome likely hinges on whether a team's own query mix is dominated by single-intent questions, since the recommendation is to route on query shape rather than pick one path for everything, and measuring that mix appears to be the intended next step.

FAQ
What is the difference between the two retrieval paths?
The Retrieve API runs one hybrid search and returns scored chunks, while the AgenticRetrieveStream API runs a planning loop that breaks the question into sub-queries and can search again if evidence is insufficient.
When does the agentic path start breaking a question into sub-queries?
maxAgentIteration defaults to five, but at two or three the planner emits no sub-queries; decomposition begins at four.
Does agentic retrieval return relevance scores like standard retrieval?
No. Standard retrieval gives each chunk a typed score field, but agentic results carry content, metadata, and sourceRetriever with no equivalent typed field, so code reading result["score"] gets nothing.
Amazon AI BlogRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleCohere's North 2 adds token caps, rebuilt agent orchestration