Monday, 05 October 2026 PDT | 03:58 PM
The 1 News Alt Logo Text Smart News for Global Indians

Agentic retrieval with LangChain and Amazon Bedrock Knowledge Bases

Finance October 05, 2026 11:00 PM
Agentic retrieval with LangChain and Amazon Bedrock Knowledge Bases

When a user asks the support assistant, a Retrieval Augmented Generation (RAG) application built with LangChain to compare two products across three dimensions, they’re effectively posing six questions simultaneously. Similarity search uses a single query vector to encapsulate all the intents. The retriever then generates the best approximation of the average of those intents. The resulting answer comes back concise. The search executes without errors. The relevance scores look reasonable. Yet the retrieved chunks, while topically relevant, only cover a fraction of what the question actually asked.

In this post, we showcase a RAG application on Amazon Bedrock Managed Knowledge Base with LangChain. We run the same multi-part question through standard and agentic retrieval, and read the trace events to see the plan the model produced. We also cover what the two retrieval paths cost and when the cheaper one is the right choice.

Agentic retrieval is available on Amazon Bedrock Managed Knowledge Base. Instead of one search, Amazon Bedrock Managed Knowledge Base plans the retrieval. It breaks the question into sub-queries, runs them, judges whether it has enough evidence, and searches again if it doesn’t. The langchain-aws package exposes both agentic and standard retrieval, so you can use either from a LangChain application.

Amazon Bedrock Managed Knowledge Base, the fully managed RAG capability in Amazon Bedrock, removes the self-managed vector store, embeddings, and re-ranking models from the RAG architecture. You configure a data source, and Amazon Bedrock Managed Knowledge Bases handles chunking, embedding, storage, and retrieval. This walkthrough uses Amazon Simple Storage Service (Amazon S3).

Amazon Bedrock Managed Knowledge Bases provides two APIs. We briefly discuss those differences in this post. The Retrieve API runs one hybrid search and returns scored chunks. The AgenticRetrieveStream API runs a planning loop and streams the steps back to you as trace events. In the langchain-aws package, the first is a standard LangChain retriever you can drop into a chain. The second is a function retrieval directly from a knowledge base.

The following diagram shows the solution architecture. The application queries Amazon Bedrock Knowledge Bases using either the Retrieve API (standard, single-shot) or the AgenticRetrieveStream API (multi-step planning loop). Both paths return document chunks from the knowledge base, which the application then uses to generate a grounded response.

Figure 1: Solution architecture for querying Amazon Bedrock Knowledge Bases with the Retrieve and AgenticRetrieveStream APIs

The following sections walk you through creating a knowledge base, querying it with both retrieval methods, and reading the trace events the agentic planner produces.

Install the packages. The Boto3 version matters: agentic_retrieve_stream did not exist before 1.43.32.

Two identities are involved and separating them is worth doing deliberately. The knowledge base assumes a service role to read your documents and call the embedding model. Your application uses an AWS Security Token Service (AWS STS) caller identity to query. Neither needs the other’s permissions.

Amazon Bedrock creates the service role for you if you let it. To supply your own, give it a trust policy that lets Amazon Bedrock assume it. Scope it with aws:SourceAccount and aws:SourceArn so that another account can’t use it as a confused deputy:

The service role also needs s3:ListBucket on your bucket and s3:GetObject on its contents, both conditioned on aws:ResourceAccount. Scope the knowledge-base/* wildcard character down to specific knowledge base IDs after you have created them.

The AWS STS caller identity needs a different set. bedrock:AgenticRetrieveStream and bedrock:InvokeModelWithResponseStream can’t be scoped to a knowledge base Amazon Resource Name (ARN). bedrock:Retrieve and bedrock:GetDocumentContent can:

bedrock:GetDocumentContent is often overlooked. Agentic retrieval calls it when a FullDocumentExpansion step decides a passage lacks the context to answer. A policy with only bedrock:Retrieve works until the planner reaches for a whole document and then fails partway through a query.

To create and manage the knowledge base itself, the calling role additionally needs bedrock:CreateKnowledgeBase on *, and the GetKnowledgeBase, UpdateKnowledgeBase, DeleteKnowledgeBase, StartIngestionJob, GetIngestionJob, and ListIngestionJobs actions on knowledge-base/*. If you’re using guardrails, add bedrock:GetGuardrail and bedrock:ApplyGuardrail.

Running this walkthrough might incur costs for document storage and ingestion in the knowledge base, retrieval calls, and foundation model (FM) inference.

For more information about pricing, see the Knowledge Bases section of Amazon Bedrock pricing.

Delete the resources when you complete this experiment.

Creating and populating the knowledge base

Create the knowledge base with a managedKnowledgeBaseConfiguration. Setting embeddingModelType to MANAGED uses the service-managed embedding model.

There’s no storageConfiguration in that request. For a self-managed knowledge base you would pass one describing your vector store. Amazon Bedrock Managed Knowledge Base does not take one, which is the clearest signal in the API that Amazon Bedrock owns the storage layer.

Attach the S3 bucket as a data source, then start an ingestion job. Ingestion is asynchronous, so poll until the job reaches a terminal state rather than sleeping for a fixed interval and hoping.

The full data source configuration and error handling are in the sample repository.

Querying with the LangChain retriever

AmazonKnowledgeBasesRetriever wraps the Retrieve API and behaves like any other LangChain retriever. For Amazon Bedrock Managed Knowledge Bases, pass managedSearchConfiguration. This is the part that trips people up: vectorSearchConfiguration is the previous path for knowledge bases where you run your own vector store. It is what most existing examples show.

Each result comes back as a LangChain Document. The relevance score is in metadata["score"], and the source document’s own metadata is under metadata["source_metadata"], renamed so it does not collide. If you want to drop low-confidence results, set min_score_confidence on the retriever instead of filtering afterward.

For a question with one clear intent, this is the right tool. It is one call. The latency is the lowest of the two options, and you keep full control of how the answer gets generated. Most of the queries a production assistant sees are this shape, and reaching for a planning loop to answer them wastes money and time.

Where single-shot retrieval runs out

Now give the same retriever a question with several parts:

Five chunks come back, ranked by hybrid score against one embedding of that whole question.

That question contains six intents: two services across three dimensions. Scoring the retrieved text for evidence of each one gives a concrete measure of what a single embedding recovers.

At five results, one embedding standing for six intents misses two of them. At ten it covers all six, with visible waste: two sub-intents are covered twice and one chunk carries none.

The retriever did its job. The limitation is structural: one vector cannot represent six intents, and there is no step in the process that asks whether the returned evidence is enough to answer the question.

Agentic retrieval is not a LangChain retriever, but a feature of Amazon Bedrock Managed Knowledge Bases. The langchain-aws package exposes it as a standalone function, agentic_retrieve, because the underlying API streams its results and doesn’t fit the synchronous BaseRetriever interface. There’s no flag on AmazonKnowledgeBasesRetriever that switches it on.

With generate_response=True, the service returns a grounded answer and citations alongside the retrieved chunks, so you get an answer without wiring up a separate model call. The function works only against Amazon Bedrock Managed Knowledge Base.

Internally, the service plans, retrieves, evaluates whether the evidence is sufficient, and iterates if it isn’t. The helper hides all of that and hands back the final chunks, which is convenient and means you can’t see the plan.

To watch the model decompose the question, call agentic_retrieve_stream on the bedrock-agent-runtime client directly. This is the one place in this walkthrough where we step around langchain-aws, because the helper discards trace events and doesn’t expose maxAgentIteration or a custom planner model.

Set generateResponse to False when you only want the retrieval behavior. The API generates a grounded answer by default, which costs an extra model call you may not need while you are inspecting the plan.

A trace event’s step tells you where the planner is. SpeculativeRetrieval runs before the first plan to cut latency and doesn’t count against your iteration budget. Planning is where the model reads the question and prior results and emits sub-queries. Retrieval fires once per sub-query. FullDocumentExpansion appears when the model decides a passage lacks the context to answer and pulls the whole document instead. Each carries a status of IN_PROGRESS, SUCCEEDED, or FAILED, plus a human-readable message.

The final chunks arrive separately. The result is its own event type rather than a fifth step, and it holds the deduplicated chunks from every iteration along with the grounded answer when response generation is on. Branch on the event key, as the preceding loop demonstrates, rather than expect a terminal step value.

The sub-query text is the part worth logging. It sits in attributes.actions[].retrieve.inputQuery.text, not in the top-level trace fields, so a handler that reads only step and status shows you that planning happened without showing you what it decided.

The following diagram shows the agentic retrieval planning loop, including the speculative retrieval, planning, sub-query retrieval, evaluation, and optional re-planning steps.

Figure 2: Steps in the agentic retrieval planning loop

Two details are worth knowing before you build on this. Deduplication applies only to the result event, so a chunk retrieved by three sub-queries appears once at the end but three times across the traces.

The second is about scores. A Retrieve response gives each chunk a typed score field holding its relevance to the query. Agentic retrieval results carry content, metadata, and sourceRetriever, with no equivalent typed field. Code that reads result["score"] after switching APIs gets nothing. If you rank or filter relevance, plan for that difference.

In production, use Amazon Bedrock Guardrails to enforce content policies and grounding checks on generated responses. Both retrieval paths support guardrails. Agentic retrieval supports guardrails through policyConfiguration.bedrockGuardrailConfiguration rather than the guardrail_config argument the LangChain retriever takes, and supports BLOCK mode only. If you rely on MASK mode, that is a reason to stay on the Retrieve API.

maxAgentIteration accepts two through ten and defaults to five. Leave it at the default. At two or three the planner runs one cycle, emits no sub-queries, and returns what the speculative retrieval step already found. This is single-shot behavior at the agentic price. Decomposition begins at four. The planner often stops early when it judges the evidence sufficient, so the ceiling is a bound rather than a target.

Comparing the two retrieval paths

For context on how this behaves at scale, AWS evaluated agentic retrieval on MuSiQue, a public multi-hop benchmark. The evaluation showed improved recall over single-shot retrieval, with the largest gains on the hardest questions. Single-hop questions saw gains under five points. That last figure matches the shape of the trade-off: decomposition helps when there is something to decompose.

For the standard retriever, the usual LangChain Expression Language (LCEL) composition works directly:

format_docs matters more than it looks. Passing Document objects straight into a prompt renders their repr, and the model gets metadata noise mixed into the context.

To put agentic retrieval in the same position, wrap it in a RunnableLambda, since it is a function rather than a retriever:

Note that generate_response is off here. The service can generate the answer itself, but inside a chain you usually want your own prompt and model, so you take the chunks and generate downstream. Use the service generation when you want one call and less code, and the wrapped version when the prompt is yours to control.

Choosing between standard and agentic retrieval

Use Retrieve for short, well-scoped questions. It is cheaper, faster, works against self-managed knowledge bases, and returns scores in the results. Most production traffic looks like this.

Use AgenticRetrieveStream when questions are multi-part, comparative, or exploratory, or when the evidence spans more than one knowledge base. It registers up to five knowledge bases in one request and routes sub-queries using a natural-language description you attach to each. The other API cannot do this at all. It costs more per call, makes several model invocations, and has the higher latency of the two.

Routing on query shape rather than picking one for everything is the pattern we recommend. A classifier or a heuristic on the question can send most traffic down the cheap path and reserve the planner for questions that need it.

Delete the knowledge base, its data source, the S3 objects and bucket, and the IAM role that you created. A knowledge base with documents in it continues to incur storage charges.

The repository includes a cleanup script that also empties the bucket and removes the role.

We showed how to build a RAG application on Amazon Bedrock Knowledge Bases with LangChain, and how agentic retrieval handles multi-part questions that single-shot retrieval answers poorly. We also showed the friction in the current integration. Agentic retrieval is a function rather than a LangChain retriever, so it needs a RunnableLambda to sit in a chain. The trace events that show the query plan require a direct boto3 call.

Agentic retrieval trades higher per-call cost for improved recall on multi-hop questions, using a built-in model for query planning. The next useful step is measuring your own query mix before you route everything through a planner.

To get started, see the Amazon Bedrock Knowledge Bases documentation and the accompanying sample code. For help applying this to your own workload, contact your AWS account team.