Reduce RAG costs on Amazon Bedrock with query
Input tokens sent to the foundation model (FM) on every call are often a meaningful part of the cost of running Retrieval Augmented Generation (RAG) at scale. Query-aware compression offers one way to reduce how many of them reach the model. Amazon Bedrock provides the foundation models and features to build RAG applications. RAG retrieval usually tunes for high recall, returning a broad set of potentially relevant chunks so the primary model has thorough source material to work with. This design helps builders feel confident that the right information is available at inference time. As workloads scale, builders often look for ways to optimize the cost-performance tradeoff by reducing the number of input tokens the primary model processes while maintaining answer quality. The open, composable architecture of Amazon Bedrock supports custom post-retrieval processing steps that refine what reaches the primary model.
In this post, we describe a post-retrieval customization pattern that achieves significant input-token reduction, and therefore cost savings, while preserving answer quality. It’s compatible with RAG retrievers on Amazon Bedrock, including Amazon Bedrock Knowledge Bases. As a secondary benefit, removing irrelevant context reduces the surface area for hallucination. After retrieval but before the final answer call, a smaller, lower-cost model on Amazon Bedrock filters retrieved chunks against the user’s query. The primary model then receives the filtered context and generates the answer.
We cover the pattern’s architecture at a high level, show the core Amazon Bedrock implementation in a AWS Lambda function, walk through the cost model and the latency tradeoff, and describe how we evaluated answer quality. We also look at how this pattern can layer on top of existing Amazon Bedrock capabilities like prompt caching, Amazon Bedrock Intelligent Prompt Routing, and the Rerank API for compounding cost savings.
To implement the solution, complete the following prerequisite steps:
How query-aware compression reduces RAG costs on Amazon Bedrock
A RAG flow using traditional RAG infrastructure or frameworks looks like:
Retrieved context scales with top-k and chunk size: retrieving 5–20 chunks at typical chunk sizes puts many technical-documentation and legal RAG workloads in the range of several thousand input tokens per query. Reducing the per-query token count can yield meaningful cost savings.
A smaller model reads the retrieved chunks alongside the query and outputs only the verbatim spans relevant to the question. We use Claude Haiku in this post, but the pattern works with other small/primary model pairs within a model family on Amazon Bedrock. Both the compression call and the primary model’s answer call run inside a single AWS Lambda function. Upstream, a retriever embeds the query and returns the top-k chunks. An Amazon Bedrock knowledge base, the fully managed RAG capability backed by Amazon OpenSearch Serverless, is one such retriever. The Lambda function receives those chunks as input and returns the final answer. The compression call is the only step added to a standard RAG flow.
Because the smaller model costs less per token than the primary model, trimming the context before the expensive answer call is where the savings come from. How large those savings are comes down to two things.
The following diagram shows the solution architecture.
Figure 1: Query-aware context compression architecture on Amazon Bedrock
The flow proceeds through the following steps:
The economics depend on two factors: the price ratio between the small and primary models on Amazon Bedrock, and the compression ratio the smaller model achieves.
For a single RAG query with R retrieved input tokens, a compression ratio of c (where c > 1), a final answer output of A tokens, and per-token prices P_small_in / P_small_out (smaller model input and output price) and P_large_in / P_large_out (primary model input and output price):
The compression call adds the input and output cost of the smaller model. Savings come from sending R/c instead of R tokens to the primary model. The economics favor compression when:
Implementation on Amazon Bedrock
The pattern fits between retrieval and the final answer call. We implement it as a single AWS Lambda function that orchestrates the two Amazon Bedrock model invocations using the Converse API. The function receives the user query and the retrieved chunks as its input event.
The compression prompt is the most important part of the implementation. It must instruct the smaller model to extract spans rather than summarize, forbid paraphrasing and rewriting, and preserve enough surrounding context for citations to remain accurate.
The function takes the user query and the retrieved chunks, reads the two model IDs from environment variables, and initializes an Amazon Bedrock Runtime client configured with adaptive retries. It then makes two calls through the Bedrock Converse API: the first to the smaller model to compress the chunks, and the second to the primary model to generate the answer from the compressed evidence. The compression call runs at temperature 0.0, which keeps the extraction deterministic so the smaller model copies spans as they appear in the source:
Step 1: Compress. The smaller model filters the retrieved chunks.
Step 2: Answer. The primary model reasons over the filtered evidence.
Note: The prompts are an example and should be adapted to your documents and question types.
Before recommending this pattern, we evaluated it empirically. The benchmark covered:
These figures describe one corpus, one domain, and one query distribution. Results on your own documents and queries will differ.
The following table summarizes the headline results across the queries, comparing the baseline, compression, and rerank + compression pipelines.
The following figures come from the benchmark. Results on your own corpus, queries, and model choices will differ. The following figure shows the average of query cost saving (left axis, percentage versus baseline) and the reduction in context sent to the primary model (right axis, times fewer tokens). Compression achieved a 33 percent cost saving or 8.6× fewer tokens. Rerank + compression reached 36 percent cost saving and 10.1× fewer tokens.
Figure 2: Cost savings and reduction in context tokens sent to the primary model
The following figure shows the LLM-judge scores (1–5) across the four answer-quality dimensions for each pipeline. Correctness stays within 0.07 of baseline across conditions. Completeness and citation accuracy are slightly lower under compression, while conciseness is slightly higher.
Figure 3: LLM-judge scores across the 4 answer-quality dimensions by pipeline
The following figure shows the hallucination rate for each pipeline, measured as the share of answers containing at least one claim not supported by the reference. The baseline is 51 percent, compression 44 percent, and rerank + compression 38 percent.
Figure 4: Hallucination rate by pipeline
The following figure shows the cost saving versus baseline for the typical-query set and the hard-query set. Compression moves from 37 percent to 26 percent, and rerank + compression from 40 percent to 30 percent between the two sets.
Figure 5: Cost savings for the typical-query and hard-query sets
Considerations for production use
Three things must be weighed before this pattern is adopted: the latency of the added compression call, the impact on answer quality, and whether the workload is a fit. Each is covered in the following sections.
Adding a smaller-model call introduces one extra step in the path. Claude Haiku is optimized for speed, and because the primary model then processes a smaller, focused context, part of that added time is recovered on the answer call. The total end-to-end latency is the compression call plus the answer call on the focused context. Net impact depends on how compute-bound the primary model is on the original context size.
For latency-critical surfaces (sub-second chat), measure with your own context sizes before deploying.
Compression involves a few factors worth deliberate engineering.
A few considerations can help guide the selection of the smaller model:
For production deployments, Amazon Bedrock Guardrails can provide additional content filtering and grounding validation as complementary controls alongside the compression pattern.
The pattern provides value when three conditions line up. Retrieved context is large, the primary model is the expensive part of the call, and the question is narrow relative to the breadth of what was retrieved. The following scenarios share these characteristics.
This pattern is likely worth a prototype if most of these are true for your workload:
If your workload is sub-second conversational chat with small retrieved context, this pattern may not be the best fit. Consider Amazon Bedrock features like prompt caching and Intelligent Prompt Routing, which provide more value with less added latency.
The pattern fits into RAG pipelines built on Amazon Bedrock: A single Lambda function between retrieval and the final model call compresses context and fine-tunes cost and quality for your specific workload.
The compression prompt handles both simple lookups and complex multi-source questions without requiring separate routing logic. Combined with existing features like prompt caching, Intelligent Prompt Routing, and the Rerank API, this approach delivers compounding optimizations across the full RAG pipeline.
You can use this pattern behind a feature flag, measure it on your real query distribution, and tune the compression prompts to match your domain.
To get started, refer to the Amazon Bedrock documentation and the Amazon Bedrock pricing page for current model costs.
Related Stories
AI News
Desmond Robinson | Artificial intelligence vs supernatural intelligence
45 minutes ago
AI News
Combating online youth deepfakes
1 hour ago
AI News
Already existing laws governing AI simply being ignored, former regulators, analysts say
1 hour ago
AI News
Utah pilot lets artificial intelligence prescribe acne treatments
1 hour ago
AI News
Artificial Goes After Our Silicon Valley Overlords. Just Not Very Well.
1 hour ago
Before you buy more AI, diagnose the gap you actually have
2 hours ago
AI News
Companies must worry about rogue AI inside and outside the tent
2 hours ago
AI News
Sharing AI progress in mathematics
3 hours ago