overfeed.news

Reduce RAG costs on Amazon Bedrock with query-aware compression

1mo

Age

Published
Collected
Image: AWS Machine Learning Blog

Input tokens are often a meaningful part of the cost of running Retrieval Augmented Generation (RAG) at scale. This post describes a query-aware context compression pattern on Amazon Bedrock: after retrieval, a smaller model filters retrieved chunks against the query before the primary model answers, reducing input tokens and cost while preserving answer quality.

Excerpt from the source

Input tokens sent to the foundation model (FM) on every call are often a meaningful part of the cost of running Retrieval Augmented Generation (RAG) at scale. Query-aware compression offers one way to reduce how many of them reach the model. Amazon Bedrock provides the foundation models and features to build RAG applications. RAG retrieval usually tunes for high recall, returning a broad set of potentially relevant chunks so the primary model has thorough source material to work with. This…

Read the full article at aws.amazon.com

overfeed.news indexes and links. We publish a short excerpt — the full article stays at AWS Machine Learning Blog.

Log in to follow this source
Reduce RAG costs on Amazon Bedrock with query-aware compression…