How Adaptive RAG Reduces Token Costs While Maintaining Accuracy in Pathway LLM Apps

Adaptive RAG cuts token costs by starting with just 2 documents and geometrically expanding retrieval only when the LLM signals uncertainty, avoiding the waste of processing large fixed contexts for every query.

Adaptive Retrieval-Augmented Generation (Adaptive RAG) transforms the traditional RAG pipeline from a static, one-size-fits-all retrieval into an iterative, confidence-driven process. In the pathwaycom/llm-app repository, this pattern is implemented in the templates/adaptive_rag template, where the AdaptiveRAGQuestionAnswerer class orchestrates dynamic context expansion to minimize token consumption without sacrificing answer quality.

How Adaptive RAG Differs from Standard RAG

Standard RAG pipelines retrieve a fixed-size chunk set—often 10 to 20 documents—regardless of query complexity. This forces the LLM to process massive contexts even for simple questions that require only a single source. Adaptive RAG reverses this logic by treating retrieval as a conditional loop.

Standard RAG Adaptive RAG
Retrieves a fixed-size chunk set for every query, regardless of difficulty. Starts with a tiny set (default 2 documents) and expands only when the LLM signals insufficient information.
Token usage is proportional to the static chunk size, leading to high average costs. Each iteration adds a geometrically larger slice (factor 2 by default) only when needed, dramatically lowering average token count.
Accuracy depends on choosing a "one-size-fits-all" document count; too few harms quality, too many wastes tokens. The algorithm stops as soon as the answer meets a confidence threshold, guaranteeing accuracy while spending minimal tokens.

The Geometric Retrieval Strategy Behind Adaptive RAG

The core mechanism resides in the answer_with_geometric_rag_strategy_from_index function within Pathway's LLM xpack. This implementation uses an exponential backoff strategy in reverse: instead of starting large and shrinking, it starts minimal and expands.

Configuration Parameters in app.yaml

The templates/adaptive_rag/app.yaml file exposes the tunable parameters that control the geometric expansion:

question_answerer: !pw.xpacks.llm.question_answering.AdaptiveRAGQuestionAnswerer
  llm: $llm
  indexer: $document_store
  n_starting_documents: 2          # Start with 2 docs

  factor: 2                        # Double docs each round

  max_iterations: 4                # Hard stop after 4 rounds

These settings create a retrieval ceiling of 32 documents (2 → 4 → 8 → 16 → 32) across four iterations, though most queries terminate after the first or second round when the LLM reports sufficient confidence.

LLM-Based Stopping Criterion

After each retrieval round, the AdaptiveRAGQuestionAnswerer prompts the LLM to generate both an answer and a self-assessment. If the model indicates uncertainty or requests additional context—detected via a structured confidence score or explicit "need more context" flag—the loop continues. Otherwise, the current answer is returned immediately, truncating further retrieval and saving tokens.

Implementing Adaptive RAG in Your Application

The templates/adaptive_rag/app.py file demonstrates how to bootstrap the adaptive pipeline. It imports SummaryQuestionAnswerer (which wraps the adaptive logic) and launches a REST server.

Setting Up the Configuration

First, ensure your app.yaml defines the adaptive question answerer with your preferred starting parameters:


# templates/adaptive_rag/app.py

from pathway.xpacks.llm.question_answering import SummaryQuestionAnswerer
from pathway.xpacks.llm.servers import QASummaryRestServer

# The App class loads app.yaml and instantiates AdaptiveRAGQuestionAnswerer

# based on the configuration shown in the previous section

Running the Service

Install Pathway with LLM extras and launch the application:

pip install pathway[all]
python templates/adaptive_rag/app.py

The service exposes the /v2/answer endpoint. When queried, it executes the geometric retrieval strategy:

curl -X POST http://0.0.0.0:8000/v2/answer \
  -H 'Content-Type: application/json' \
  -d '{"prompt": "When was the company founded?"}'

Behind the scenes, the answer_with_geometric_rag_strategy_from_index function retrieves 2 documents, checks LLM confidence, and only fetches 4, 8, or 16 additional documents if the initial context proves insufficient.

Summary

  • Adaptive RAG replaces static retrieval with an iterative, confidence-driven loop that starts with just 2 documents and geometrically expands only when necessary.
  • The geometric expansion factor (default 2) and maximum iteration limit (default 4) are configured in templates/adaptive_rag/app.yaml under the AdaptiveRAGQuestionAnswerer class.
  • Token cost reduction occurs because most queries resolve in the first or second iteration, avoiding the processing of large fixed contexts required by standard RAG.
  • The core logic resides in the answer_with_geometric_rag_strategy_from_index function within Pathway's LLM xpack, referenced in templates/adaptive_rag/README.md.

Frequently Asked Questions

What is the default starting document count in Adaptive RAG?

The default configuration in templates/adaptive_rag/app.yaml sets n_starting_documents: 2. This means the system initially retrieves only 2 documents from the index, minimizing token usage for simple queries that require minimal context.

How does Adaptive RAG determine when to stop retrieving documents?

After each retrieval iteration, the LLM evaluates its own answer confidence. If the model indicates sufficient certainty or provides a complete answer, the loop terminates. If the LLM signals uncertainty or requests additional context, the system retrieves geometrically more documents (doubling by default) until either confidence is achieved or max_iterations (default 4) is reached.

Can I adjust the geometric expansion factor in Adaptive RAG?

Yes. The expansion factor is controlled by the factor parameter in app.yaml. The default value is 2, meaning each iteration doubles the document count (2 → 4 → 8 → 16). You can increase this to retrieve more context per iteration or decrease it for finer-grained control, though the default geometric progression optimally balances token efficiency and retrieval latency.

Where is the core Adaptive RAG logic implemented in the Pathway codebase?

While the template configuration resides in templates/adaptive_rag/, the underlying implementation lives in Pathway's LLM xpack within the pathway.xpacks.llm.question_answering module. Specifically, the function answer_with_geometric_rag_strategy_from_index handles the iterative retrieval and confidence checking logic, as referenced in the template's README.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →