# How Adaptive RAG Reduces Token Costs While Maintaining Accuracy in Pathway LLM Apps

> Discover how Adaptive RAG slashes token costs for LLM apps. Get accurate results by retrieving only necessary documents, unlike standard RAG. Read more.

- Repository: [Pathway/llm-app](https://github.com/pathwaycom/llm-app)
- Tags: deep-dive
- Published: 2026-03-07

---

**Adaptive RAG cuts token costs by starting with just 2 documents and geometrically expanding retrieval only when the LLM signals uncertainty, avoiding the waste of processing large fixed contexts for every query.**

Adaptive Retrieval-Augmented Generation (Adaptive RAG) transforms the traditional RAG pipeline from a static, one-size-fits-all retrieval into an iterative, confidence-driven process. In the `pathwaycom/llm-app` repository, this pattern is implemented in the `templates/adaptive_rag` template, where the `AdaptiveRAGQuestionAnswerer` class orchestrates dynamic context expansion to minimize token consumption without sacrificing answer quality.

## How Adaptive RAG Differs from Standard RAG

Standard RAG pipelines retrieve a fixed-size chunk set—often 10 to 20 documents—regardless of query complexity. This forces the LLM to process massive contexts even for simple questions that require only a single source. Adaptive RAG reverses this logic by treating retrieval as a conditional loop.

| Standard RAG | Adaptive RAG |
|--------------|--------------|
| Retrieves a fixed-size chunk set for every query, regardless of difficulty. | Starts with a tiny set (default 2 documents) and expands only when the LLM signals insufficient information. |
| Token usage is proportional to the static chunk size, leading to high average costs. | Each iteration adds a geometrically larger slice (factor 2 by default) only when needed, dramatically lowering average token count. |
| Accuracy depends on choosing a "one-size-fits-all" document count; too few harms quality, too many wastes tokens. | The algorithm stops as soon as the answer meets a confidence threshold, guaranteeing accuracy while spending minimal tokens. |

## The Geometric Retrieval Strategy Behind Adaptive RAG

The core mechanism resides in the `answer_with_geometric_rag_strategy_from_index` function within Pathway's LLM xpack. This implementation uses an exponential backoff strategy in reverse: instead of starting large and shrinking, it starts minimal and expands.

### Configuration Parameters in app.yaml

The [`templates/adaptive_rag/app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/templates/adaptive_rag/app.yaml) file exposes the tunable parameters that control the geometric expansion:

```yaml
question_answerer: !pw.xpacks.llm.question_answering.AdaptiveRAGQuestionAnswerer
  llm: $llm
  indexer: $document_store
  n_starting_documents: 2          # Start with 2 docs

  factor: 2                        # Double docs each round

  max_iterations: 4                # Hard stop after 4 rounds

```

These settings create a retrieval ceiling of 32 documents (2 → 4 → 8 → 16 → 32) across four iterations, though most queries terminate after the first or second round when the LLM reports sufficient confidence.

### LLM-Based Stopping Criterion

After each retrieval round, the `AdaptiveRAGQuestionAnswerer` prompts the LLM to generate both an answer and a self-assessment. If the model indicates uncertainty or requests additional context—detected via a structured confidence score or explicit "need more context" flag—the loop continues. Otherwise, the current answer is returned immediately, truncating further retrieval and saving tokens.

## Implementing Adaptive RAG in Your Application

The [`templates/adaptive_rag/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/adaptive_rag/app.py) file demonstrates how to bootstrap the adaptive pipeline. It imports `SummaryQuestionAnswerer` (which wraps the adaptive logic) and launches a REST server.

### Setting Up the Configuration

First, ensure your [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml) defines the adaptive question answerer with your preferred starting parameters:

```python

# templates/adaptive_rag/app.py

from pathway.xpacks.llm.question_answering import SummaryQuestionAnswerer
from pathway.xpacks.llm.servers import QASummaryRestServer

# The App class loads app.yaml and instantiates AdaptiveRAGQuestionAnswerer

# based on the configuration shown in the previous section

```

### Running the Service

Install Pathway with LLM extras and launch the application:

```bash
pip install pathway[all]
python templates/adaptive_rag/app.py

```

The service exposes the `/v2/answer` endpoint. When queried, it executes the geometric retrieval strategy:

```bash
curl -X POST http://0.0.0.0:8000/v2/answer \
  -H 'Content-Type: application/json' \
  -d '{"prompt": "When was the company founded?"}'

```

Behind the scenes, the `answer_with_geometric_rag_strategy_from_index` function retrieves 2 documents, checks LLM confidence, and only fetches 4, 8, or 16 additional documents if the initial context proves insufficient.

## Summary

- **Adaptive RAG** replaces static retrieval with an iterative, confidence-driven loop that starts with just **2 documents** and geometrically expands only when necessary.
- The **geometric expansion factor** (default 2) and **maximum iteration limit** (default 4) are configured in [`templates/adaptive_rag/app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/templates/adaptive_rag/app.yaml) under the `AdaptiveRAGQuestionAnswerer` class.
- **Token cost reduction** occurs because most queries resolve in the first or second iteration, avoiding the processing of large fixed contexts required by standard RAG.
- The core logic resides in the `answer_with_geometric_rag_strategy_from_index` function within Pathway's LLM xpack, referenced in [`templates/adaptive_rag/README.md`](https://github.com/pathwaycom/llm-app/blob/main/templates/adaptive_rag/README.md).

## Frequently Asked Questions

### What is the default starting document count in Adaptive RAG?

The default configuration in [`templates/adaptive_rag/app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/templates/adaptive_rag/app.yaml) sets `n_starting_documents: 2`. This means the system initially retrieves only 2 documents from the index, minimizing token usage for simple queries that require minimal context.

### How does Adaptive RAG determine when to stop retrieving documents?

After each retrieval iteration, the LLM evaluates its own answer confidence. If the model indicates sufficient certainty or provides a complete answer, the loop terminates. If the LLM signals uncertainty or requests additional context, the system retrieves geometrically more documents (doubling by default) until either confidence is achieved or `max_iterations` (default 4) is reached.

### Can I adjust the geometric expansion factor in Adaptive RAG?

Yes. The expansion factor is controlled by the `factor` parameter in [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml). The default value is `2`, meaning each iteration doubles the document count (2 → 4 → 8 → 16). You can increase this to retrieve more context per iteration or decrease it for finer-grained control, though the default geometric progression optimally balances token efficiency and retrieval latency.

### Where is the core Adaptive RAG logic implemented in the Pathway codebase?

While the template configuration resides in `templates/adaptive_rag/`, the underlying implementation lives in Pathway's LLM xpack within the `pathway.xpacks.llm.question_answering` module. Specifically, the function `answer_with_geometric_rag_strategy_from_index` handles the iterative retrieval and confidence checking logic, as referenced in the template's README.md.