# Key Differences Between question_answering_rag, document_indexing, and adaptive_rag Templates in Pathway LLM-App

> Understand the key differences between Pathway LLM-App's question_answering_rag, document_indexing, and adaptive_rag templates. Choose the best RAG solution for your needs.

- Repository: [Pathway/llm-app](https://github.com/pathwaycom/llm-app)
- Tags: deep-dive
- Published: 2026-03-07

---

**The question_answering_rag template provides end-to-end Retrieval-Augmented Generation (RAG) with LLM-based answers, document_indexing serves as a pure vector storage and retrieval service without LLM generation, and adaptive_rag extends standard RAG with a geometric expansion strategy that dynamically grows context size to optimize token usage.**

All three templates reside in the `templates/` directory of the **pathwaycom/llm-app** repository and share a common Pathway-based backbone for document processing. While they reuse identical persistence and parsing infrastructure, each targets a distinct operational mode—ranging from full conversational AI to bare-metal vector search to token-efficient adaptive retrieval.

## Architectural Overview

Each template follows a configuration-driven architecture centered on an `App` class defined in [`app.py`](https://github.com/pathwaycom/llm-app/blob/main/app.py). The templates share common persistence logic (`pw.PersistenceMode.UDF_CACHING`) and filesystem backend configuration (lines 31–61 across all three files), but diverge in their core Pathway components and REST server implementations.

**Shared infrastructure includes:**
- YAML-based configuration via [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml)
- Pathway document parsing and embedding pipelines
- Optional filesystem caching (`./Cache`)
- Docker Compose orchestration files

## Core Differences by Template

### Question-Answering RAG (question_answering_rag)

**Purpose:** End-to-end RAG pipeline that indexes documents **and** answers natural language questions using a hosted LLM.

**Key Components:**
- **`SummaryQuestionAnswerer`** – Wraps a `DocumentStore` with LLM generation capabilities
- **`QASummaryRestServer`** – Exposes conversational endpoints including summarization
- **`App` class** – Instantiates `question_answerer: InstanceOf[SummaryQuestionAnswerer]`

**Implementation:** In [`templates/question_answering_rag/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/question_answering_rag/app.py), the application constructs a `QASummaryRestServer` that routes requests to the `SummaryQuestionAnswerer`, providing complete document ingestion and chat functionality in a single service.

### Document Indexing (document_indexing)

**Purpose:** Standalone document store service that **only** indexes documents and provides retrieval endpoints—no LLM generation involved.

**Key Components:**
- **`DocumentStore`** – Pure vector index with similarity search and statistics
- **`DocumentStoreServer`** – Lightweight server exposing only retrieval-related routes
- **`App` class** – Instantiates `document_store: InstanceOf[DocumentStore]`

**Implementation:** The [`templates/document_indexing/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/document_indexing/app.py) file creates a `DocumentStoreServer` rather than a QA server. This template is ideal when you need to embed documents for downstream consumers or want to decouple retrieval from generation.

### Adaptive RAG (adaptive_rag)

**Purpose:** Token-efficient RAG that uses **adaptive retrieval** to minimize LLM API costs while maintaining answer quality.

**Key Components:**
- **`SummaryQuestionAnswerer`** – Same class as standard QA-RAG, but configured with adaptive parameters
- **`answer_with_geometric_rag_strategy_from_index`** – Dynamically expands retrieved chunks geometrically until confidence thresholds are met
- **Geometric expansion parameters** – `start_documents` and `geometric_factor` tunable in [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml)

**Implementation:** The [`templates/adaptive_rag/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/adaptive_rag/app.py) file overrides the standard `answer` method to call `answer_with_geometric_rag_strategy_from_index` (see lines 31–33 in the README). This strategy starts with a small retrieval set and fetches additional chunks only when the LLM indicates insufficient context, reducing token consumption compared to the fixed-context approach in standard QA-RAG.

## Configuration and Implementation Details

All three pipelines are driven by [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml), but their configuration schemas differ based on server type.

**Question-Answering RAG and Adaptive RAG require:**
- LLM configuration (OpenAI, Azure, or local)
- Embedder settings
- Index parameters
- Document sources

**Document Indexing requires only:**
- DocumentStore configuration (sources, parser, splitter, retriever)

**App Class Structure:**
- **QA-RAG & Adaptive:** `class App(BaseModel): question_answerer: InstanceOf[SummaryQuestionAnswerer]`
- **Document Indexing:** `class App(BaseModel): document_store: InstanceOf[DocumentStore]`

Swapping between QA-RAG and Adaptive RAG requires only changing the template directory—both use identical [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml) structures, with Adaptive RAG adding optional adaptive parameters like `start_documents` and `geometric_factor`.

## API Endpoints Comparison

| Template | Available Endpoints |
|----------|---------------------|
| **question_answering_rag** | `/v1/retrieve` (similarity search)<br>`/v2/answer` (RAG Q&A)<br>`/v2/summarize` (document summarization)<br>`/v1/statistics`<br>`/v2/list_documents` |
| **document_indexing** | `/v1/retrieve`<br>`/v1/statistics`<br>`/v2/list_documents` |
| **adaptive_rag** | Same as QA-RAG, but `/v2/answer` executes geometric expansion strategy under the hood |

The **Document Indexing** template intentionally omits `/v2/answer` and `/v2/summarize` because it lacks an LLM integration. The **Adaptive RAG** template exposes the same interface as standard QA-RAG, making it a drop-in replacement for token-constrained environments.

## Practical Usage Examples

### Running Question-Answering RAG

Navigate to the template directory and launch via Docker Compose:

```bash
cd templates/question_answering_rag
docker compose up --build

```

Query the RAG endpoint:

```bash
curl -X POST http://localhost:8000/v2/answer \
  -H 'Content-Type: application/json' \
  -d '{"prompt":"What is the start date of the contract?"}'

```

### Running Document Indexing

Start the vector storage service:

```bash
cd templates/document_indexing
docker compose up --build

```

Retrieve similar chunks without LLM generation:

```bash
curl -X POST http://localhost:8000/v1/retrieve \
  -H 'Content-Type: application/json' \
  -d '{"query":"latest earnings report", "k":5}'

```

### Running Adaptive RAG

Launch the token-efficient variant:

```bash
cd templates/adaptive_rag
docker compose up --build

```

Submit a question with adaptive retrieval:

```bash
curl -X POST http://localhost:8000/v2/answer \
  -H 'Content-Type: application/json' \
  -d '{"prompt":"Explain the main risk factors in the 10-K filing."}'

```

*Behind the scenes, `answer_with_geometric_rag_strategy_from_index` retrieves an initial small chunk set, evaluates LLM confidence, and expands the context window geometrically only when the answer quality is insufficient.*

## Summary

- **question_answering_rag** provides a complete RAG stack with `SummaryQuestionAnswerer` and `QASummaryRestServer` for document indexing plus conversational AI.
- **document_indexing** offers a lightweight `DocumentStoreServer` for pure vector retrieval scenarios without LLM dependencies.
- **adaptive_rag** uses identical infrastructure to QA-RAG but implements `answer_with_geometric_rag_strategy_from_index` to minimize token usage through dynamic context expansion.
- All three templates share persistence configuration and Pathway document processing pipelines, differing only in their REST server implementation and retrieval strategy.

## Frequently Asked Questions

### When should I use document_indexing instead of question_answering_rag?

Use **document_indexing** when you need only vector search and retrieval capabilities without LLM generation—such as feeding embeddings to a separate downstream service or building a custom RAG pipeline where you control the LLM invocation yourself. It runs lighter and avoids LLM API costs entirely.

### How does adaptive_rag reduce token consumption compared to standard RAG?

The **adaptive_rag** template calls `answer_with_geometric_rag_strategy_from_index` (defined in `pathway.xpacks.llm.question_answering`), which starts with a small number of retrieved documents (configured via `start_documents`) and expands the result set geometrically only if the LLM signals low confidence. This avoids the "over-fetching" problem where standard RAG retrieves a fixed large chunk set for every query regardless of actual need.

### Can I switch between question_answering_rag and adaptive_rag without code changes?

Yes. Both templates use the same `App` class structure and [`app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/app.yaml) schema. You can copy your configuration from [`templates/question_answering_rag/app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/templates/question_answering_rag/app.yaml) to [`templates/adaptive_rag/app.yaml`](https://github.com/pathwaycom/llm-app/blob/main/templates/adaptive_rag/app.yaml), adjust the adaptive-specific parameters (`geometric_factor`, `start_documents`), and run the Adaptive RAG template without modifying Python code.

### What files define the core logic for each template?

The primary implementation files are:
- [`templates/question_answering_rag/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/question_answering_rag/app.py) – Instantiates `QASummaryRestServer` with `SummaryQuestionAnswerer`
- [`templates/document_indexing/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/document_indexing/app.py) – Instantiates `DocumentStoreServer` with `DocumentStore`
- [`templates/adaptive_rag/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/adaptive_rag/app.py) – Identical server setup to QA-RAG but configures the adaptive retrieval strategy

These files are accessible directly in the [pathwaycom/llm-app](https://github.com/pathwaycom/llm-app) repository under their respective template directories.