Key Differences Between question_answering_rag, document_indexing, and adaptive_rag Templates in Pathway LLM-App

The question_answering_rag template provides end-to-end Retrieval-Augmented Generation (RAG) with LLM-based answers, document_indexing serves as a pure vector storage and retrieval service without LLM generation, and adaptive_rag extends standard RAG with a geometric expansion strategy that dynamically grows context size to optimize token usage.

All three templates reside in the templates/ directory of the pathwaycom/llm-app repository and share a common Pathway-based backbone for document processing. While they reuse identical persistence and parsing infrastructure, each targets a distinct operational mode—ranging from full conversational AI to bare-metal vector search to token-efficient adaptive retrieval.

Architectural Overview

Each template follows a configuration-driven architecture centered on an App class defined in app.py. The templates share common persistence logic (pw.PersistenceMode.UDF_CACHING) and filesystem backend configuration (lines 31–61 across all three files), but diverge in their core Pathway components and REST server implementations.

Shared infrastructure includes:

  • YAML-based configuration via app.yaml
  • Pathway document parsing and embedding pipelines
  • Optional filesystem caching (./Cache)
  • Docker Compose orchestration files

Core Differences by Template

Question-Answering RAG (question_answering_rag)

Purpose: End-to-end RAG pipeline that indexes documents and answers natural language questions using a hosted LLM.

Key Components:

  • SummaryQuestionAnswerer – Wraps a DocumentStore with LLM generation capabilities
  • QASummaryRestServer – Exposes conversational endpoints including summarization
  • App class – Instantiates question_answerer: InstanceOf[SummaryQuestionAnswerer]

Implementation: In templates/question_answering_rag/app.py, the application constructs a QASummaryRestServer that routes requests to the SummaryQuestionAnswerer, providing complete document ingestion and chat functionality in a single service.

Document Indexing (document_indexing)

Purpose: Standalone document store service that only indexes documents and provides retrieval endpoints—no LLM generation involved.

Key Components:

  • DocumentStore – Pure vector index with similarity search and statistics
  • DocumentStoreServer – Lightweight server exposing only retrieval-related routes
  • App class – Instantiates document_store: InstanceOf[DocumentStore]

Implementation: The templates/document_indexing/app.py file creates a DocumentStoreServer rather than a QA server. This template is ideal when you need to embed documents for downstream consumers or want to decouple retrieval from generation.

Adaptive RAG (adaptive_rag)

Purpose: Token-efficient RAG that uses adaptive retrieval to minimize LLM API costs while maintaining answer quality.

Key Components:

  • SummaryQuestionAnswerer – Same class as standard QA-RAG, but configured with adaptive parameters
  • answer_with_geometric_rag_strategy_from_index – Dynamically expands retrieved chunks geometrically until confidence thresholds are met
  • Geometric expansion parameters – start_documents and geometric_factor tunable in app.yaml

Implementation: The templates/adaptive_rag/app.py file overrides the standard answer method to call answer_with_geometric_rag_strategy_from_index (see lines 31–33 in the README). This strategy starts with a small retrieval set and fetches additional chunks only when the LLM indicates insufficient context, reducing token consumption compared to the fixed-context approach in standard QA-RAG.

Configuration and Implementation Details

All three pipelines are driven by app.yaml, but their configuration schemas differ based on server type.

Question-Answering RAG and Adaptive RAG require:

  • LLM configuration (OpenAI, Azure, or local)
  • Embedder settings
  • Index parameters
  • Document sources

Document Indexing requires only:

  • DocumentStore configuration (sources, parser, splitter, retriever)

App Class Structure:

  • QA-RAG & Adaptive: class App(BaseModel): question_answerer: InstanceOf[SummaryQuestionAnswerer]
  • Document Indexing: class App(BaseModel): document_store: InstanceOf[DocumentStore]

Swapping between QA-RAG and Adaptive RAG requires only changing the template directory—both use identical app.yaml structures, with Adaptive RAG adding optional adaptive parameters like start_documents and geometric_factor.

API Endpoints Comparison

Template Available Endpoints
question_answering_rag /v1/retrieve (similarity search)/v2/answer (RAG Q&A)/v2/summarize (document summarization)/v1/statistics/v2/list_documents
document_indexing /v1/retrieve/v1/statistics/v2/list_documents
adaptive_rag Same as QA-RAG, but /v2/answer executes geometric expansion strategy under the hood

The Document Indexing template intentionally omits /v2/answer and /v2/summarize because it lacks an LLM integration. The Adaptive RAG template exposes the same interface as standard QA-RAG, making it a drop-in replacement for token-constrained environments.

Practical Usage Examples

Running Question-Answering RAG

Navigate to the template directory and launch via Docker Compose:

cd templates/question_answering_rag
docker compose up --build

Query the RAG endpoint:

curl -X POST http://localhost:8000/v2/answer \
  -H 'Content-Type: application/json' \
  -d '{"prompt":"What is the start date of the contract?"}'

Running Document Indexing

Start the vector storage service:

cd templates/document_indexing
docker compose up --build

Retrieve similar chunks without LLM generation:

curl -X POST http://localhost:8000/v1/retrieve \
  -H 'Content-Type: application/json' \
  -d '{"query":"latest earnings report", "k":5}'

Running Adaptive RAG

Launch the token-efficient variant:

cd templates/adaptive_rag
docker compose up --build

Submit a question with adaptive retrieval:

curl -X POST http://localhost:8000/v2/answer \
  -H 'Content-Type: application/json' \
  -d '{"prompt":"Explain the main risk factors in the 10-K filing."}'

Behind the scenes, answer_with_geometric_rag_strategy_from_index retrieves an initial small chunk set, evaluates LLM confidence, and expands the context window geometrically only when the answer quality is insufficient.

Summary

  • question_answering_rag provides a complete RAG stack with SummaryQuestionAnswerer and QASummaryRestServer for document indexing plus conversational AI.
  • document_indexing offers a lightweight DocumentStoreServer for pure vector retrieval scenarios without LLM dependencies.
  • adaptive_rag uses identical infrastructure to QA-RAG but implements answer_with_geometric_rag_strategy_from_index to minimize token usage through dynamic context expansion.
  • All three templates share persistence configuration and Pathway document processing pipelines, differing only in their REST server implementation and retrieval strategy.

Frequently Asked Questions

When should I use document_indexing instead of question_answering_rag?

Use document_indexing when you need only vector search and retrieval capabilities without LLM generation—such as feeding embeddings to a separate downstream service or building a custom RAG pipeline where you control the LLM invocation yourself. It runs lighter and avoids LLM API costs entirely.

How does adaptive_rag reduce token consumption compared to standard RAG?

The adaptive_rag template calls answer_with_geometric_rag_strategy_from_index (defined in pathway.xpacks.llm.question_answering), which starts with a small number of retrieved documents (configured via start_documents) and expands the result set geometrically only if the LLM signals low confidence. This avoids the "over-fetching" problem where standard RAG retrieves a fixed large chunk set for every query regardless of actual need.

Can I switch between question_answering_rag and adaptive_rag without code changes?

Yes. Both templates use the same App class structure and app.yaml schema. You can copy your configuration from templates/question_answering_rag/app.yaml to templates/adaptive_rag/app.yaml, adjust the adaptive-specific parameters (geometric_factor, start_documents), and run the Adaptive RAG template without modifying Python code.

What files define the core logic for each template?

The primary implementation files are:

These files are accessible directly in the pathwaycom/llm-app repository under their respective template directories.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →