Key Differences Between question_answering_rag, document_indexing, and adaptive_rag Templates in Pathway LLM-App
The question_answering_rag template provides end-to-end Retrieval-Augmented Generation (RAG) with LLM-based answers, document_indexing serves as a pure vector storage and retrieval service without LLM generation, and adaptive_rag extends standard RAG with a geometric expansion strategy that dynamically grows context size to optimize token usage.
All three templates reside in the templates/ directory of the pathwaycom/llm-app repository and share a common Pathway-based backbone for document processing. While they reuse identical persistence and parsing infrastructure, each targets a distinct operational mode—ranging from full conversational AI to bare-metal vector search to token-efficient adaptive retrieval.
Architectural Overview
Each template follows a configuration-driven architecture centered on an App class defined in app.py. The templates share common persistence logic (pw.PersistenceMode.UDF_CACHING) and filesystem backend configuration (lines 31–61 across all three files), but diverge in their core Pathway components and REST server implementations.
Shared infrastructure includes:
- YAML-based configuration via
app.yaml - Pathway document parsing and embedding pipelines
- Optional filesystem caching (
./Cache) - Docker Compose orchestration files
Core Differences by Template
Question-Answering RAG (question_answering_rag)
Purpose: End-to-end RAG pipeline that indexes documents and answers natural language questions using a hosted LLM.
Key Components:
SummaryQuestionAnswerer– Wraps aDocumentStorewith LLM generation capabilitiesQASummaryRestServer– Exposes conversational endpoints including summarizationAppclass – Instantiatesquestion_answerer: InstanceOf[SummaryQuestionAnswerer]
Implementation: In templates/question_answering_rag/app.py, the application constructs a QASummaryRestServer that routes requests to the SummaryQuestionAnswerer, providing complete document ingestion and chat functionality in a single service.
Document Indexing (document_indexing)
Purpose: Standalone document store service that only indexes documents and provides retrieval endpoints—no LLM generation involved.
Key Components:
DocumentStore– Pure vector index with similarity search and statisticsDocumentStoreServer– Lightweight server exposing only retrieval-related routesAppclass – Instantiatesdocument_store: InstanceOf[DocumentStore]
Implementation: The templates/document_indexing/app.py file creates a DocumentStoreServer rather than a QA server. This template is ideal when you need to embed documents for downstream consumers or want to decouple retrieval from generation.
Adaptive RAG (adaptive_rag)
Purpose: Token-efficient RAG that uses adaptive retrieval to minimize LLM API costs while maintaining answer quality.
Key Components:
SummaryQuestionAnswerer– Same class as standard QA-RAG, but configured with adaptive parametersanswer_with_geometric_rag_strategy_from_index– Dynamically expands retrieved chunks geometrically until confidence thresholds are met- Geometric expansion parameters –
start_documentsandgeometric_factortunable inapp.yaml
Implementation: The templates/adaptive_rag/app.py file overrides the standard answer method to call answer_with_geometric_rag_strategy_from_index (see lines 31–33 in the README). This strategy starts with a small retrieval set and fetches additional chunks only when the LLM indicates insufficient context, reducing token consumption compared to the fixed-context approach in standard QA-RAG.
Configuration and Implementation Details
All three pipelines are driven by app.yaml, but their configuration schemas differ based on server type.
Question-Answering RAG and Adaptive RAG require:
- LLM configuration (OpenAI, Azure, or local)
- Embedder settings
- Index parameters
- Document sources
Document Indexing requires only:
- DocumentStore configuration (sources, parser, splitter, retriever)
App Class Structure:
- QA-RAG & Adaptive:
class App(BaseModel): question_answerer: InstanceOf[SummaryQuestionAnswerer] - Document Indexing:
class App(BaseModel): document_store: InstanceOf[DocumentStore]
Swapping between QA-RAG and Adaptive RAG requires only changing the template directory—both use identical app.yaml structures, with Adaptive RAG adding optional adaptive parameters like start_documents and geometric_factor.
API Endpoints Comparison
| Template | Available Endpoints |
|---|---|
| question_answering_rag | /v1/retrieve (similarity search)/v2/answer (RAG Q&A)/v2/summarize (document summarization)/v1/statistics/v2/list_documents |
| document_indexing | /v1/retrieve/v1/statistics/v2/list_documents |
| adaptive_rag | Same as QA-RAG, but /v2/answer executes geometric expansion strategy under the hood |
The Document Indexing template intentionally omits /v2/answer and /v2/summarize because it lacks an LLM integration. The Adaptive RAG template exposes the same interface as standard QA-RAG, making it a drop-in replacement for token-constrained environments.
Practical Usage Examples
Running Question-Answering RAG
Navigate to the template directory and launch via Docker Compose:
cd templates/question_answering_rag
docker compose up --build
Query the RAG endpoint:
curl -X POST http://localhost:8000/v2/answer \
-H 'Content-Type: application/json' \
-d '{"prompt":"What is the start date of the contract?"}'
Running Document Indexing
Start the vector storage service:
cd templates/document_indexing
docker compose up --build
Retrieve similar chunks without LLM generation:
curl -X POST http://localhost:8000/v1/retrieve \
-H 'Content-Type: application/json' \
-d '{"query":"latest earnings report", "k":5}'
Running Adaptive RAG
Launch the token-efficient variant:
cd templates/adaptive_rag
docker compose up --build
Submit a question with adaptive retrieval:
curl -X POST http://localhost:8000/v2/answer \
-H 'Content-Type: application/json' \
-d '{"prompt":"Explain the main risk factors in the 10-K filing."}'
Behind the scenes, answer_with_geometric_rag_strategy_from_index retrieves an initial small chunk set, evaluates LLM confidence, and expands the context window geometrically only when the answer quality is insufficient.
Summary
- question_answering_rag provides a complete RAG stack with
SummaryQuestionAnswererandQASummaryRestServerfor document indexing plus conversational AI. - document_indexing offers a lightweight
DocumentStoreServerfor pure vector retrieval scenarios without LLM dependencies. - adaptive_rag uses identical infrastructure to QA-RAG but implements
answer_with_geometric_rag_strategy_from_indexto minimize token usage through dynamic context expansion. - All three templates share persistence configuration and Pathway document processing pipelines, differing only in their REST server implementation and retrieval strategy.
Frequently Asked Questions
When should I use document_indexing instead of question_answering_rag?
Use document_indexing when you need only vector search and retrieval capabilities without LLM generation—such as feeding embeddings to a separate downstream service or building a custom RAG pipeline where you control the LLM invocation yourself. It runs lighter and avoids LLM API costs entirely.
How does adaptive_rag reduce token consumption compared to standard RAG?
The adaptive_rag template calls answer_with_geometric_rag_strategy_from_index (defined in pathway.xpacks.llm.question_answering), which starts with a small number of retrieved documents (configured via start_documents) and expands the result set geometrically only if the LLM signals low confidence. This avoids the "over-fetching" problem where standard RAG retrieves a fixed large chunk set for every query regardless of actual need.
Can I switch between question_answering_rag and adaptive_rag without code changes?
Yes. Both templates use the same App class structure and app.yaml schema. You can copy your configuration from templates/question_answering_rag/app.yaml to templates/adaptive_rag/app.yaml, adjust the adaptive-specific parameters (geometric_factor, start_documents), and run the Adaptive RAG template without modifying Python code.
What files define the core logic for each template?
The primary implementation files are:
templates/question_answering_rag/app.py– InstantiatesQASummaryRestServerwithSummaryQuestionAnswerertemplates/document_indexing/app.py– InstantiatesDocumentStoreServerwithDocumentStoretemplates/adaptive_rag/app.py– Identical server setup to QA-RAG but configures the adaptive retrieval strategy
These files are accessible directly in the pathwaycom/llm-app repository under their respective template directories.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →