Document Augmentation Strategies for RAG: Boost Retrieval with Synthetic Questions

Document augmentation enriches RAG vector stores by generating synthetic question-style documents for each text chunk, dramatically improving query-to-context alignment without requiring runtime LLM calls during retrieval.

The NirDiamant/RAG_Techniques repository implements a production-ready document augmentation pipeline that transforms sparse corpora into densely searchable knowledge bases. By automatically generating diverse interrogative sentences for every source chunk, this technique bridges the semantic gap between how users ask questions and how information is originally phrased.

Why Document Augmentation Matters for RAG

Standard RAG pipelines index raw text chunks directly, assuming user queries will semantically align with the stored prose. In practice, users ask questions ("What causes climate change?") while documents contain statements ("Climate change is caused by greenhouse gases"). This formulation mismatch reduces retrieval accuracy.

The document augmentation strategy implemented in all_rag_techniques_runnable_scripts/document_augmentation.py solves this by populating the vector space with question-style embeddings that mirror natural user queries, increasing the probability of high-similarity matches during retrieval.

The Three-Stage Document Augmentation Pipeline

Stage 1: Intelligent Chunking with Context Overlap

The pipeline begins by splitting source material into overlapping chunks to preserve context across boundaries. In document_augmentation.py, the split_document function processes raw text using two configurable strategies:

  • Document-level chunking: Splits text into chunks up to 4,000 tokens with DOCUMENT_OVERLAP_TOKENS overlap
  • Fragment-level chunking: Creates smaller 128-token chunks with FRAGMENT_OVERLAP_TOKENS for granular retrieval

This overlap ensures that questions generated at chunk boundaries retain necessary surrounding context.

Stage 2: Synthetic Question Generation

For each chunk, the generate_questions function (lines 61-73) uses ChatOpenAI to synthesize at least 40 diverse interrogative sentences. The generation process follows strict quality controls:

  1. Prompt engineering: Forces the LLM to output only question-style sentences targeting the chunk's core concepts
  2. Deduplication: The clean_and_filter_questions function removes near-duplicate queries using semantic similarity filtering
  3. Metadata tagging: Each generated question receives metadata marking it as type AUGMENTED and linking it to its parent chunk via the "text" field

The QUESTION_GENERATION enum controls whether questions are generated from full documents or smaller fragments, allowing granular tuning based on source material density.

Stage 3: Building the Augmented FAISS Index

The final stage constructs a unified vector store containing both original and synthetic content. In document_augmentation.py (lines 46-50), the pipeline:

  1. Wraps original chunks as Document objects with metadata={"type": "ORIGINAL"}
  2. Creates separate Document objects for each generated question with metadata={"type": "AUGMENTED", "text": <original_chunk_content>}
  3. Embeds all documents using OpenAIEmbeddings and indexes them with FAISS.from_documents

At retrieval time, the system fetches the most similar augmented document (typically a synthetic question), then extracts the original parent chunk from the "text" metadata field to provide full context to the generation LLM.

Implementation Walkthrough

To implement document augmentation in your RAG pipeline using the NirDiamant/RAG_Techniques approach:


# Step 1: Initialize the document processor

from helper_functions import read_pdf_to_string
from document_augmentation import DocumentProcessor, OpenAIEmbeddings

# Load source material

content = read_pdf_to_string("data/research_paper.pdf")
embeddings = OpenAIEmbeddings()

# Build the augmented retriever (handles chunking, question generation, and indexing)

processor = DocumentProcessor(content, embeddings)
retriever = processor.run()

# Step 2: Query the augmented index

query = "What are the thermal effects on marine ecosystems?"

# Retrieve relevant documents (returns augmented questions and original chunks)

docs = retriever.get_relevant_documents(query)

# Extract the original context from metadata

original_context = docs[0].metadata["text"]
doc_type = docs[0].metadata["type"]  # "AUGMENTED" or "ORIGINAL"

# Step 3: Generate the final answer

from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4o-mini")
response = llm.invoke(
    f"Context: {original_context}\n\n"
    f"Question: {query}\n\n"
    f"Provide a concise, accurate answer based solely on the context."
)

print(response.content)

The DocumentProcessor class encapsulates the entire pipeline, automatically handling the split_document logic, generate_questions calls, and FAISS indexing with proper metadata tagging.

Why This Approach Works

Document augmentation improves RAG performance through three key mechanisms:

  • Query-question alignment: By embedding 40+ synthetic questions per chunk, the vector space becomes densely populated with vectors that naturally align with how users actually phrase queries, increasing cosine similarity scores for relevant matches.

  • Context preservation: The overlapping chunk strategy (DOCUMENT_OVERLAP_TOKENS and FRAGMENT_OVERLAP_TOKENS) ensures that synthetic questions reference complete semantic units rather than isolated sentence fragments.

  • Scalable retrieval: FAISS indexing supports sub-second nearest-neighbor search across thousands of augmented documents, maintaining low latency even as the index grows through the OpenAIEmbeddings vectorization process.

Summary

  • Document augmentation generates synthetic question-style documents for each source chunk, bridging the semantic gap between user queries and indexed content.
  • The NirDiamant/RAG_Techniques implementation uses a three-stage pipeline: intelligent chunking with overlap, LLM-based question generation (40+ per chunk), and unified FAISS indexing with metadata linking.
  • Key files include all_rag_techniques_runnable_scripts/document_augmentation.py for the core logic and helper_functions.py for data ingestion.
  • The technique improves retrieval accuracy by aligning vector embeddings with natural question phrasing while maintaining sub-second search latency through FAISS.

Frequently Asked Questions

How many synthetic questions does the document augmentation pipeline generate per chunk?

The pipeline generates at least 40 diverse questions per chunk through the generate_questions function in document_augmentation.py. These questions undergo deduplication via clean_and_filter_questions to ensure semantic variety before indexing.

What is the difference between ORIGINAL and AUGMENTED metadata types in the FAISS index?

Documents with metadata={"type": "ORIGINAL"} represent the actual source text chunks containing full contextual passages. Documents with metadata={"type": "AUGMENTED"} are the synthetic questions generated from those chunks, which store the parent chunk's content in the "text" field for retrieval resolution.

Does document augmentation increase retrieval latency?

No, the technique maintains sub-second retrieval latency because all synthetic questions are pre-computed and embedded using OpenAIEmbeddings during indexing. The FAISS vector store performs efficient nearest-neighbor search across both original and augmented documents without requiring runtime LLM calls during the retrieval phase.

Can I adjust the chunk size and overlap for different document types?

Yes, the DocumentProcessor class exposes configuration constants for both document-level and fragment-level chunking. You can modify DOCUMENT_OVERLAP_TOKENS and FRAGMENT_OVERLAP_TOKENS in document_augmentation.py to optimize context preservation for dense technical documents versus sparse narrative text.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →