Document Augmentation Strategies for RAG: Boost Retrieval with Synthetic Questions
Document augmentation enriches RAG vector stores by generating synthetic question-style documents for each text chunk, dramatically improving query-to-context alignment without requiring runtime LLM calls during retrieval.
The NirDiamant/RAG_Techniques repository implements a production-ready document augmentation pipeline that transforms sparse corpora into densely searchable knowledge bases. By automatically generating diverse interrogative sentences for every source chunk, this technique bridges the semantic gap between how users ask questions and how information is originally phrased.
Why Document Augmentation Matters for RAG
Standard RAG pipelines index raw text chunks directly, assuming user queries will semantically align with the stored prose. In practice, users ask questions ("What causes climate change?") while documents contain statements ("Climate change is caused by greenhouse gases"). This formulation mismatch reduces retrieval accuracy.
The document augmentation strategy implemented in all_rag_techniques_runnable_scripts/document_augmentation.py solves this by populating the vector space with question-style embeddings that mirror natural user queries, increasing the probability of high-similarity matches during retrieval.
The Three-Stage Document Augmentation Pipeline
Stage 1: Intelligent Chunking with Context Overlap
The pipeline begins by splitting source material into overlapping chunks to preserve context across boundaries. In document_augmentation.py, the split_document function processes raw text using two configurable strategies:
- Document-level chunking: Splits text into chunks up to 4,000 tokens with
DOCUMENT_OVERLAP_TOKENSoverlap - Fragment-level chunking: Creates smaller 128-token chunks with
FRAGMENT_OVERLAP_TOKENSfor granular retrieval
This overlap ensures that questions generated at chunk boundaries retain necessary surrounding context.
Stage 2: Synthetic Question Generation
For each chunk, the generate_questions function (lines 61-73) uses ChatOpenAI to synthesize at least 40 diverse interrogative sentences. The generation process follows strict quality controls:
- Prompt engineering: Forces the LLM to output only question-style sentences targeting the chunk's core concepts
- Deduplication: The
clean_and_filter_questionsfunction removes near-duplicate queries using semantic similarity filtering - Metadata tagging: Each generated question receives metadata marking it as type
AUGMENTEDand linking it to its parent chunk via the"text"field
The QUESTION_GENERATION enum controls whether questions are generated from full documents or smaller fragments, allowing granular tuning based on source material density.
Stage 3: Building the Augmented FAISS Index
The final stage constructs a unified vector store containing both original and synthetic content. In document_augmentation.py (lines 46-50), the pipeline:
- Wraps original chunks as
Documentobjects withmetadata={"type": "ORIGINAL"} - Creates separate
Documentobjects for each generated question withmetadata={"type": "AUGMENTED", "text": <original_chunk_content>} - Embeds all documents using
OpenAIEmbeddingsand indexes them withFAISS.from_documents
At retrieval time, the system fetches the most similar augmented document (typically a synthetic question), then extracts the original parent chunk from the "text" metadata field to provide full context to the generation LLM.
Implementation Walkthrough
To implement document augmentation in your RAG pipeline using the NirDiamant/RAG_Techniques approach:
# Step 1: Initialize the document processor
from helper_functions import read_pdf_to_string
from document_augmentation import DocumentProcessor, OpenAIEmbeddings
# Load source material
content = read_pdf_to_string("data/research_paper.pdf")
embeddings = OpenAIEmbeddings()
# Build the augmented retriever (handles chunking, question generation, and indexing)
processor = DocumentProcessor(content, embeddings)
retriever = processor.run()
# Step 2: Query the augmented index
query = "What are the thermal effects on marine ecosystems?"
# Retrieve relevant documents (returns augmented questions and original chunks)
docs = retriever.get_relevant_documents(query)
# Extract the original context from metadata
original_context = docs[0].metadata["text"]
doc_type = docs[0].metadata["type"] # "AUGMENTED" or "ORIGINAL"
# Step 3: Generate the final answer
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(model="gpt-4o-mini")
response = llm.invoke(
f"Context: {original_context}\n\n"
f"Question: {query}\n\n"
f"Provide a concise, accurate answer based solely on the context."
)
print(response.content)
The DocumentProcessor class encapsulates the entire pipeline, automatically handling the split_document logic, generate_questions calls, and FAISS indexing with proper metadata tagging.
Why This Approach Works
Document augmentation improves RAG performance through three key mechanisms:
-
Query-question alignment: By embedding 40+ synthetic questions per chunk, the vector space becomes densely populated with vectors that naturally align with how users actually phrase queries, increasing cosine similarity scores for relevant matches.
-
Context preservation: The overlapping chunk strategy (
DOCUMENT_OVERLAP_TOKENSandFRAGMENT_OVERLAP_TOKENS) ensures that synthetic questions reference complete semantic units rather than isolated sentence fragments. -
Scalable retrieval: FAISS indexing supports sub-second nearest-neighbor search across thousands of augmented documents, maintaining low latency even as the index grows through the
OpenAIEmbeddingsvectorization process.
Summary
- Document augmentation generates synthetic question-style documents for each source chunk, bridging the semantic gap between user queries and indexed content.
- The
NirDiamant/RAG_Techniquesimplementation uses a three-stage pipeline: intelligent chunking with overlap, LLM-based question generation (40+ per chunk), and unified FAISS indexing with metadata linking. - Key files include
all_rag_techniques_runnable_scripts/document_augmentation.pyfor the core logic andhelper_functions.pyfor data ingestion. - The technique improves retrieval accuracy by aligning vector embeddings with natural question phrasing while maintaining sub-second search latency through FAISS.
Frequently Asked Questions
How many synthetic questions does the document augmentation pipeline generate per chunk?
The pipeline generates at least 40 diverse questions per chunk through the generate_questions function in document_augmentation.py. These questions undergo deduplication via clean_and_filter_questions to ensure semantic variety before indexing.
What is the difference between ORIGINAL and AUGMENTED metadata types in the FAISS index?
Documents with metadata={"type": "ORIGINAL"} represent the actual source text chunks containing full contextual passages. Documents with metadata={"type": "AUGMENTED"} are the synthetic questions generated from those chunks, which store the parent chunk's content in the "text" field for retrieval resolution.
Does document augmentation increase retrieval latency?
No, the technique maintains sub-second retrieval latency because all synthetic questions are pre-computed and embedded using OpenAIEmbeddings during indexing. The FAISS vector store performs efficient nearest-neighbor search across both original and augmented documents without requiring runtime LLM calls during the retrieval phase.
Can I adjust the chunk size and overlap for different document types?
Yes, the DocumentProcessor class exposes configuration constants for both document-level and fragment-level chunking. You can modify DOCUMENT_OVERLAP_TOKENS and FRAGMENT_OVERLAP_TOKENS in document_augmentation.py to optimize context preservation for dense technical documents versus sparse narrative text.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →