# Document Augmentation Strategies for RAG: Boost Retrieval with Synthetic Questions

> Boost RAG retrieval with document augmentation. Generate synthetic questions for each text chunk to improve query-to-context alignment. Discover efficient RAG techniques today.

- Repository: [NirDiamant/RAG_Techniques](https://github.com/nirdiamant/rag_techniques)
- Tags: deep-dive
- Published: 2026-02-19

---

**Document augmentation enriches RAG vector stores by generating synthetic question-style documents for each text chunk, dramatically improving query-to-context alignment without requiring runtime LLM calls during retrieval.**

The `NirDiamant/RAG_Techniques` repository implements a production-ready document augmentation pipeline that transforms sparse corpora into densely searchable knowledge bases. By automatically generating diverse interrogative sentences for every source chunk, this technique bridges the semantic gap between how users ask questions and how information is originally phrased.

## Why Document Augmentation Matters for RAG

Standard RAG pipelines index raw text chunks directly, assuming user queries will semantically align with the stored prose. In practice, users ask questions ("What causes climate change?") while documents contain statements ("Climate change is caused by greenhouse gases"). This formulation mismatch reduces retrieval accuracy.

The document augmentation strategy implemented in [`all_rag_techniques_runnable_scripts/document_augmentation.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/document_augmentation.py) solves this by populating the vector space with **question-style embeddings** that mirror natural user queries, increasing the probability of high-similarity matches during retrieval.

## The Three-Stage Document Augmentation Pipeline

### Stage 1: Intelligent Chunking with Context Overlap

The pipeline begins by splitting source material into overlapping chunks to preserve context across boundaries. In [`document_augmentation.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/document_augmentation.py), the `split_document` function processes raw text using two configurable strategies:

- **Document-level chunking**: Splits text into chunks up to 4,000 tokens with `DOCUMENT_OVERLAP_TOKENS` overlap
- **Fragment-level chunking**: Creates smaller 128-token chunks with `FRAGMENT_OVERLAP_TOKENS` for granular retrieval

This overlap ensures that questions generated at chunk boundaries retain necessary surrounding context.

### Stage 2: Synthetic Question Generation

For each chunk, the `generate_questions` function (lines 61-73) uses `ChatOpenAI` to synthesize at least 40 diverse interrogative sentences. The generation process follows strict quality controls:

1. **Prompt engineering**: Forces the LLM to output only question-style sentences targeting the chunk's core concepts
2. **Deduplication**: The `clean_and_filter_questions` function removes near-duplicate queries using semantic similarity filtering
3. **Metadata tagging**: Each generated question receives metadata marking it as type `AUGMENTED` and linking it to its parent chunk via the `"text"` field

The `QUESTION_GENERATION` enum controls whether questions are generated from full documents or smaller fragments, allowing granular tuning based on source material density.

### Stage 3: Building the Augmented FAISS Index

The final stage constructs a unified vector store containing both original and synthetic content. In [`document_augmentation.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/document_augmentation.py) (lines 46-50), the pipeline:

1. Wraps original chunks as `Document` objects with `metadata={"type": "ORIGINAL"}`
2. Creates separate `Document` objects for each generated question with `metadata={"type": "AUGMENTED", "text": <original_chunk_content>}`
3. Embeds all documents using `OpenAIEmbeddings` and indexes them with `FAISS.from_documents`

At retrieval time, the system fetches the most similar augmented document (typically a synthetic question), then extracts the original parent chunk from the `"text"` metadata field to provide full context to the generation LLM.

## Implementation Walkthrough

To implement document augmentation in your RAG pipeline using the `NirDiamant/RAG_Techniques` approach:

```python

# Step 1: Initialize the document processor

from helper_functions import read_pdf_to_string
from document_augmentation import DocumentProcessor, OpenAIEmbeddings

# Load source material

content = read_pdf_to_string("data/research_paper.pdf")
embeddings = OpenAIEmbeddings()

# Build the augmented retriever (handles chunking, question generation, and indexing)

processor = DocumentProcessor(content, embeddings)
retriever = processor.run()

```

```python

# Step 2: Query the augmented index

query = "What are the thermal effects on marine ecosystems?"

# Retrieve relevant documents (returns augmented questions and original chunks)

docs = retriever.get_relevant_documents(query)

# Extract the original context from metadata

original_context = docs[0].metadata["text"]
doc_type = docs[0].metadata["type"]  # "AUGMENTED" or "ORIGINAL"

```

```python

# Step 3: Generate the final answer

from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4o-mini")
response = llm.invoke(
    f"Context: {original_context}\n\n"
    f"Question: {query}\n\n"
    f"Provide a concise, accurate answer based solely on the context."
)

print(response.content)

```

The `DocumentProcessor` class encapsulates the entire pipeline, automatically handling the `split_document` logic, `generate_questions` calls, and FAISS indexing with proper metadata tagging.

## Why This Approach Works

Document augmentation improves RAG performance through three key mechanisms:

- **Query-question alignment**: By embedding 40+ synthetic questions per chunk, the vector space becomes densely populated with vectors that naturally align with how users actually phrase queries, increasing cosine similarity scores for relevant matches.

- **Context preservation**: The overlapping chunk strategy (`DOCUMENT_OVERLAP_TOKENS` and `FRAGMENT_OVERLAP_TOKENS`) ensures that synthetic questions reference complete semantic units rather than isolated sentence fragments.

- **Scalable retrieval**: FAISS indexing supports sub-second nearest-neighbor search across thousands of augmented documents, maintaining low latency even as the index grows through the `OpenAIEmbeddings` vectorization process.

## Summary

- Document augmentation generates synthetic question-style documents for each source chunk, bridging the semantic gap between user queries and indexed content.
- The `NirDiamant/RAG_Techniques` implementation uses a three-stage pipeline: intelligent chunking with overlap, LLM-based question generation (40+ per chunk), and unified FAISS indexing with metadata linking.
- Key files include [`all_rag_techniques_runnable_scripts/document_augmentation.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/all_rag_techniques_runnable_scripts/document_augmentation.py) for the core logic and [`helper_functions.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/helper_functions.py) for data ingestion.
- The technique improves retrieval accuracy by aligning vector embeddings with natural question phrasing while maintaining sub-second search latency through FAISS.

## Frequently Asked Questions

### How many synthetic questions does the document augmentation pipeline generate per chunk?

The pipeline generates at least 40 diverse questions per chunk through the `generate_questions` function in [`document_augmentation.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/document_augmentation.py). These questions undergo deduplication via `clean_and_filter_questions` to ensure semantic variety before indexing.

### What is the difference between ORIGINAL and AUGMENTED metadata types in the FAISS index?

Documents with `metadata={"type": "ORIGINAL"}` represent the actual source text chunks containing full contextual passages. Documents with `metadata={"type": "AUGMENTED"}` are the synthetic questions generated from those chunks, which store the parent chunk's content in the `"text"` field for retrieval resolution.

### Does document augmentation increase retrieval latency?

No, the technique maintains sub-second retrieval latency because all synthetic questions are pre-computed and embedded using `OpenAIEmbeddings` during indexing. The FAISS vector store performs efficient nearest-neighbor search across both original and augmented documents without requiring runtime LLM calls during the retrieval phase.

### Can I adjust the chunk size and overlap for different document types?

Yes, the `DocumentProcessor` class exposes configuration constants for both document-level and fragment-level chunking. You can modify `DOCUMENT_OVERLAP_TOKENS` and `FRAGMENT_OVERLAP_TOKENS` in [`document_augmentation.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/document_augmentation.py) to optimize context preservation for dense technical documents versus sparse narrative text.