# Context Window Enhancement in RAG: Techniques for Improving Retrieval Context

> Enhance RAG context window with neighboring segments for better narrative flow and answer coherence. Discover techniques implemented in NirDiamant/RAG_Techniques.

- Repository: [NirDiamant/RAG_Techniques](https://github.com/nirdiamant/rag_techniques)
- Tags: deep-dive
- Published: 2026-02-19

---

**Context Window Enhancement in RAG expands retrieved text chunks with neighboring segments to preserve narrative flow and improve answer coherence, implemented in the NirDiamant/RAG_Techniques repository through the `retrieve_with_context_overlap` function.**

Retrieval-Augmented Generation (RAG) pipelines normally retrieve isolated vector chunks based on semantic similarity, often fragmenting the narrative flow essential for accurate language model comprehension. The **Context Window Enhancement** technique solves this limitation by stitching retrieved chunks with their surrounding neighbors, reconstructing coherent windows that maintain the original document's structure. This implementation is available in the open-source **NirDiamant/RAG_Techniques** repository.

## The Limitations of Standard Chunk Retrieval in RAG

Standard RAG implementations split documents into fixed-size chunks and retrieve the top-k most similar vectors using approximate nearest neighbor search. When chunks are short and disconnected, the LLM receives fragmented context that lacks narrative continuity, often resulting in incomplete answers that miss critical surrounding details. This fragmentation particularly impacts documents where meaning spans across chunk boundaries, such as legal contracts, scientific papers, or narrative texts.

## How Context Window Enhancement Works

This technique retrieves the semantically relevant chunk then expands it with neighboring chunks to reconstruct a coherent window that preserves the original document's structure.

### Document Processing and Indexing

The pipeline begins in [`helper_functions.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/helper_functions.py) where `read_pdf_to_string` uses **PyMuPDF** (`fitz`) to extract raw text from PDFs. The `split_text_to_chunks_with_indices` function then creates overlapping chunks using LangChain's `RecursiveCharacterTextSplitter` with configurable parameters (default 400 tokens with 200 token overlap). Each chunk receives a metadata index tracking its original position in the document sequence.

### The Context Overlap Retrieval Algorithm

The `retrieve_with_context_overlap` function, defined in `all_rag_techniques/context_enrichment_window.ipynb`, executes the enhancement through a specific neighbor-fetching workflow:

1. Retrieves top-k chunks using `retriever.get_relevant_documents`
2. Extracts each chunk's positional index from `chunk.metadata["index"]`
3. Calculates neighbor range using `max(0, idx - num_neighbors)` to handle document start boundaries
4. Fetches preceding and succeeding chunks via `get_chunk_by_index`
5. Sorts by index and concatenates, trimming the `chunk_overlap` region to eliminate redundancy

This yields seamless windows containing the target content plus essential surrounding context, improving coherence for downstream LLM generation.

## Implementation in NirDiamant/RAG_Techniques

The repository provides a complete working implementation using FAISS for vector storage and LangChain for orchestration. The core logic resides in the notebook `all_rag_techniques/context_enrichment_window.ipynb`, while shared utilities live in [`helper_functions.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/helper_functions.py).

### Building the Vector Store

```python
from helper_functions import read_pdf_to_string, split_text_to_chunks_with_indices
from langchain_openai import OpenAIEmbeddings
from langchain_community.vectorstores import FAISS

pdf_path = "data/Understanding_Climate_Change.pdf"
content = read_pdf_to_string(pdf_path)

# Chunk size 400 tokens, overlap 200 tokens (as used in the notebook)

docs = split_text_to_chunks_with_indices(content, chunk_size=400, chunk_overlap=200)

embeddings = OpenAIEmbeddings()
vectorstore = FAISS.from_documents(docs, embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 1})

```

### Context-Enriched Retrieval Function

```python
def get_chunk_by_index(vectorstore, target_index: int):
    """Return the chunk whose metadata['index'] matches target_index."""
    all_docs = vectorstore.similarity_search("", k=vectorstore.index.ntotal)
    for doc in all_docs:
        if doc.metadata.get("index") == target_index:
            return doc
    return None


def retrieve_with_context_overlap(vectorstore, retriever, query,
                                 num_neighbors: int = 1,
                                 chunk_overlap: int = 200) -> list[str]:
    """Enrich retrieved chunks with surrounding context."""
    relevant_chunks = retriever.get_relevant_documents(query)
    enriched = []

    for chunk in relevant_chunks:
        idx = chunk.metadata.get("index")
        if idx is None:
            continue

        # neighbour range (inclusive)

        start = max(0, idx - num_neighbors)
        end = idx + num_neighbors + 1

        # fetch neighbour chunks

        neighbours = [get_chunk_by_index(vectorstore, i) for i in range(start, end)]
        neighbours = [c for c in neighbours if c]               # drop Nones

        neighbours.sort(key=lambda d: d.metadata["index"])

        # concatenate while respecting overlap

        txt = neighbours[0].page_content
        for nxt in neighbours[1:]:
            overlap_start = max(0, len(txt) - chunk_overlap)
            txt = txt[:overlap_start] + nxt.page_content
        enriched.append(txt)

    return enriched

```

### Generating Answers with Enriched Context

```python
from langchain_core.prompts import PromptTemplate
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4o-mini")
prompt = PromptTemplate(
    template="Answer the question using ONLY the following context:\n{context}\n\nQuestion: {question}",
    input_variables=["context", "question"],
)

def answer_question(question: str, context: str) -> str:
    chain = prompt | llm
    return chain.invoke({"context": context, "question": question}).content

# Example usage

q = "What role do deforestation and fossil fuels play in climate change?"
enriched = retrieve_with_context_overlap(vectorstore, retriever, q)[0]
print(answer_question(q, enriched))

```

## Architectural Benefits of Context Window Enhancement

Implementing **Context Window Enhancement** provides specific advantages over standard retrieval:

- **Preserves Narrative Flow**: By stitching adjacent chunks using the overlap trimming logic in `retrieve_with_context_overlap`, the model receives full sentences and paragraphs rather than disjoint fragments.
- **Minimal Retrieval Overhead**: Only a small, fixed number of additional chunks (typically 1-2 per side, controlled by the `num_neighbors` parameter) are fetched, keeping query latency low compared to retrieving larger initial chunks.
- **Configurable Window Size**: Parameters like `num_neighbors`, `chunk_size`, and `chunk_overlap` allow tuning per domain, whether processing legal contracts requiring large windows or scientific articles needing precise granularity.
- **Model-Agnostic**: The enriched text output can be supplied to any LLM (ChatGPT, Claude, Llama-2, etc.) without altering the underlying FAISS retrieval logic in [`helper_functions.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/helper_functions.py).

## Summary

- **Context Window Enhancement** solves the fragmentation problem in standard RAG by expanding retrieved chunks with their neighbors to reconstruct coherent passages.
- The `NirDiamant/RAG_Techniques` implementation uses `retrieve_with_context_overlap` to fetch surrounding chunks via metadata indices and concatenate them with overlap trimming.
- Key configuration parameters include `num_neighbors` (window radius), `chunk_size` (token count), and `chunk_overlap` (redundancy elimination).
- The technique preserves narrative continuity while adding minimal latency, making it suitable for production RAG pipelines using FAISS and LangChain.

## Frequently Asked Questions

### What is Context Window Enhancement in RAG?

**Context Window Enhancement** is a retrieval technique that expands semantically relevant chunks with their neighboring document segments before passing them to the language model. This approach reconstructs broader contextual windows that preserve the original document's narrative flow, addressing the fragmentation issues inherent in standard chunk-based retrieval where isolated segments lose contextual continuity.

### How does the context overlap algorithm handle document boundaries?

The `retrieve_with_context_overlap` function handles document start boundaries by using `max(0, idx - num_neighbors)` to prevent negative indices. For end-of-document scenarios, the algorithm filters out None values returned by `get_chunk_by_index` when indices exceed the document length, effectively dropping non-existent neighbors while preserving valid surrounding chunks within the available range.

### What are the performance implications of fetching neighboring chunks?

Fetching neighboring chunks adds minimal overhead because the system retrieves only a small, fixed number of additional chunks (typically 1-2 per side, controlled by the `num_neighbors` parameter). Since these lookups use exact index matches via `get_chunk_by_index` rather than expensive similarity searches, latency remains low while significantly improving context coherence for downstream LLM generation.

### Can Context Window Enhancement work with other retrieval methods?

Yes, **Context Window Enhancement** is retrieval-method agnostic. While the NirDiamant/RAG_Techniques implementation uses FAISS for vector search, the `retrieve_with_context_overlap` function can wrap any LangChain retriever—including those using Chroma, Pinecone, or Weaviate—as long as the retrieved chunks retain metadata indices tracking their original document positions.