# Build RAG Pipelines with Chunking and Reranking: A Pure Python Implementation

> Build RAG pipelines with Python using sliding-window chunking, TF-IDF embeddings, and cosine-similarity reranking. Get relevant context for generation with this pure Python implementation.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-26

---

**The rohitg00/ai-engineering-from-scratch repository implements a complete Retrieval-Augmented Generation (RAG) pipeline using sliding-window chunking, TF-IDF embeddings, and cosine-similarity reranking to retrieve and rank relevant context before generation.**

Building RAG pipelines with chunking and reranking from first principles demystifies how production retrieval systems operate under the hood. This educational implementation in [`phases/11-llm-engineering/06-rag/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/06-rag/code/main.py) breaks the pipeline into four transparent stages—chunking, embedding, reranking, and generation—using pure Python without external vector database dependencies.

## Pipeline Architecture

The reference implementation organizes the RAG workflow into three core components that mirror enterprise architectures. **Chunking** splits raw documents into overlapping windows via the `chunk_text` function. **Embedding** converts these chunks into TF-IDF vectors using `tfidf_embed`, keeping the linear algebra explicit and debuggable. **Reranking** occurs during the `search` operation, which computes cosine similarity between query and chunk vectors to return the top-k most relevant segments. Finally, `build_rag_prompt` formats retrieved chunks for the generation stage.

This modular design allows you to replace individual components—swapping TF-IDF for `text-embedding-3-small` or the in-memory store for FAISS—without rewriting the orchestration logic.

## Document Chunking Strategy

Effective chunking balances context preservation with embedding granularity. The implementation uses a sliding-window approach that prevents semantic boundaries from being severed while maintaining uniform chunk sizes.

```python
def chunk_text(text, chunk_size=200, overlap=50):
    words = text.split()
    chunks = []
    start = 0
    while start < len(words):
        end = start + chunk_size
        chunk = " ".join(words[start:end])
        chunks.append(chunk)
        start += chunk_size - overlap
    return chunks

```

The default configuration creates 200-word chunks with 50-word overlap, ensuring continuity between segments. This parameterization lives in the `RAGPipeline` class initialization and directly impacts retrieval recall—larger chunks preserve more context but reduce specificity, while smaller chunks increase granularity but risk fragmenting key concepts.

## TF-IDF Embedding and Vocabulary Management

Before storing chunks, the pipeline builds a global vocabulary and computes inverse document frequency (IDF) weights. This happens during the indexing phase through three coordinated functions in [`phases/11-llm-engineering/06-rag/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/06-rag/code/main.py).

```python
def build_vocabulary(documents):
    vocab = set()
    for doc in documents:
        vocab.update(doc.lower().split())
    return sorted(vocab)

def compute_tf(text, vocab):
    words = text.lower().split()
    count = Counter(words)
    total = len(words)
    return [count.get(word, 0) / total for word in vocab]

def compute_idf(documents, vocab):
    n = len(documents)
    idf = []
    for word in vocab:
        doc_count = sum(1 for doc in documents if word in doc.lower().split())
        idf.append(math.log((n + 1) / (doc_count + 1)) + 1)
    return idf

def tfidf_embed(text, vocab, idf):
    tf = compute_tf(text, vocab)
    return [t * i for t, i in zip(tf, idf)]

```

The `tfidf_embed` function generates sparse vectors where term frequency is weighted by rarity across the corpus. While neural embeddings (e.g., SentenceTransformers) capture semantic relationships better, this TF-IDF approach requires zero external dependencies and makes the vector math inspectable for educational purposes.

## Reranking with Cosine Similarity

The reranking stage determines which chunks most closely match the user query. The `search` function performs a brute-force comparison between the query vector and all stored chunk vectors, sorting by similarity score.

```python
def cosine_similarity(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    norm_a = math.sqrt(sum(x * x for x in a))
    norm_b = math.sqrt(sum(x * x for x in b))
    if norm_a == 0 or norm_b == 0:
        return 0.0
    return dot / (norm_a * norm_b)

def search(query_embedding, stored_embeddings, top_k=5):
    scores = [(i, cosine_similarity(query_embedding, emb))
              for i, emb in enumerate(stored_embeddings)]
    scores.sort(key=lambda x: x[1], reverse=True)
    return scores[:top_k]

```

This explicit reranking step returns the top-k chunks (default 5) ordered by descending similarity. In production systems, this logic often moves to dedicated reranking models (cross-encoders) after initial retrieval, but the cosine-similarity approach here demonstrates the fundamental scoring mechanism that orders candidate documents.

## Prompt Construction and Generation

Once reranked chunks are selected, `build_rag_prompt` injects them into a structured template that constrains the model to use only provided context.

```python
def build_rag_prompt(query, retrieved_chunks):
    context = "\n\n---\n\n".join(
        f"[Source {i+1}]\n{chunk}" for i, chunk in enumerate(retrieved_chunks)
    )
    return f"""Answer the question based ONLY on the following context.
If the context doesn't contain enough information, say "I don't have enough information."

Context:
{context}

Question: {query}

Answer:"""

```

The `simple_generate` function simulates LLM behavior by selecting the sentence with maximum word overlap with the query. In a production deployment, you would replace this with an actual API call to GPT-4, Claude, or another LLM while preserving the same prompt structure.

## End-to-End Pipeline Usage

The `RAGPipeline` class orchestrates indexing and querying through two primary methods: `index` for offline document processing and `query` for online retrieval.

```python
from phases.11_llm_engineering.06_rag.code.main import RAGPipeline

docs = ["Enterprise refund policy details...", "Customer service guidelines..."]
source_names = ["policy_doc", "service_doc"]

pipeline = RAGPipeline(chunk_size=50, overlap=10, top_k=3)
pipeline.index(docs, source_names)

result = pipeline.query("What is the refund policy for enterprise customers?")
print("Answer:", result["answer"])
print("Retrieved chunks:", result["retrieved"])

```

The `index` method handles chunking, vocabulary construction, and vector storage, while `query` embeds the input, executes the reranking search, and returns both the generated response and source chunks for provenance.

## Summary

- **Chunking** uses sliding windows with configurable overlap (default 200 words, 50 overlap) to preserve context boundaries while creating embeddable units.
- **TF-IDF embedding** provides a transparent, dependency-free vectorization method via `tfidf_embed`, though neural embeddings can be substituted for semantic search.
- **Reranking** occurs through cosine similarity scoring in the `search` function, explicitly ordering candidates by relevance before prompt construction.
- **Modular architecture** allows individual components to be swapped for production-grade alternatives (FAISS, Pinecone, OpenAI embeddings) without changing the pipeline flow.

## Frequently Asked Questions

### What is the optimal chunk size for RAG pipelines?

The optimal chunk size depends on your document structure and embedding model context windows. The implementation defaults to 200 words with 50-word overlap, which works well for TF-IDF vectors. For neural embeddings like `text-embedding-3-large`, 512-1024 tokens often performs better, preserving semantic coherence while maintaining retrieval precision.

### How does reranking improve retrieval accuracy?

Reranking refines the candidate set by scoring chunks against the specific query rather than relying solely on approximate nearest neighbors. The repository's `search` function uses cosine similarity to reorder all indexed chunks by relevance, ensuring the top-k results provided to the LLM are the most contextually similar to the user's question.

### Can this implementation scale to production workloads?

The current brute-force search and in-memory storage work for thousands of chunks but will bottleneck at scale. For production, replace the `search` function's linear scan with FAISS or ChromaDB for approximate nearest neighbors, and swap the TF-IDF embedder for a neural model. The `RAGPipeline` class structure remains valid regardless of backend changes.

### Why use TF-IDF instead of neural embeddings?

TF-IDF provides interpretable, deterministic vectors that make debugging and educational analysis straightforward. The repository uses it to demonstrate the mathematical foundations of retrieval without requiring GPU resources or external API calls. For production RAG systems, you should migrate to neural embeddings (OpenAI, Cohere, or open-source SentenceTransformers) to capture semantic meaning beyond exact term matching.