Build RAG Pipelines with Chunking and Reranking: A Pure Python Implementation
The rohitg00/ai-engineering-from-scratch repository implements a complete Retrieval-Augmented Generation (RAG) pipeline using sliding-window chunking, TF-IDF embeddings, and cosine-similarity reranking to retrieve and rank relevant context before generation.
Building RAG pipelines with chunking and reranking from first principles demystifies how production retrieval systems operate under the hood. This educational implementation in phases/11-llm-engineering/06-rag/code/main.py breaks the pipeline into four transparent stages—chunking, embedding, reranking, and generation—using pure Python without external vector database dependencies.
Pipeline Architecture
The reference implementation organizes the RAG workflow into three core components that mirror enterprise architectures. Chunking splits raw documents into overlapping windows via the chunk_text function. Embedding converts these chunks into TF-IDF vectors using tfidf_embed, keeping the linear algebra explicit and debuggable. Reranking occurs during the search operation, which computes cosine similarity between query and chunk vectors to return the top-k most relevant segments. Finally, build_rag_prompt formats retrieved chunks for the generation stage.
This modular design allows you to replace individual components—swapping TF-IDF for text-embedding-3-small or the in-memory store for FAISS—without rewriting the orchestration logic.
Document Chunking Strategy
Effective chunking balances context preservation with embedding granularity. The implementation uses a sliding-window approach that prevents semantic boundaries from being severed while maintaining uniform chunk sizes.
def chunk_text(text, chunk_size=200, overlap=50):
words = text.split()
chunks = []
start = 0
while start < len(words):
end = start + chunk_size
chunk = " ".join(words[start:end])
chunks.append(chunk)
start += chunk_size - overlap
return chunks
The default configuration creates 200-word chunks with 50-word overlap, ensuring continuity between segments. This parameterization lives in the RAGPipeline class initialization and directly impacts retrieval recall—larger chunks preserve more context but reduce specificity, while smaller chunks increase granularity but risk fragmenting key concepts.
TF-IDF Embedding and Vocabulary Management
Before storing chunks, the pipeline builds a global vocabulary and computes inverse document frequency (IDF) weights. This happens during the indexing phase through three coordinated functions in phases/11-llm-engineering/06-rag/code/main.py.
def build_vocabulary(documents):
vocab = set()
for doc in documents:
vocab.update(doc.lower().split())
return sorted(vocab)
def compute_tf(text, vocab):
words = text.lower().split()
count = Counter(words)
total = len(words)
return [count.get(word, 0) / total for word in vocab]
def compute_idf(documents, vocab):
n = len(documents)
idf = []
for word in vocab:
doc_count = sum(1 for doc in documents if word in doc.lower().split())
idf.append(math.log((n + 1) / (doc_count + 1)) + 1)
return idf
def tfidf_embed(text, vocab, idf):
tf = compute_tf(text, vocab)
return [t * i for t, i in zip(tf, idf)]
The tfidf_embed function generates sparse vectors where term frequency is weighted by rarity across the corpus. While neural embeddings (e.g., SentenceTransformers) capture semantic relationships better, this TF-IDF approach requires zero external dependencies and makes the vector math inspectable for educational purposes.
Reranking with Cosine Similarity
The reranking stage determines which chunks most closely match the user query. The search function performs a brute-force comparison between the query vector and all stored chunk vectors, sorting by similarity score.
def cosine_similarity(a, b):
dot = sum(x * y for x, y in zip(a, b))
norm_a = math.sqrt(sum(x * x for x in a))
norm_b = math.sqrt(sum(x * x for x in b))
if norm_a == 0 or norm_b == 0:
return 0.0
return dot / (norm_a * norm_b)
def search(query_embedding, stored_embeddings, top_k=5):
scores = [(i, cosine_similarity(query_embedding, emb))
for i, emb in enumerate(stored_embeddings)]
scores.sort(key=lambda x: x[1], reverse=True)
return scores[:top_k]
This explicit reranking step returns the top-k chunks (default 5) ordered by descending similarity. In production systems, this logic often moves to dedicated reranking models (cross-encoders) after initial retrieval, but the cosine-similarity approach here demonstrates the fundamental scoring mechanism that orders candidate documents.
Prompt Construction and Generation
Once reranked chunks are selected, build_rag_prompt injects them into a structured template that constrains the model to use only provided context.
def build_rag_prompt(query, retrieved_chunks):
context = "\n\n---\n\n".join(
f"[Source {i+1}]\n{chunk}" for i, chunk in enumerate(retrieved_chunks)
)
return f"""Answer the question based ONLY on the following context.
If the context doesn't contain enough information, say "I don't have enough information."
Context:
{context}
Question: {query}
Answer:"""
The simple_generate function simulates LLM behavior by selecting the sentence with maximum word overlap with the query. In a production deployment, you would replace this with an actual API call to GPT-4, Claude, or another LLM while preserving the same prompt structure.
End-to-End Pipeline Usage
The RAGPipeline class orchestrates indexing and querying through two primary methods: index for offline document processing and query for online retrieval.
from phases.11_llm_engineering.06_rag.code.main import RAGPipeline
docs = ["Enterprise refund policy details...", "Customer service guidelines..."]
source_names = ["policy_doc", "service_doc"]
pipeline = RAGPipeline(chunk_size=50, overlap=10, top_k=3)
pipeline.index(docs, source_names)
result = pipeline.query("What is the refund policy for enterprise customers?")
print("Answer:", result["answer"])
print("Retrieved chunks:", result["retrieved"])
The index method handles chunking, vocabulary construction, and vector storage, while query embeds the input, executes the reranking search, and returns both the generated response and source chunks for provenance.
Summary
- Chunking uses sliding windows with configurable overlap (default 200 words, 50 overlap) to preserve context boundaries while creating embeddable units.
- TF-IDF embedding provides a transparent, dependency-free vectorization method via
tfidf_embed, though neural embeddings can be substituted for semantic search. - Reranking occurs through cosine similarity scoring in the
searchfunction, explicitly ordering candidates by relevance before prompt construction. - Modular architecture allows individual components to be swapped for production-grade alternatives (FAISS, Pinecone, OpenAI embeddings) without changing the pipeline flow.
Frequently Asked Questions
What is the optimal chunk size for RAG pipelines?
The optimal chunk size depends on your document structure and embedding model context windows. The implementation defaults to 200 words with 50-word overlap, which works well for TF-IDF vectors. For neural embeddings like text-embedding-3-large, 512-1024 tokens often performs better, preserving semantic coherence while maintaining retrieval precision.
How does reranking improve retrieval accuracy?
Reranking refines the candidate set by scoring chunks against the specific query rather than relying solely on approximate nearest neighbors. The repository's search function uses cosine similarity to reorder all indexed chunks by relevance, ensuring the top-k results provided to the LLM are the most contextually similar to the user's question.
Can this implementation scale to production workloads?
The current brute-force search and in-memory storage work for thousands of chunks but will bottleneck at scale. For production, replace the search function's linear scan with FAISS or ChromaDB for approximate nearest neighbors, and swap the TF-IDF embedder for a neural model. The RAGPipeline class structure remains valid regardless of backend changes.
Why use TF-IDF instead of neural embeddings?
TF-IDF provides interpretable, deterministic vectors that make debugging and educational analysis straightforward. The repository uses it to demonstrate the mathematical foundations of retrieval without requiring GPU resources or external API calls. For production RAG systems, you should migrate to neural embeddings (OpenAI, Cohere, or open-source SentenceTransformers) to capture semantic meaning beyond exact term matching.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →