Implementing RAG with Vector Databases from Scratch: A Complete Technical Guide

Build a production-ready Retrieval-Augmented Generation pipeline using only NumPy and standard Python libraries, without external vector database services.

Implementing RAG with vector databases from scratch requires understanding how to bridge document retrieval with language model generation using minimal dependencies. The rohitg00/ai-engineering-from-scratch repository provides a comprehensive curriculum that constructs this pipeline using only allowed Python libraries: numpy, torch, h5py, zstandard, and safetensors. This guide distills the repository's seven-stage architecture into actionable code you can run immediately.

The 7-Stage RAG Architecture

The curriculum breaks down RAG implementation into discrete, swappable components. Each stage is documented in specific lesson files within the repository's phased structure.

Stage 1: Document Ingestion and Chunking

Raw documents undergo sentence-aware chunking to preserve context while respecting embedding model constraints. The repository emphasizes chunk size as a critical hyper-parameter—too small loses semantic context, too large dilutes relevance.

According to phases/19-capstone-projects/64-chunking-strategies-advanced/docs/en.md, the recommended approach splits text at sentence boundaries using regex patterns:

import re

def chunk_text(text: str, max_len: int = 500) -> list[str]:
    """Split `text` into chunks no longer than `max_len` characters,
    respecting sentence boundaries."""
    sentences = re.split(r'(?<=[.!?])\s+', text)
    chunks, cur = [], ''
    for s in sentences:
        if len(cur) + len(s) + 1 <= max_len:
            cur = f'{cur} {s}'.strip()
        else:
            chunks.append(cur)
            cur = s
    if cur:
        chunks.append(cur)
    return chunks

Stage 2: Embedding Generation

Each chunk converts to a dense vector representation using encoder models. The repository's site/data.js defines embeddings as "dense vector representations" and references implementations in phases/11-llm-engineering/04-embeddings/docs/en.md. While the curriculum includes pretrained sentence-transformers, the from-scratch approach uses NumPy-based encoders for pedagogical clarity.

Stage 3: Vector Store Construction

The vector database implementation stores embeddings in a simple in-memory NumPy array rather than external services. As defined in site/data.js, retrieval relies on approximate nearest-neighbor search using brute-force cosine similarity.

This approach appears in phases/11-llm-engineering/06-rag-fundamentals/docs/en.md, which establishes the foundational pattern of dense vector storage and similarity-based retrieval.

Given a user query, the system embeds the query using the same encoder, then computes cosine similarity against the vector store to fetch top-k chunks. The repository extends this basic retrieval with hybrid search (BM25 + dense vectors) documented in phases/19-capstone-projects/65-hybrid-retrieval-bm25-dense/docs/en.md.

Stage 5: Reranking with Cross-Encoders

Retrieved chunks undergo cross-encoder reranking to improve precision. This stage, detailed in phases/19-capstone-projects/66-reranker-cross-encoder/docs/en.md, uses a more expensive model to re-score the initial retrieval results, significantly improving answer relevance before generation.

Stage 6: Prompt Construction and LLM Generation

Selected chunks concatenate into a structured prompt with citations, then feed into a language model. The "RAG pipeline" entry in site/data.js links to phases/19-capstone-projects/69-end-to-end-rag-system/docs/en.md, which demonstrates wiring retrieval output to generation input using a minimal LLaMA-style model built from scratch.

Stage 7: Evaluation and Metrics

The pipeline closes with rigorous evaluation using Recall@k, MRR (Mean Reciprocal Rank), nDCG, and RAGAS-style faithfulness metrics. Lesson 68 (phases/19-capstone-projects/68-rag-eval-precision-recall/docs/en.md) provides the evaluation harness for measuring retrieval quality and answer relevance.

Building a Vector Store from Scratch

The repository's pedagogical approach avoids external vector databases in favor of NumPy-based implementations. Here is the complete retrieval system using only standard scientific Python libraries:

Cosine Similarity Retrieval Engine

import numpy as np

# 1. Encode chunks (here we use a dummy encoder; replace with a real model)

def dummy_encoder(chunk: str) -> np.ndarray:
    rng = np.random.default_rng(abs(hash(chunk)) % (2**32))
    return rng.normal(size=768)          # 768-dim embedding

# 2. Build the store

def build_store(chunks: list[str]) -> tuple[np.ndarray, list[str]]:
    vectors = np.stack([dummy_encoder(c) for c in chunks])
    return vectors, chunks

# 3. Retrieve top-k most similar chunks for a query

def retrieve(query: str, vectors: np.ndarray, chunks: list[str], k: int = 5) -> list[str]:
    q_vec = dummy_encoder(query)
    # Cosine similarity (dot / norms)

    sims = vectors @ q_vec / (np.linalg.norm(vectors, axis=1) * np.linalg.norm(q_vec))
    top_idx = np.argsort(sims)[-k:][::-1]
    return [chunks[i] for i in top_idx]

The build_store function creates a vectors.npy matrix where row indices map to chunk indices, enabling O(n) memory footprint with O(n) query time—sufficient for educational datasets before scaling to FAISS or HNSW.

End-to-End RAG Pipeline Implementation

The complete pipeline from raw documents to generated answers appears in phases/19-capstone-projects/69-end-to-end-rag-system/code/main.py. This implementation wires together all seven stages:

def rag_pipeline(query: str, raw_docs: list[str]) -> str:
    # 0️⃣ Chunk all documents

    all_chunks = []
    for doc in raw_docs:
        all_chunks.extend(chunk_text(doc))

    # 1️⃣ Build vector store

    vectors, chunks = build_store(all_chunks)

    # 2️⃣ Retrieve relevant pieces

    retrieved = retrieve(query, vectors, chunks, k=3)

    # 3️⃣ Assemble prompt (simple concatenation)

    prompt = f"Answer the question using the following excerpts:\n\n"
    for i, c in enumerate(retrieved, 1):
        prompt += f"[{i}] {c}\n"
    prompt += f"\nQuestion: {query}\nAnswer:"

    # 4️⃣ Generate (here we just echo the prompt; replace with LLM call)

    return prompt   # in practice feed `prompt` to a LLM (e.g., a toy transformer from the repo)

Key Source Files in the Curriculum

File Path Purpose Lesson Reference
phases/11-llm-engineering/06-rag-fundamentals/docs/en.md Core RAG concepts and retrieval patterns Lesson 06
phases/11-llm-engineering/04-embeddings/docs/en.md Dense vector generation theory Lesson 04
phases/19-capstone-projects/64-chunking-strategies-advanced/docs/en.md Advanced text segmentation strategies Lesson 64
phases/19-capstone-projects/65-hybrid-retrieval-bm25-dense/docs/en.md Lexical + semantic retrieval fusion Lesson 65
phases/19-capstone-projects/66-reranker-cross-encoder/docs/en.md Cross-encoder reranking implementation Lesson 66
phases/19-capstone-projects/68-rag-eval-precision-recall/docs/en.md Evaluation metrics and benchmarking Lesson 68
phases/19-capstone-projects/69-end-to-end-rag-system/code/main.py Reference implementation of complete pipeline Lesson 69
site/data.js Metadata definitions for vector databases and RAG patterns Curriculum index

Summary

  • RAG combines retrieval and generation by fetching relevant document chunks before prompting a language model, reducing hallucinations and grounding answers in source material.
  • Vector stores can be built from scratch using NumPy arrays and cosine similarity, as demonstrated in the repository's Lesson 06 and site/data.js definitions.
  • Chunking strategy critically impacts retrieval quality—the curriculum provides sentence-aware splitting functions in Lesson 64 to balance context preservation with embedding constraints.
  • Hybrid retrieval and reranking significantly improve precision; Lessons 65 and 66 teach BM25+dense fusion and cross-encoder scoring before generation.
  • Evaluation requires specific metrics like Recall@k and faithfulness scores, implemented in Lesson 68's evaluation harness.

Frequently Asked Questions

What dependencies are required to implement RAG from scratch according to this curriculum?

The rohitg00/ai-engineering-from-scratch repository restricts implementations to five allowed libraries: numpy for vector operations, torch for neural network components, h5py for storage serialization, zstandard for compression, and safetensors for model weight handling. No external vector database services like Pinecone or Weaviate are required for the foundational implementation.

How does the repository handle document chunking for RAG?

The curriculum implements sentence-boundary chunking using regex split patterns that respect punctuation while enforcing maximum character lengths. Located in phases/19-capstone-projects/64-chunking-strategies-advanced/docs/en.md, this approach prevents mid-sentence cuts that degrade retrieval quality, using a sliding window approach where chunks overlap to preserve context across boundaries.

Can this from-scratch implementation scale to production workloads?

The NumPy-based vector store provides O(n) query complexity suitable for educational datasets and prototyping. The repository explicitly designs components to be swappable—once you understand the brute-force cosine similarity implementation in phases/11-llm-engineering/06-rag-fundamentals/docs/en.md, you can replace the build_store and retrieve functions with FAISS or HNSW indices while maintaining the same chunking, reranking, and evaluation pipeline established in the capstone lessons.

What evaluation metrics does the curriculum use for RAG systems?

Lesson 68 (phases/19-capstone-projects/68-rag-eval-precision-recall/docs/en.md) implements Recall@k, MRR (Mean Reciprocal Rank), and nDCG for retrieval evaluation, plus faithfulness and answer relevance scores for generation quality. These metrics measure both whether the system retrieves the correct chunks and whether the final generated answer accurately reflects the retrieved content without hallucination.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →