# Implementing RAG with Vector Databases from Scratch: A Complete Technical Guide

> Implement RAG with vector databases from scratch using NumPy and Python. Build a production-ready pipeline without external services. Your complete technical guide for AI engineering.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-26

---

**Build a production-ready Retrieval-Augmented Generation pipeline using only NumPy and standard Python libraries, without external vector database services.**

Implementing RAG with vector databases from scratch requires understanding how to bridge document retrieval with language model generation using minimal dependencies. The `rohitg00/ai-engineering-from-scratch` repository provides a comprehensive curriculum that constructs this pipeline using only allowed Python libraries: `numpy`, `torch`, `h5py`, `zstandard`, and `safetensors`. This guide distills the repository's seven-stage architecture into actionable code you can run immediately.

## The 7-Stage RAG Architecture

The curriculum breaks down RAG implementation into discrete, swappable components. Each stage is documented in specific lesson files within the repository's phased structure.

### Stage 1: Document Ingestion and Chunking

Raw documents undergo **sentence-aware chunking** to preserve context while respecting embedding model constraints. The repository emphasizes chunk size as a critical hyper-parameter—too small loses semantic context, too large dilutes relevance.

According to [`phases/19-capstone-projects/64-chunking-strategies-advanced/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/64-chunking-strategies-advanced/docs/en.md), the recommended approach splits text at sentence boundaries using regex patterns:

```python
import re

def chunk_text(text: str, max_len: int = 500) -> list[str]:
    """Split `text` into chunks no longer than `max_len` characters,
    respecting sentence boundaries."""
    sentences = re.split(r'(?<=[.!?])\s+', text)
    chunks, cur = [], ''
    for s in sentences:
        if len(cur) + len(s) + 1 <= max_len:
            cur = f'{cur} {s}'.strip()
        else:
            chunks.append(cur)
            cur = s
    if cur:
        chunks.append(cur)
    return chunks

```

### Stage 2: Embedding Generation

Each chunk converts to a dense vector representation using encoder models. The repository's [`site/data.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/data.js) defines embeddings as "dense vector representations" and references implementations in [`phases/11-llm-engineering/04-embeddings/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/04-embeddings/docs/en.md). While the curriculum includes pretrained sentence-transformers, the from-scratch approach uses NumPy-based encoders for pedagogical clarity.

### Stage 3: Vector Store Construction

The **vector database** implementation stores embeddings in a simple in-memory NumPy array rather than external services. As defined in [`site/data.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/data.js), retrieval relies on **approximate nearest-neighbor search** using brute-force cosine similarity.

This approach appears in [`phases/11-llm-engineering/06-rag-fundamentals/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/06-rag-fundamentals/docs/en.md), which establishes the foundational pattern of dense vector storage and similarity-based retrieval.

### Stage 4: Dense Retrieval and Hybrid Search

Given a user query, the system embeds the query using the same encoder, then computes cosine similarity against the vector store to fetch **top-k chunks**. The repository extends this basic retrieval with **hybrid search** (BM25 + dense vectors) documented in [`phases/19-capstone-projects/65-hybrid-retrieval-bm25-dense/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/65-hybrid-retrieval-bm25-dense/docs/en.md).

### Stage 5: Reranking with Cross-Encoders

Retrieved chunks undergo **cross-encoder reranking** to improve precision. This stage, detailed in [`phases/19-capstone-projects/66-reranker-cross-encoder/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/66-reranker-cross-encoder/docs/en.md), uses a more expensive model to re-score the initial retrieval results, significantly improving answer relevance before generation.

### Stage 6: Prompt Construction and LLM Generation

Selected chunks concatenate into a structured prompt with citations, then feed into a language model. The "RAG pipeline" entry in [`site/data.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/data.js) links to [`phases/19-capstone-projects/69-end-to-end-rag-system/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/69-end-to-end-rag-system/docs/en.md), which demonstrates wiring retrieval output to generation input using a minimal LLaMA-style model built from scratch.

### Stage 7: Evaluation and Metrics

The pipeline closes with rigorous evaluation using **Recall@k**, **MRR** (Mean Reciprocal Rank), **nDCG**, and **RAGAS-style faithfulness** metrics. Lesson 68 ([`phases/19-capstone-projects/68-rag-eval-precision-recall/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/68-rag-eval-precision-recall/docs/en.md)) provides the evaluation harness for measuring retrieval quality and answer relevance.

## Building a Vector Store from Scratch

The repository's pedagogical approach avoids external vector databases in favor of NumPy-based implementations. Here is the complete retrieval system using only standard scientific Python libraries:

### Cosine Similarity Retrieval Engine

```python
import numpy as np

# 1. Encode chunks (here we use a dummy encoder; replace with a real model)

def dummy_encoder(chunk: str) -> np.ndarray:
    rng = np.random.default_rng(abs(hash(chunk)) % (2**32))
    return rng.normal(size=768)          # 768-dim embedding

# 2. Build the store

def build_store(chunks: list[str]) -> tuple[np.ndarray, list[str]]:
    vectors = np.stack([dummy_encoder(c) for c in chunks])
    return vectors, chunks

# 3. Retrieve top-k most similar chunks for a query

def retrieve(query: str, vectors: np.ndarray, chunks: list[str], k: int = 5) -> list[str]:
    q_vec = dummy_encoder(query)
    # Cosine similarity (dot / norms)

    sims = vectors @ q_vec / (np.linalg.norm(vectors, axis=1) * np.linalg.norm(q_vec))
    top_idx = np.argsort(sims)[-k:][::-1]
    return [chunks[i] for i in top_idx]

```

The `build_store` function creates a `vectors.npy` matrix where row indices map to chunk indices, enabling O(n) memory footprint with O(n) query time—sufficient for educational datasets before scaling to FAISS or HNSW.

## End-to-End RAG Pipeline Implementation

The complete pipeline from raw documents to generated answers appears in [`phases/19-capstone-projects/69-end-to-end-rag-system/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/69-end-to-end-rag-system/code/main.py). This implementation wires together all seven stages:

```python
def rag_pipeline(query: str, raw_docs: list[str]) -> str:
    # 0️⃣ Chunk all documents

    all_chunks = []
    for doc in raw_docs:
        all_chunks.extend(chunk_text(doc))

    # 1️⃣ Build vector store

    vectors, chunks = build_store(all_chunks)

    # 2️⃣ Retrieve relevant pieces

    retrieved = retrieve(query, vectors, chunks, k=3)

    # 3️⃣ Assemble prompt (simple concatenation)

    prompt = f"Answer the question using the following excerpts:\n\n"
    for i, c in enumerate(retrieved, 1):
        prompt += f"[{i}] {c}\n"
    prompt += f"\nQuestion: {query}\nAnswer:"

    # 4️⃣ Generate (here we just echo the prompt; replace with LLM call)

    return prompt   # in practice feed `prompt` to a LLM (e.g., a toy transformer from the repo)

```

## Key Source Files in the Curriculum

| File Path | Purpose | Lesson Reference |
|-----------|---------|------------------|
| [`phases/11-llm-engineering/06-rag-fundamentals/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/06-rag-fundamentals/docs/en.md) | Core RAG concepts and retrieval patterns | Lesson 06 |
| [`phases/11-llm-engineering/04-embeddings/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/04-embeddings/docs/en.md) | Dense vector generation theory | Lesson 04 |
| [`phases/19-capstone-projects/64-chunking-strategies-advanced/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/64-chunking-strategies-advanced/docs/en.md) | Advanced text segmentation strategies | Lesson 64 |
| [`phases/19-capstone-projects/65-hybrid-retrieval-bm25-dense/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/65-hybrid-retrieval-bm25-dense/docs/en.md) | Lexical + semantic retrieval fusion | Lesson 65 |
| [`phases/19-capstone-projects/66-reranker-cross-encoder/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/66-reranker-cross-encoder/docs/en.md) | Cross-encoder reranking implementation | Lesson 66 |
| [`phases/19-capstone-projects/68-rag-eval-precision-recall/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/68-rag-eval-precision-recall/docs/en.md) | Evaluation metrics and benchmarking | Lesson 68 |
| [`phases/19-capstone-projects/69-end-to-end-rag-system/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/69-end-to-end-rag-system/code/main.py) | Reference implementation of complete pipeline | Lesson 69 |
| [`site/data.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/data.js) | Metadata definitions for vector databases and RAG patterns | Curriculum index |

## Summary

- **RAG combines retrieval and generation** by fetching relevant document chunks before prompting a language model, reducing hallucinations and grounding answers in source material.
- **Vector stores can be built from scratch** using NumPy arrays and cosine similarity, as demonstrated in the repository's Lesson 06 and [`site/data.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/data.js) definitions.
- **Chunking strategy critically impacts retrieval quality**—the curriculum provides sentence-aware splitting functions in Lesson 64 to balance context preservation with embedding constraints.
- **Hybrid retrieval and reranking** significantly improve precision; Lessons 65 and 66 teach BM25+dense fusion and cross-encoder scoring before generation.
- **Evaluation requires specific metrics** like Recall@k and faithfulness scores, implemented in Lesson 68's evaluation harness.

## Frequently Asked Questions

### What dependencies are required to implement RAG from scratch according to this curriculum?

The `rohitg00/ai-engineering-from-scratch` repository restricts implementations to five allowed libraries: `numpy` for vector operations, `torch` for neural network components, `h5py` for storage serialization, `zstandard` for compression, and `safetensors` for model weight handling. No external vector database services like Pinecone or Weaviate are required for the foundational implementation.

### How does the repository handle document chunking for RAG?

The curriculum implements **sentence-boundary chunking** using regex split patterns that respect punctuation while enforcing maximum character lengths. Located in [`phases/19-capstone-projects/64-chunking-strategies-advanced/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/64-chunking-strategies-advanced/docs/en.md), this approach prevents mid-sentence cuts that degrade retrieval quality, using a sliding window approach where chunks overlap to preserve context across boundaries.

### Can this from-scratch implementation scale to production workloads?

The NumPy-based vector store provides O(n) query complexity suitable for educational datasets and prototyping. The repository explicitly designs components to be swappable—once you understand the brute-force cosine similarity implementation in [`phases/11-llm-engineering/06-rag-fundamentals/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/06-rag-fundamentals/docs/en.md), you can replace the `build_store` and `retrieve` functions with FAISS or HNSW indices while maintaining the same chunking, reranking, and evaluation pipeline established in the capstone lessons.

### What evaluation metrics does the curriculum use for RAG systems?

Lesson 68 ([`phases/19-capstone-projects/68-rag-eval-precision-recall/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/68-rag-eval-precision-recall/docs/en.md)) implements **Recall@k**, **MRR** (Mean Reciprocal Rank), and **nDCG** for retrieval evaluation, plus **faithfulness** and **answer relevance** scores for generation quality. These metrics measure both whether the system retrieves the correct chunks and whether the final generated answer accurately reflects the retrieved content without hallucination.