# Implementing Embedding-Based Similarity Search for RAG: From Mock Vectors to Production

> Learn to implement embedding-based similarity search for RAG. Map text to vectors and scan for relevant documents to enhance LLM context. Start from scratch for production readiness.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-26

---

**Embedding-based similarity search for RAG uses deterministic hashing to map text into normalized 96-dimensional vectors, then performs brute-force cosine similarity scans to retrieve the top-k most relevant documents for LLM context augmentation.**

The `rohitg00/ai-engineering-from-scratch` repository demonstrates how to build a complete **embedding-based similarity search** pipeline from first principles without external model dependencies. This implementation provides a self-contained dense retriever that can run entirely offline using hash-based mock embeddings, while maintaining a thin API surface that allows seamless swapping for production vector databases like FAISS, HNSW, or Milvus when scaling beyond experimental datasets.

## Deterministic Mock Embeddings for Offline Development

At the core of this RAG implementation sits a **deterministic mock embedding** function that generates reproducible 96-dimensional vectors from any input string. As found in [`phases/19-capstone-projects/69-end-to-end-rag-system/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/69-end-to-end-rag-system/code/main.py) at line 66, the `mock_embed(text, dim=96)` function uses SHA-256 hashing mixed with a linear congruential generator to produce normalized vectors without calling external APIs.

This approach serves two critical purposes in the learning pipeline. First, it makes the entire RAG loop deterministic and testable—running the same query always yields the same embedding vector. Second, it eliminates network dependencies and API costs, allowing students to experiment with vector search mechanics using pure Python standard library code.

```python
def mock_embed(text: str, dim: int = 96) -> list[float]:
    """
    Produce a reproducible 96-dimensional vector from any string.
    The implementation hashes each token, mixes the hash with a
    simple linear congruential generator and normalises the result.
    """
    import hashlib, math
    rng = int(hashlib.sha256(text.encode()).hexdigest(), 16)
    vec = [(rng >> i) & 0xFF for i in range(dim)]
    norm = math.sqrt(sum(v * v for v in vec)) or 1.0
    return [v / norm for v in vec]

```

## Building an In-Memory Vector Store with Cosine Similarity

The repository implements a **flat index** vector store that stores `(doc_id, vector)` tuples in a Python list and performs brute-force similarity computations for each query. According to [`phases/11-llm-engineering/06-rag/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/11-llm-engineering/06-rag/docs/en.md), this design mirrors a FAISS "flat" index and remains performant for up to approximately 100,000 vectors before latency degrades.

The `VectorStore` class exposes a minimal API surface with two primary methods: `add(id, text)` to embed and store documents, and `search(query, k)` to retrieve the top-k matches. The similarity calculation leverages the fact that `mock_embed` outputs L2-normalized vectors, allowing the implementation to use simple dot products as a proxy for cosine similarity.

```python
from typing import List, Tuple
import math

class VectorStore:
    """Flat index storing (doc_id, vector) pairs."""
    def __init__(self) -> None:
        self._data: List[Tuple[str, List[float]]] = []

    def add(self, doc_id: str, text: str) -> None:
        """Embed `text` and store it."""
        self._data.append((doc_id, mock_embed(text)))

    @staticmethod
    def _cosine(a: List[float], b: List[float]) -> float:
        dot = sum(x * y for x, y in zip(a, b))
        # vectors are already L2-normalised by mock_embed

        return dot

    def search(self, query: str, top_k: int = 5) -> List[Tuple[str, float]]:
        """Return the `top_k` doc_ids sorted by descending cosine similarity."""
        q_vec = mock_embed(query)
        sims = [(doc_id, self._cosine(q_vec, vec)) for doc_id, vec in self._data]
        sims.sort(key=lambda x: x[1], reverse=True)
        return sims[:top_k]

```

## Cosine Similarity vs. L2 Distance in Vector Retrieval

The foundation lesson in [`phases/01-math-foundations/14-norms-and-distances/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/01-math-foundations/14-norms-and-distances/docs/en.md) explains the mathematical distinction between **cosine similarity** and **L2 distance** that drives retrieval behavior. Cosine similarity measures angular alignment between vectors—ideal for capturing semantic "direction" regardless of vector magnitude—while L2 distance measures raw Euclidean proximity in the embedding space.

In the provided implementation, vectors arrive pre-normalized from `mock_embed`, meaning the dot product calculation in `_cosine()` mathematically equals the cosine similarity score. This optimization eliminates costly square root operations during search time, reducing the brute-force scan to simple arithmetic operations over the 96-dimensional vectors.

## Hybrid Retrieval: Combining Dense and Sparse Signals

Advanced capstone projects extend the basic embedding search into **hybrid retrieval** pipelines that combine dense vector similarity with sparse lexical matching. As documented in [`phases/19-capstone-projects/65-hybrid-retrieval-bm25-dense/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/65-hybrid-retrieval-bm25-dense/docs/en.md), this approach fuses the semantic understanding of embeddings with the keyword precision of BM25 algorithms.

The hybrid architecture reuses the same deterministic embedding function and vector store class, adding a parallel BM25 index built over the raw text. At query time, the system retrieves candidates from both stores and applies score fusion—typically reciprocal rank fusion or linear weighting—to produce a final ranked list that captures both semantic and lexical relevance.

## Production Migration Path: Swapping the Embedder

Because the `VectorStore` class depends only on the embedding interface—not the implementation—students can migrate to production systems by replacing `mock_embed` with real models. The repository suggests dropping in `sentence-transformers` models, OpenAI's `text-embedding-3-small`, or any other encoder that returns normalized float vectors. Since the store expects only `add(id, text)` and `search(query, k)` methods, it serves as a drop-in replacement for more sophisticated backends like Pinecone, Weaviate, or pgvector.

To migrate, simply override the embedding call inside `add()` and `search()` while preserving the cosine similarity logic, or replace the entire `_data` structure with a FAISS index that implements its own approximate nearest neighbor search.

```python

# Example usage with the mock embedder

store = VectorStore()
store.add("doc1", "The budget abort threshold is 0.5")
store.add("doc2", "How to handle multipart upload failures")
store.add("doc3", "Policy for retry budgets in S3")

results = store.search("What happens when the upload budget is exhausted?")
for doc_id, score in results:
    print(f"{doc_id}: {score:.3f}")

```

## Summary

- **Deterministic mock embeddings** enable offline RAG development using SHA-256 hashing to generate reproducible 96-dimensional vectors without external API calls.
- **Brute-force similarity search** implemented in `VectorStore` provides O(N) retrieval suitable for prototyping with fewer than 100,000 documents.
- **Cosine similarity** captures angular alignment between query and document vectors; the implementation optimizes this by using dot products on pre-normalized vectors.
- **Hybrid retrieval** combines dense embedding search with BM25 sparse signals in later capstone projects for robust production-grade retrieval.
- **Thin API abstraction** allows seamless replacement of the mock embedder with production models like sentence-transformers or OpenAI embeddings while preserving the search interface.

## Frequently Asked Questions

### Why use a mock embedding function instead of a real model?

The **mock embedder** eliminates external dependencies and API costs, allowing the RAG pipeline to run entirely offline. Since it produces deterministic outputs based on SHA-256 hashing, it enables reproducible debugging and testing of the similarity search logic before integrating expensive embedding models.

### When should I replace the brute-force search with FAISS?

Replace the flat index when your corpus exceeds roughly 100,000 vectors or when query latency exceeds acceptable thresholds. The brute-force approach in `VectorStore` requires scanning the entire dataset for each query—O(N) complexity—whereas FAISS implements approximate nearest neighbor search with sub-linear complexity, enabling millisecond-scale retrieval over millions of vectors.

### How does cosine similarity differ from dot product in this implementation?

Mathematically, cosine similarity equals the dot product divided by the product of the vector magnitudes. Because `mock_embed` explicitly L2-normalizes all output vectors to unit length (norm = 1.0), the division step becomes unnecessary, and the raw dot product equals the cosine similarity score. This optimization reduces computational overhead during the brute-force scan.

### Can I use this vector store with OpenAI or sentence-transformer embeddings?

Yes. The `VectorStore` class architecture is backend-agnostic. Replace the `mock_embed(text)` calls in both the `add()` and `search()` methods with calls to `openai.embeddings.create()` or `sentence_transformers.encode()`, ensuring the resulting vectors are normalized before storage. The cosine similarity calculation remains valid as long as vectors share the same dimensionality and normalization scheme.