Implementing Embedding-Based Similarity Search for RAG: From Mock Vectors to Production

Embedding-based similarity search for RAG uses deterministic hashing to map text into normalized 96-dimensional vectors, then performs brute-force cosine similarity scans to retrieve the top-k most relevant documents for LLM context augmentation.

The rohitg00/ai-engineering-from-scratch repository demonstrates how to build a complete embedding-based similarity search pipeline from first principles without external model dependencies. This implementation provides a self-contained dense retriever that can run entirely offline using hash-based mock embeddings, while maintaining a thin API surface that allows seamless swapping for production vector databases like FAISS, HNSW, or Milvus when scaling beyond experimental datasets.

Deterministic Mock Embeddings for Offline Development

At the core of this RAG implementation sits a deterministic mock embedding function that generates reproducible 96-dimensional vectors from any input string. As found in phases/19-capstone-projects/69-end-to-end-rag-system/code/main.py at line 66, the mock_embed(text, dim=96) function uses SHA-256 hashing mixed with a linear congruential generator to produce normalized vectors without calling external APIs.

This approach serves two critical purposes in the learning pipeline. First, it makes the entire RAG loop deterministic and testable—running the same query always yields the same embedding vector. Second, it eliminates network dependencies and API costs, allowing students to experiment with vector search mechanics using pure Python standard library code.

def mock_embed(text: str, dim: int = 96) -> list[float]:
    """
    Produce a reproducible 96-dimensional vector from any string.
    The implementation hashes each token, mixes the hash with a
    simple linear congruential generator and normalises the result.
    """
    import hashlib, math
    rng = int(hashlib.sha256(text.encode()).hexdigest(), 16)
    vec = [(rng >> i) & 0xFF for i in range(dim)]
    norm = math.sqrt(sum(v * v for v in vec)) or 1.0
    return [v / norm for v in vec]

Building an In-Memory Vector Store with Cosine Similarity

The repository implements a flat index vector store that stores (doc_id, vector) tuples in a Python list and performs brute-force similarity computations for each query. According to phases/11-llm-engineering/06-rag/docs/en.md, this design mirrors a FAISS "flat" index and remains performant for up to approximately 100,000 vectors before latency degrades.

The VectorStore class exposes a minimal API surface with two primary methods: add(id, text) to embed and store documents, and search(query, k) to retrieve the top-k matches. The similarity calculation leverages the fact that mock_embed outputs L2-normalized vectors, allowing the implementation to use simple dot products as a proxy for cosine similarity.

from typing import List, Tuple
import math

class VectorStore:
    """Flat index storing (doc_id, vector) pairs."""
    def __init__(self) -> None:
        self._data: List[Tuple[str, List[float]]] = []

    def add(self, doc_id: str, text: str) -> None:
        """Embed `text` and store it."""
        self._data.append((doc_id, mock_embed(text)))

    @staticmethod
    def _cosine(a: List[float], b: List[float]) -> float:
        dot = sum(x * y for x, y in zip(a, b))
        # vectors are already L2-normalised by mock_embed

        return dot

    def search(self, query: str, top_k: int = 5) -> List[Tuple[str, float]]:
        """Return the `top_k` doc_ids sorted by descending cosine similarity."""
        q_vec = mock_embed(query)
        sims = [(doc_id, self._cosine(q_vec, vec)) for doc_id, vec in self._data]
        sims.sort(key=lambda x: x[1], reverse=True)
        return sims[:top_k]

Cosine Similarity vs. L2 Distance in Vector Retrieval

The foundation lesson in phases/01-math-foundations/14-norms-and-distances/docs/en.md explains the mathematical distinction between cosine similarity and L2 distance that drives retrieval behavior. Cosine similarity measures angular alignment between vectors—ideal for capturing semantic "direction" regardless of vector magnitude—while L2 distance measures raw Euclidean proximity in the embedding space.

In the provided implementation, vectors arrive pre-normalized from mock_embed, meaning the dot product calculation in _cosine() mathematically equals the cosine similarity score. This optimization eliminates costly square root operations during search time, reducing the brute-force scan to simple arithmetic operations over the 96-dimensional vectors.

Hybrid Retrieval: Combining Dense and Sparse Signals

Advanced capstone projects extend the basic embedding search into hybrid retrieval pipelines that combine dense vector similarity with sparse lexical matching. As documented in phases/19-capstone-projects/65-hybrid-retrieval-bm25-dense/docs/en.md, this approach fuses the semantic understanding of embeddings with the keyword precision of BM25 algorithms.

The hybrid architecture reuses the same deterministic embedding function and vector store class, adding a parallel BM25 index built over the raw text. At query time, the system retrieves candidates from both stores and applies score fusion—typically reciprocal rank fusion or linear weighting—to produce a final ranked list that captures both semantic and lexical relevance.

Production Migration Path: Swapping the Embedder

Because the VectorStore class depends only on the embedding interface—not the implementation—students can migrate to production systems by replacing mock_embed with real models. The repository suggests dropping in sentence-transformers models, OpenAI's text-embedding-3-small, or any other encoder that returns normalized float vectors. Since the store expects only add(id, text) and search(query, k) methods, it serves as a drop-in replacement for more sophisticated backends like Pinecone, Weaviate, or pgvector.

To migrate, simply override the embedding call inside add() and search() while preserving the cosine similarity logic, or replace the entire _data structure with a FAISS index that implements its own approximate nearest neighbor search.


# Example usage with the mock embedder

store = VectorStore()
store.add("doc1", "The budget abort threshold is 0.5")
store.add("doc2", "How to handle multipart upload failures")
store.add("doc3", "Policy for retry budgets in S3")

results = store.search("What happens when the upload budget is exhausted?")
for doc_id, score in results:
    print(f"{doc_id}: {score:.3f}")

Summary

  • Deterministic mock embeddings enable offline RAG development using SHA-256 hashing to generate reproducible 96-dimensional vectors without external API calls.
  • Brute-force similarity search implemented in VectorStore provides O(N) retrieval suitable for prototyping with fewer than 100,000 documents.
  • Cosine similarity captures angular alignment between query and document vectors; the implementation optimizes this by using dot products on pre-normalized vectors.
  • Hybrid retrieval combines dense embedding search with BM25 sparse signals in later capstone projects for robust production-grade retrieval.
  • Thin API abstraction allows seamless replacement of the mock embedder with production models like sentence-transformers or OpenAI embeddings while preserving the search interface.

Frequently Asked Questions

Why use a mock embedding function instead of a real model?

The mock embedder eliminates external dependencies and API costs, allowing the RAG pipeline to run entirely offline. Since it produces deterministic outputs based on SHA-256 hashing, it enables reproducible debugging and testing of the similarity search logic before integrating expensive embedding models.

When should I replace the brute-force search with FAISS?

Replace the flat index when your corpus exceeds roughly 100,000 vectors or when query latency exceeds acceptable thresholds. The brute-force approach in VectorStore requires scanning the entire dataset for each query—O(N) complexity—whereas FAISS implements approximate nearest neighbor search with sub-linear complexity, enabling millisecond-scale retrieval over millions of vectors.

How does cosine similarity differ from dot product in this implementation?

Mathematically, cosine similarity equals the dot product divided by the product of the vector magnitudes. Because mock_embed explicitly L2-normalizes all output vectors to unit length (norm = 1.0), the division step becomes unnecessary, and the raw dot product equals the cosine similarity score. This optimization reduces computational overhead during the brute-force scan.

Can I use this vector store with OpenAI or sentence-transformer embeddings?

Yes. The VectorStore class architecture is backend-agnostic. Replace the mock_embed(text) calls in both the add() and search() methods with calls to openai.embeddings.create() or sentence_transformers.encode(), ensuring the resulting vectors are normalized before storage. The cosine similarity calculation remains valid as long as vectors share the same dimensionality and normalization scheme.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →