# Why QMD Requires Separate Embedding Generation for Vector Search

> Discover why QMD needs separate embedding generation for vector search. Understand the cost-effective approach to processing expensive embeddings once and performing lightweight SQL lookups repeatedly.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: internals
- Published: 2026-02-16

---

**QMD separates embedding generation from vector search because computing embeddings is computationally expensive and must be performed once per document, while vector search is a lightweight SQL lookup that runs repeatedly against pre-computed indexes.**

QMD (Query Markdown) is an open-source semantic search engine that stores document embeddings in SQLite. Unlike systems that generate embeddings on-the-fly during search, QMD explicitly splits the workflow into two distinct phases. This architectural decision is driven by performance constraints, cost optimization, and the technical limitations of SQLite-vec virtual tables.

## The Two-Phase Architecture

QMD’s vector search implementation is deliberately bifurcated into **embedding generation** and **vector retrieval**. Understanding this separation requires examining how each phase interacts with the underlying storage layer.

### Phase 1: Embedding Generation

During indexing, QMD processes raw document chunks through an embedding model. The default configuration uses `embeddinggemma`, which produces 768-dimensional float arrays. This operation occurs in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) via the `embed` and `embedBatch` functions (lines 97-124), which interface with the underlying LLM runtime.

The generated vectors are then persisted to two locations in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts):
- The `content_vectors` table stores metadata about the embedding
- The `vectors_vec` virtual table (managed by SQLite-vec) stores the actual vector data for indexed retrieval

The `insertEmbedding` function (lines 82-96) handles this dual-write operation:

```typescript
// From src/store.ts - insertEmbedding (lines 82-96)
store.insertEmbedding(
  hash,           // Document hash
  0,              // Sequence number (chunk index)
  0,              // Position within chunk
  new Float32Array(result.embedding),  // The 768-dim vector
  "embeddinggemma",                    // Model identifier
  new Date().toISOString()             // Timestamp
);

```

### Phase 2: Vector Search

When a user submits a query, QMD does **not** re-embed the entire document corpus. Instead, it generates a single embedding for the query string, then performs a nearest-neighbor lookup against the pre-computed `vectors_vec` table.

The `searchVec` function in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) (lines 2056-2140) implements this logic:

1. **Query Embedding**: The function calls `getEmbedding` (which uses `llm.embed`) to convert the search string into a 768-dimensional vector
2. **Vector Retrieval**: It executes a SQLite-vec query using cosine distance: `SELECT ... FROM vectors_vec WHERE embedding MATCH ?`
3. **Document Join**: Because SQLite-vec virtual tables cannot be combined with standard SQL joins in the same statement, the implementation performs a separate join against document metadata using the returned `hash_seq` values

```typescript
// Conceptual flow from src/store.ts searchVec
const queryVector = await getEmbedding(query, model);  // Step 1: Single embedding
const vectorMatches = await db.all(
  `SELECT rowid, distance FROM vectors_vec WHERE embedding MATCH ? LIMIT ?`,
  [JSON.stringify(queryVector), limit]
);  // Step 2: Pure SQL vector search

```

## Why Separation Is Mandatory

The architectural split is not merely stylistic—it addresses three technical constraints inherent to semantic search systems.

### Computational Cost and Reuse

Embedding generation is **expensive**. Each call to `llm.embed` invokes a GPU or CPU-intensive neural network inference. For a corpus of 10,000 documents, on-the-fly embedding would require 10,000 inference operations per search query.

By separating the phases, QMD performs inference **once** during indexing. The resulting vectors are stored in `content_vectors` and reused for every subsequent search. This transforms search complexity from O(N) × inference_cost to O(1) × inference_cost (for the query) + O(log N) (for the SQL index lookup).

### SQLite-vec Technical Limitations

The `vectors_vec` table is a **virtual table** provided by the SQLite-vec extension. As noted in the source code comments within `searchVec`, SQLite-vec virtual tables cannot participate in standard SQL JOIN operations within the same query statement.

This constraint forces a two-step retrieval process:
1. Query the virtual table to get matching vector rowids and distances
2. Join those results with the `content` and `content_vectors` tables to retrieve document metadata

This technical boundary reinforces the architectural separation: vector storage/retrieval must remain distinct from document metadata operations.

### Batch Processing and Concurrency

QMD optimizes embedding generation through **context pooling** implemented in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) (lines 19-34). The `ensureEmbedContexts` function creates multiple embedding contexts that process document chunks concurrently.

This batching is only possible because embedding is a distinct phase. During indexing, QMD can:
- Collect all documents needing embeddings via `getHashesForEmbedding`
- Process them in parallel batches using `embedBatch`
- Store results transactionally

If embedding were coupled to search, this optimization would be impossible, and every search would incur the latency of sequential document processing.

## Implementation Example

The following TypeScript example demonstrates the complete workflow, from initial embedding generation to repeated vector searches:

```typescript
import { createStore } from "./store";
import { withLLMSession } from "./llm";

// Initialize the store (creates SQLite tables including vectors_vec)
const store = createStore();

// Step 1: Ensure vector table exists with correct dimensions (768 for embeddinggemma)
store.ensureVecTable(768);

// Step 2: Generate embeddings for documents that lack them
await withLLMSession(async session => {
  const toEmbed = store.getHashesForEmbedding();  // Documents needing vectors
  
  for (const { hash, body, path } of toEmbed) {
    const result = await session.embed(body, { model: "embeddinggemma" });
    if (result?.embedding) {
      // Insert into both content_vectors and vectors_vec tables
      store.insertEmbedding(
        hash, 0, 0, 
        new Float32Array(result.embedding), 
        "embeddinggemma",
        new Date().toISOString()
      );
    }
  }
});

// Step 3: Perform vector searches (can run thousands of times without LLM calls)
const query = "How to configure QMD collections?";
const results = store.searchVec(query, "embeddinggemma", 10);

// Results contain filepath, score, and metadata
console.log(results.map(r => `${r.filepath}: ${r.score.toFixed(3)}`));

```

In this workflow, the expensive embedding generation (Step 2) occurs once during indexing, while the lightweight vector search (Step 3) executes repeatedly via pure SQL queries against the `vectors_vec` virtual table.

## Summary

- **Embedding generation is computationally expensive**: QMD uses `llm.embed` and `llm.embedBatch` (in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts)) to create 768-dimensional vectors via neural network inference, which requires significant GPU/CPU resources.
- **Pre-computed embeddings enable scalable search**: By storing vectors in the `vectors_vec` SQLite-vec virtual table and `content_vectors` table (managed in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts)), QMD avoids re-embedding documents on every query.
- **Technical constraints enforce separation**: SQLite-vec virtual tables cannot participate in standard SQL JOINs, requiring a two-phase retrieval process (vector lookup followed by document join) implemented in `searchVec` (lines 2056-2140 of [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts)).
- **Batch processing optimizes indexing**: The separate phase allows `ensureEmbedContexts` (in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts)) to parallelize embedding generation across multiple contexts, dramatically speeding up initial index builds.

## Frequently Asked Questions

### Why can't QMD generate embeddings during the search query?

Generating embeddings during search would require invoking the LLM inference engine for every document in the corpus, resulting in O(N) complexity with high per-operation latency. By separating the phases, QMD performs inference once during indexing (using `insertEmbedding` in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts)) and stores results in the `vectors_vec` table, allowing subsequent searches to execute as pure SQL queries without model runtime overhead.

### What embedding model does QMD use by default?

QMD defaults to `embeddinggemma`, which produces 768-dimensional float vectors. This is configured in the LLM layer ([`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts)) and referenced throughout the store implementation ([`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts)) when calling `session.embed()` or `store.searchVec()`. The model identifier is stored alongside vectors in the `content_vectors` table to ensure retrieval uses compatible embeddings.

### How does QMD handle batch embedding for large document collections?

QMD implements context pooling via `ensureEmbedContexts` in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) (lines 19-34), which creates multiple concurrent embedding contexts. During indexing, `getHashesForEmbedding` identifies documents lacking vectors, and the system processes them in parallel batches using `embedBatch`. This separation of concerns allows the expensive embedding phase to scale across CPU/GPU resources while keeping the search phase lightweight and SQLite-bound.

### Why does QMD use a separate `vectors_vec` virtual table instead of storing vectors in the main content table?

QMD uses the `vectors_vec` virtual table provided by the SQLite-vec extension to enable specialized vector indexing and cosine-similarity search. As noted in the `searchVec` implementation ([`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts), lines 2056-2140), SQLite-vec virtual tables cannot be combined with standard SQL JOINs in a single statement. This technical constraint necessitates a two-step query process: first retrieving nearest neighbors from `vectors_vec`, then joining with the `content` table via `hash_seq` to fetch document metadata.