Why QMD Requires Separate Embedding Generation for Vector Search
QMD separates embedding generation from vector search because computing embeddings is computationally expensive and must be performed once per document, while vector search is a lightweight SQL lookup that runs repeatedly against pre-computed indexes.
QMD (Query Markdown) is an open-source semantic search engine that stores document embeddings in SQLite. Unlike systems that generate embeddings on-the-fly during search, QMD explicitly splits the workflow into two distinct phases. This architectural decision is driven by performance constraints, cost optimization, and the technical limitations of SQLite-vec virtual tables.
The Two-Phase Architecture
QMD’s vector search implementation is deliberately bifurcated into embedding generation and vector retrieval. Understanding this separation requires examining how each phase interacts with the underlying storage layer.
Phase 1: Embedding Generation
During indexing, QMD processes raw document chunks through an embedding model. The default configuration uses embeddinggemma, which produces 768-dimensional float arrays. This operation occurs in src/llm.ts via the embed and embedBatch functions (lines 97-124), which interface with the underlying LLM runtime.
The generated vectors are then persisted to two locations in src/store.ts:
- The
content_vectorstable stores metadata about the embedding - The
vectors_vecvirtual table (managed by SQLite-vec) stores the actual vector data for indexed retrieval
The insertEmbedding function (lines 82-96) handles this dual-write operation:
// From src/store.ts - insertEmbedding (lines 82-96)
store.insertEmbedding(
hash, // Document hash
0, // Sequence number (chunk index)
0, // Position within chunk
new Float32Array(result.embedding), // The 768-dim vector
"embeddinggemma", // Model identifier
new Date().toISOString() // Timestamp
);
Phase 2: Vector Search
When a user submits a query, QMD does not re-embed the entire document corpus. Instead, it generates a single embedding for the query string, then performs a nearest-neighbor lookup against the pre-computed vectors_vec table.
The searchVec function in src/store.ts (lines 2056-2140) implements this logic:
- Query Embedding: The function calls
getEmbedding(which usesllm.embed) to convert the search string into a 768-dimensional vector - Vector Retrieval: It executes a SQLite-vec query using cosine distance:
SELECT ... FROM vectors_vec WHERE embedding MATCH ? - Document Join: Because SQLite-vec virtual tables cannot be combined with standard SQL joins in the same statement, the implementation performs a separate join against document metadata using the returned
hash_seqvalues
// Conceptual flow from src/store.ts searchVec
const queryVector = await getEmbedding(query, model); // Step 1: Single embedding
const vectorMatches = await db.all(
`SELECT rowid, distance FROM vectors_vec WHERE embedding MATCH ? LIMIT ?`,
[JSON.stringify(queryVector), limit]
); // Step 2: Pure SQL vector search
Why Separation Is Mandatory
The architectural split is not merely stylistic—it addresses three technical constraints inherent to semantic search systems.
Computational Cost and Reuse
Embedding generation is expensive. Each call to llm.embed invokes a GPU or CPU-intensive neural network inference. For a corpus of 10,000 documents, on-the-fly embedding would require 10,000 inference operations per search query.
By separating the phases, QMD performs inference once during indexing. The resulting vectors are stored in content_vectors and reused for every subsequent search. This transforms search complexity from O(N) × inference_cost to O(1) × inference_cost (for the query) + O(log N) (for the SQL index lookup).
SQLite-vec Technical Limitations
The vectors_vec table is a virtual table provided by the SQLite-vec extension. As noted in the source code comments within searchVec, SQLite-vec virtual tables cannot participate in standard SQL JOIN operations within the same query statement.
This constraint forces a two-step retrieval process:
- Query the virtual table to get matching vector rowids and distances
- Join those results with the
contentandcontent_vectorstables to retrieve document metadata
This technical boundary reinforces the architectural separation: vector storage/retrieval must remain distinct from document metadata operations.
Batch Processing and Concurrency
QMD optimizes embedding generation through context pooling implemented in src/llm.ts (lines 19-34). The ensureEmbedContexts function creates multiple embedding contexts that process document chunks concurrently.
This batching is only possible because embedding is a distinct phase. During indexing, QMD can:
- Collect all documents needing embeddings via
getHashesForEmbedding - Process them in parallel batches using
embedBatch - Store results transactionally
If embedding were coupled to search, this optimization would be impossible, and every search would incur the latency of sequential document processing.
Implementation Example
The following TypeScript example demonstrates the complete workflow, from initial embedding generation to repeated vector searches:
import { createStore } from "./store";
import { withLLMSession } from "./llm";
// Initialize the store (creates SQLite tables including vectors_vec)
const store = createStore();
// Step 1: Ensure vector table exists with correct dimensions (768 for embeddinggemma)
store.ensureVecTable(768);
// Step 2: Generate embeddings for documents that lack them
await withLLMSession(async session => {
const toEmbed = store.getHashesForEmbedding(); // Documents needing vectors
for (const { hash, body, path } of toEmbed) {
const result = await session.embed(body, { model: "embeddinggemma" });
if (result?.embedding) {
// Insert into both content_vectors and vectors_vec tables
store.insertEmbedding(
hash, 0, 0,
new Float32Array(result.embedding),
"embeddinggemma",
new Date().toISOString()
);
}
}
});
// Step 3: Perform vector searches (can run thousands of times without LLM calls)
const query = "How to configure QMD collections?";
const results = store.searchVec(query, "embeddinggemma", 10);
// Results contain filepath, score, and metadata
console.log(results.map(r => `${r.filepath}: ${r.score.toFixed(3)}`));
In this workflow, the expensive embedding generation (Step 2) occurs once during indexing, while the lightweight vector search (Step 3) executes repeatedly via pure SQL queries against the vectors_vec virtual table.
Summary
- Embedding generation is computationally expensive: QMD uses
llm.embedandllm.embedBatch(insrc/llm.ts) to create 768-dimensional vectors via neural network inference, which requires significant GPU/CPU resources. - Pre-computed embeddings enable scalable search: By storing vectors in the
vectors_vecSQLite-vec virtual table andcontent_vectorstable (managed insrc/store.ts), QMD avoids re-embedding documents on every query. - Technical constraints enforce separation: SQLite-vec virtual tables cannot participate in standard SQL JOINs, requiring a two-phase retrieval process (vector lookup followed by document join) implemented in
searchVec(lines 2056-2140 ofsrc/store.ts). - Batch processing optimizes indexing: The separate phase allows
ensureEmbedContexts(insrc/llm.ts) to parallelize embedding generation across multiple contexts, dramatically speeding up initial index builds.
Frequently Asked Questions
Why can't QMD generate embeddings during the search query?
Generating embeddings during search would require invoking the LLM inference engine for every document in the corpus, resulting in O(N) complexity with high per-operation latency. By separating the phases, QMD performs inference once during indexing (using insertEmbedding in src/store.ts) and stores results in the vectors_vec table, allowing subsequent searches to execute as pure SQL queries without model runtime overhead.
What embedding model does QMD use by default?
QMD defaults to embeddinggemma, which produces 768-dimensional float vectors. This is configured in the LLM layer (src/llm.ts) and referenced throughout the store implementation (src/store.ts) when calling session.embed() or store.searchVec(). The model identifier is stored alongside vectors in the content_vectors table to ensure retrieval uses compatible embeddings.
How does QMD handle batch embedding for large document collections?
QMD implements context pooling via ensureEmbedContexts in src/llm.ts (lines 19-34), which creates multiple concurrent embedding contexts. During indexing, getHashesForEmbedding identifies documents lacking vectors, and the system processes them in parallel batches using embedBatch. This separation of concerns allows the expensive embedding phase to scale across CPU/GPU resources while keeping the search phase lightweight and SQLite-bound.
Why does QMD use a separate vectors_vec virtual table instead of storing vectors in the main content table?
QMD uses the vectors_vec virtual table provided by the SQLite-vec extension to enable specialized vector indexing and cosine-similarity search. As noted in the searchVec implementation (src/store.ts, lines 2056-2140), SQLite-vec virtual tables cannot be combined with standard SQL JOINs in a single statement. This technical constraint necessitates a two-step query process: first retrieving nearest neighbors from vectors_vec, then joining with the content table via hash_seq to fetch document metadata.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →