QMD Search Commands Explained: Differences Between search, vsearch, and query

The search command performs pure BM25 full-text retrieval, vsearch executes vector-only semantic search, and query runs a hybrid pipeline combining BM25, vector search, LLM query expansion, and re-ranking.

QMD (Query Markdown Database) is an open-source retrieval system by Tobias Lütke (tobi/qmd) that provides three distinct CLI commands for querying indexed documents. Each command wraps a different retrieval strategy implemented in src/qmd.ts and delegates to specific algorithms in src/store.ts. Understanding the architectural differences between these commands helps you select the right tool for keyword-heavy lookups, semantic similarity searches, or high-precision hybrid retrieval.

Overview of QMD's Three Search Strategies

The QMD CLI exposes three entry points that map directly to internal TypeScript functions:

Each strategy makes different trade-offs between speed, recall, and semantic understanding.

The search Command: Pure BM25 Full-Text Retrieval

The search command implements traditional keyword-based information retrieval using SQLite's FTS5 extension with BM25 ranking.

Implementation Details

In src/qmd.ts at line 1905, the search function invokes searchFTS from src/store.ts. This function scans the FTS5 inverted index, matches literal terms against document content, and returns results ranked by the BM25 algorithm. The implementation performs no vector embedding lookups, no query expansion, and no neural re-ranking.

Use search when you need fast, deterministic keyword matching without semantic ambiguity. This command excels for exact terminology, code identifiers, or when you know the precise words appearing in your documents.


# Fast keyword search – BM25 only

qmd search "authentication flow" -c notes

The vsearch command performs dense retrieval using vector embeddings and cosine similarity, bypassing traditional keyword indexing entirely.

Implementation Details

Defined at line 1960 in src/qmd.ts, the async function vectorSearch calls vectorSearchQuery in src/store.ts. This pipeline uses the SQLite vec extension to query a vector index populated by an embedding model (defaulting to embeddinggemma). It retrieves nearest neighbors based on cosine similarity between the query embedding and document chunk embeddings. The process includes no BM25 fallback, no LLM query expansion, and only optional reranking within the vector pipeline itself.

When to Use vsearch

Use vsearch when you need conceptual similarity rather than lexical matching. This command finds documents related to the meaning of your query even when they use different terminology or synonyms.


# Semantic vector search – no BM25, no expansion

qmd vsearch "how to login" -c docs --limit 5

The query Command: Hybrid Search with LLM Enhancement

The query command implements a sophisticated multi-stage retrieval pipeline that combines the strengths of keyword and vector search with large language model augmentation.

Implementation Details

Located at line 2021 in src/qmd.ts, the async function querySearch orchestrates hybridQuery from src/store.ts. The algorithm executes four distinct phases:

  1. BM25 Pass: Runs an initial fast full-text search to retrieve candidate documents
  2. Query Expansion: If the BM25 signal is weak, expands the original query using an LLM (Qwen3) and performs additional vector searches
  3. Re-ranking: Combines results from previous stages and applies a neural re-ranking model (Qwen3-reranker) to optimize result ordering
  4. Deduplication: Returns the best chunk per document to provide concise, non-redundant snippets

When to Use query

Use query when you require the highest possible result quality and can tolerate additional latency. This command maximizes recall and precision by leveraging multiple retrieval signals and neural re-ranking.


# Hybrid search with LLM expansion and reranking

qmd query "user authentication" -c all --json

Programmatic Usage Examples

You can invoke these retrieval strategies directly in Node.js or Bun applications without using the CLI. The src/store.ts module exports the underlying functions that power each command:

import { createStore, searchFTS, vectorSearchQuery, hybridQuery } from "./src/store.js";

const store = await createStore();
const db = store.db;

// 1️⃣ BM25 only (equivalent to CLI 'search')
const bm25 = searchFTS(db, "authentication flow", 10);
console.log("BM25 results:", bm25);

// 2️⃣ Vector only (equivalent to CLI 'vsearch')
const vec = await vectorSearchQuery(store, "how to login", { limit: 5 });
console.log("Vector results:", vec);

// 3️⃣ Hybrid (equivalent to CLI 'query')
const hybrid = await hybridQuery(store, "user authentication", { limit: 10 });
console.log("Hybrid results:", hybrid);

Summary

  • search executes pure BM25 full-text retrieval via searchFTS in src/store.ts, ideal for fast keyword matching without semantic processing.
  • vsearch performs vector-only semantic search via vectorSearchQuery, using the SQLite vec extension and cosine similarity for conceptual retrieval.
  • query runs a hybrid pipeline via hybridQuery that combines BM25, vector search, LLM query expansion (Qwen3), and neural re-ranking (Qwen3-reranker) for maximum accuracy.

Frequently Asked Questions

Which QMD command should I use for keyword-heavy searches?

Use the search command. It executes a pure BM25 retrieval through the FTS5 index without the overhead of embedding generation or neural re-ranking, making it the fastest option for exact terminology and code identifiers.

Does vsearch use BM25 as a fallback?

No. The vsearch command performs vector-only semantic search using cosine similarity against document embeddings. It does not fall back to BM25 keyword matching or use the FTS5 index, which means it may miss documents containing exact keywords but lacking semantic similarity in the embedding space.

What makes the query command slower than search or vsearch?

The query command implements a multi-stage hybrid pipeline that executes BM25 retrieval, conditional LLM query expansion using Qwen3, vector similarity search, and neural re-ranking with Qwen3-reranker. Each stage adds computational overhead compared to the single-pass algorithms used by search and vsearch.

Can I use these commands programmatically in my own scripts?

Yes. Each CLI command wraps functions exported from src/store.ts that you can import directly into Node.js or Bun applications. Import searchFTS for BM25 retrieval, vectorSearchQuery for vector search, and hybridQuery for the full hybrid pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →