# How QMD Combines BM25, Vector Search, and LLM Reranking for Hybrid Search

> Discover how QMD's eight-stage hybrid search pipeline uniquely merges BM25, vector search, and LLM reranking. Achieve superior relevance with advanced fusion and LLM techniques. Learn more today

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: how-to-guide
- Published: 2026-02-16

---

**QMD implements an eight-stage hybrid search pipeline that probes BM25 scores to trigger conditional LLM query expansion, fuses lexical and vector results using Reciprocal Rank Fusion, and applies a final LLM reranker with position-aware score blending to maximize result relevance.**

QMD (Query Markdown Database) is an open-source semantic search engine developed by tobi that merges classical information retrieval with modern AI techniques. According to the `tobi/qmd` source code, the system orchestrates BM25 keyword matching, dense vector similarity, and LLM-based reranking through a sophisticated pipeline defined in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts).

## The Eight-Stage Hybrid Search Pipeline

The core logic resides in the `hybridQuery` method within [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts). The process flows through eight distinct stages, each optimizing different aspects of retrieval quality and computational efficiency.

### Stage 1: BM25 Probe and Early Exit

The pipeline begins with an initial BM25 probe using `searchFTS` on the raw user query. If the top BM25 score is sufficiently high and the gap to the second result exceeds a defined threshold, QMD treats the query as "already well-covered" and skips the expensive LLM expansion phase entirely.

This optimization appears at lines `L2823-L2830` in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts), preventing unnecessary compute for straightforward keyword matches.

### Stage 2: LLM Query Expansion

When the BM25 probe indicates weak coverage, QMD invokes `llm.expandQuery` to generate typed query variants:

- **lex**: Lexical variations for BM25 search
- **vec**: Semantic variations for vector search  
- **hyde**: Hypothetical document embeddings (LLM-generated answer-like text)

Duplicates of the original query are removed before routing. This expansion logic is implemented at `L2842-L2846`.

### Stage 3: Type-Routed Search

QMD routes expanded queries to their respective search backends based on variant type:

- **Vector search (`searchVec`)**: Receives the original query plus any `vec` or `hyde` variants
- **BM25 search (`searchFTS`)**: Receives all `lex` variants

All results are collected into `rankedLists` for subsequent fusion. This routing occurs at `L2859-L2871`.

### Stage 4: Reciprocal Rank Fusion (RRF)

The separate result lists are merged using Reciprocal Rank Fusion, giving higher weight to the first two lists (original BM25 and first vector results). The fused list is trimmed to `candidateLimit` before proceeding to reranking.

The RRF implementation appears at `L29001-L29008`.

### Stage 5: Chunking and Best-Chunk Selection

Each candidate document is split into approximately 900-token chunks using `chunkDocument`. The system selects the chunk sharing the most query keywords for reranking, ensuring the LLM evaluates the most relevant passage within each document.

This selection logic is found at `L29018-L29031`.

### Stage 6: LLM Reranking

The selected chunks are fed to `llm.rerank`, which evaluates relevance using a local GGUF model. Reranking operates on chunks rather than full bodies to minimize token consumption while maintaining context relevance.

The reranking call occurs at `L29035-L29039`.

### Stage 7: Position-Aware Score Blending

QMD blends the RRF rank (converted to `1 / rank`) with the LLM reranker score using dynamic weights based on result position:

- **Top 3 results**: Weight 0.75
- **Positions 4-10**: Weight 0.60  
- **Remaining results**: Weight 0.40

The final `blendedScore` is computed at `L29040-L29055`.

### Stage 8: Deduplication and Thresholding

Duplicate file paths are removed, results below `minScore` are filtered out, and the list is sliced to the user-requested `limit` before returning final results.

This final filtering happens at `L29075-L29084`.

## BM25 Score Normalization

BM25 scores returned by SQLite FTS5 are negative values. Before UI rendering, QMD normalizes these to a 0-1 range using a sigmoid-like function, enabling intuitive percentage-based relevance displays.

This normalization is implemented in [`src/qmd.ts`](https://github.com/tobi/qmd/blob/main/src/qmd.ts) within the `normalizeBM25` function (lines `1735-1743`).

## Vector Search Implementation Details

The `searchVec` function in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) (lines `2056-2120`) handles dense retrieval by first computing or reusing query embeddings, then performing a two-step SQLite-vec lookup to avoid JOIN-related performance issues (as noted in comments referencing PR #23). Nearest-vector chunks are turned into document-level results, deduplicated by file path, and scored by Euclidean distance.

## LLM Reranking API

The reranker receives a list of `{file, text}` chunks, calls `llm.rerank`, caches per-chunk results, and returns scores sorted descending. This interface is defined in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) at lines `2237-2260`.

## Practical Usage Examples

### CLI Hybrid Search (Full Pipeline)

```bash

# Triggers BM25 probe, expansion, RRF, and LLM reranking

qmd query "how to backup a postgres database"

```

### Pure BM25 Search

```bash

# Fast keyword search without LLM expansion or reranking

qmd search "postgres backup"

```

### Pure Vector Search

```bash

# Semantic search using embeddings only

qmd vsearch "backup strategy for postgres"

```

### Programmatic API

```typescript
import { Store } from "./src/store.js";

async function demo() {
  const store = await Store.open();
  const results = await store.hybridQuery(
    "how to backup a postgres database",
    { limit: 5, minScore: 0.2 }
  );

  for (const r of results) {
    console.log(`${r.title} — score: ${r.score.toFixed(3)}`);
    console.log(r.bestChunk.slice(0, 200) + "…");
  }
}
demo();

```

## Summary

- **QMD** implements an eight-stage hybrid search pipeline in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) that intelligently balances speed and accuracy through conditional LLM expansion.
- **BM25 probing** at `L2823-L2830` determines whether to trigger expensive LLM query generation, optimizing for straightforward keyword matches.
- **Type-routed search** dispatches lexical variants to `searchFTS` and semantic variants to `searchVec`, collecting results for fusion.
- **Reciprocal Rank Fusion** at `L29001-L29008` merges heterogeneous result lists with position-aware weighting before candidate trimming.
- **LLM reranking** operates on 900-token chunks selected via keyword overlap, minimizing token usage while maximizing relevance precision at `L29035-L29039`.
- **Position-aware blending** dynamically weights the final score based on result position (0.75 for top 3, 0.60 for 4-10, 0.40 for remainder) at `L29040-L29055`.

## Frequently Asked Questions

### What is Reciprocal Rank Fusion (RRF) in QMD?

Reciprocal Rank Fusion is the algorithm QMD uses to combine result lists from BM25 and vector search into a single ranked list. It assigns a score of `1 / rank` to each result and sums these across lists, giving higher weight to the original BM25 and first vector results. This approach effectively balances lexical and semantic relevance signals without requiring score normalization between the disparate systems.

### How does QMD decide when to use LLM query expansion?

QMD uses an initial BM25 probe to evaluate query coverage before committing to LLM expansion. If the top BM25 result achieves a high score and the gap to the second result exceeds a defined threshold, the system considers the query "well-covered" and skips LLM expansion to save compute. Only when the BM25 probe indicates weak or ambiguous coverage does QMD invoke `llm.expandQuery` to generate lexical, vector, and hypothetical document variants.

### Why does QMD rerank chunks instead of full documents?

QMD splits documents into approximately 900-token chunks and selects the chunk sharing the most query keywords for LLM reranking. This approach minimizes token consumption during the expensive LLM inference phase while ensuring the model evaluates the most relevant passage within each document. Reranking chunks rather than full bodies allows QMD to process larger candidate sets efficiently without exceeding context window limits or incurring excessive latency.

### What is the difference between `searchFTS` and `searchVec` in QMD?

`searchFTS` performs BM25 keyword search using SQLite FTS5 on lexical query variants, returning results ranked by classical term frequency-inverse document frequency scores. `searchVec` handles dense retrieval by computing query embeddings and performing vector similarity search using SQLite-vec, scoring results by Euclidean distance on semantic variants including hypothetical document embeddings. The hybrid pipeline combines these complementary retrieval methods through Reciprocal Rank Fusion to leverage both exact matching and semantic understanding.