# How QMD Implements Query Expansion to Improve Search Results

> Discover how QMD implements query expansion using a grammar guided LLM to generate typed query variations for enhanced search results. Explore lexical, vector, and hypothetical document expansions.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: internals
- Published: 2026-02-16

---

**QMD uses a grammar-guided LLM pipeline in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) to generate typed query variations—lexical, vector, and hypothetical document—then caches these expansions in SQLite to power hybrid search without redundant LLM calls.**

Query expansion is a retrieval technique that enriches a user's original search term with semantically related variations to improve recall and relevance. In the `tobi/qmd` repository, this process is implemented as a lightweight, local LLM-driven pipeline that transforms a single input into multiple typed query "flavors" before executing retrieval.

## LLM-Driven Query Expansion Pipeline

The core expansion logic resides in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) within the `expandQuery` method. This asynchronous function accepts a raw query string and optional configuration, then orchestrates a structured generation process using a local Llama.cpp model.

### Grammar-Guided Output Generation

To ensure predictable, parseable output from the LLM, QMD employs a custom grammar that constrains the model to emit strictly formatted lines. The grammar definition forces the model to produce lines matching the pattern `type: content`, where type must be one of `lex`, `vec`, or `hyde`.

```typescript
const grammar = await llama.createGrammar({
  grammar: `
    root ::= line+
    line ::= type ": " content "\\n"
    type ::= "lex" | "vec" | "hyde"
    content ::= [^\\n]+
  `,
});

```

This structured approach yields three distinct query variations:
- **lex**: Lexical variants optimized for full-text search (FTS)
- **vec**: Semantic paraphrases suited for vector embedding search
- **hyde**: Hypothetical document embeddings that simulate an ideal answer document

### Safety Filters and Validation

After parsing the grammar-constrained output, QMD applies several validation filters to ensure quality. The system discards empty lines, rejects unknown type identifiers, and filters out expansions that fail to contain at least one token from the original query. This prevents the LLM from generating completely unrelated search terms that could pollute results.

Additionally, callers can control lexical expansion generation via the `includeLexical` option. When set to `false`, the system filters out `lex` variants, which are computationally cheap but sometimes noisy, allowing the pipeline to focus on semantic and hypothetical document variations.

### Fallback Strategies

The implementation includes multiple fallback layers to ensure robustness. If the LLM fails to generate any valid expansions, or if the generation process throws an error, `expandQuery` returns a default set of queryables:

```typescript
const fallback: Queryable[] = [
  { type: "hyde", text: `Information about ${query}` },
  { type: "lex", text: query },
  { type: "vec", text: query },
];

```

In catastrophic failure scenarios where the LLM is unavailable, the method falls back to returning the original query as a vector search term, optionally prepending the lexical variant if `includeLexical` is enabled.

## Caching Query Expansions for Performance

To avoid redundant LLM invocations for repeated queries, QMD implements a caching layer in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts). The `expandQuery` wrapper function generates a deterministic cache key based on the query string and model identifier, then checks an SQLite-backed cache before invoking the LLM.

```typescript
export async function expandQuery(
  query: string,
  model: string = DEFAULT_QUERY_MODEL,
  db: Database
): Promise<ExpandedQuery[]> {
  const cacheKey = getCacheKey("expandQuery", { query, model });
  const cached = getCachedResult(db, cacheKey);
  if (cached) {
    try { return JSON.parse(cached) as ExpandedQuery[]; } catch {}
  }

  const llm = getDefaultLlamaCpp();
  const results = await llm.expandQuery(query);

  const expanded = results
    .filter((r) => r.text !== query)
    .map((r) => ({ type: r.type, text: r.text }));

  if (expanded.length) setCachedResult(db, cacheKey, JSON.stringify(expanded));
  return expanded;
}

```

This caching strategy significantly reduces latency for common queries and minimizes computational costs associated with local LLM inference.

## Integration with Hybrid Search

Query expansion integrates seamlessly into QMD's hybrid retrieval pipeline. During search execution in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts), the system first probes the index using BM25 to detect strong signals. If the initial lexical search yields high-confidence results, the pipeline skips expansion to conserve resources.

When expansion is required, the typed query variants are routed to their respective backends: `lex` expansions execute against the full-text search index, `vec` expansions generate embeddings for vector similarity search, and `hyde` expansions create hypothetical documents that bridge the lexical-semantic gap. The results from all variants are then fused using Reciprocal Rank Fusion (RRF) and reranked at the chunk level.

## Programmatic API and CLI Usage

QMD exposes query expansion through multiple interfaces. Command-line users automatically benefit from expansion when running:

```bash
qmd query "how to bake sourdough"

```

Developers can access the expansion layer directly through two primary APIs. The low-level LLM interface in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) returns raw `Queryable` objects:

```typescript
import { getDefaultLlamaCpp } from "./llm.js";

const llm = getDefaultLlamaCpp();
const variants = await llm.expandQuery("quantum computing algorithms");

// Returns typed variants: [{type: 'lex', text: '...'}, {type: 'vec', text: '...'}, ...]
console.log(variants);

```

For cached, type-safe access, the store layer in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) provides the recommended entry point:

```typescript
import { openStore } from "./store.js";

const store = await openStore("my-collection");
const expansions = await store.expandQuery("docker compose volumes");
// => [{type: "lex", text: "docker compose volumes guide"}, ...]
console.log(expansions);

```

## Summary

- QMD implements **query expansion** through a grammar-guided LLM pipeline in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) that generates three typed variants: **lexical** (`lex`), **vector** (`vec`), and **hypothetical document** (`hyde`).
- A **custom grammar** constrains the local Llama.cpp model to output parseable `type: content` lines, ensuring reliable structured generation without complex post-processing.
- **Safety filters** discard expansions that lack tokens from the original query, while **fallback strategies** ensure the search pipeline remains functional even when LLM generation fails.
- The **caching layer** in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) stores expansion results in SQLite, eliminating redundant LLM calls for repeated queries and significantly reducing latency.
- Query expansion integrates into the **hybrid search pipeline**, where typed variants route to specific backends (FTS, vector search, hypothetical document generation) and results are fused using Reciprocal Rank Fusion.

## Frequently Asked Questions

### What types of query expansions does QMD generate?

QMD generates three distinct expansion types through its LLM pipeline: **lexical** (`lex`) variants optimized for full-text search, **vector** (`vec`) paraphrases suited for semantic similarity search, and **hypothetical document** (`hyde`) expansions that simulate ideal answer documents to bridge lexical-semantic gaps.

### How does QMD ensure the LLM produces parseable query expansions?

QMD uses a **custom grammar** defined in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) that constrains the Llama.cpp model to emit strictly formatted lines matching `type: content`, where type must be `lex`, `vec`, or `hyde`. This grammar-guided generation eliminates the need for complex regex parsing or JSON validation of unstructured LLM output.

### What happens if the LLM fails to generate valid query expansions?

The `expandQuery` method implements multiple fallback layers. If the LLM returns no valid expansions or throws an error, the system returns a default set including a hypothetical document variant, the original query as a lexical term, and the original query as a vector term. In catastrophic failures where the LLM is unavailable, it falls back to the original query alone.

### Does QMD cache query expansions to improve performance?

Yes, QMD implements an SQLite-backed caching layer in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) that stores expansion results keyed by query string and model identifier. Subsequent identical queries retrieve cached results instead of invoking the LLM, dramatically reducing latency and computational overhead for repeated searches.