# How the LLM Cache in QMD Delivers 10-30× Performance Improvements

> Discover how the QMD LLM cache achieves 10-30x performance gains by storing responses locally, slashing latency from seconds to milliseconds and eliminating redundant network calls.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: performance
- Published: 2026-02-16

---

**The LLM cache in QMD stores raw model responses in a local SQLite database keyed by SHA-256 hashes of requests, eliminating redundant network calls and reducing query expansion and reranking latency from seconds to milliseconds.**

QMD (Query Markup Documents) is an open-source framework that orchestrates large language model (LLM) operations—including query expansion, semantic reranking, and vector embedding—to enhance search and retrieval pipelines. Because every call to an LLM backend incurs significant network latency and consumes remote compute resources, the LLM cache in QMD provides a deterministic deduplication layer that transforms expensive HTTP round-trips into inexpensive local SQLite lookups.

## How the LLM Cache in QMD Works

### Cache Key Generation

At the heart of the caching mechanism is a deterministic hash function that generates unique cache keys based on the exact request parameters. Located in **[src/store.ts#L1104-L1110]**, the `getCacheKey` function creates a SHA-256 hash of both the endpoint URL and the serialized request body:

```typescript
export function getCacheKey(url: string, body: object): string {
  const hash = createHash("sha256");
  hash.update(url);                         // the endpoint (e.g. "expandQuery")
  hash.update(JSON.stringify(body));        // request payload (query, model, etc.)
  return hash.digest("hex");
}

```

By incorporating both the endpoint name and the full request payload into the hash, QMD guarantees that only *exact* duplicate calls hit the cache, preventing false positives when parameters differ slightly.

### Reading and Writing Cached Results

The cache interacts with a local SQLite database through two core functions defined in **[src/store.ts#L1111-L1116]** and **[src/store.ts#L1116-L1121]**:

```typescript
export function getCachedResult(db: Database, cacheKey: string): string | null {
  const row = db.prepare(`SELECT result FROM llm_cache WHERE key = ?`).get(cacheKey);
  return row?.result ?? null;
}

export function setCachedResult(db: Database, cacheKey: string, result: string): void {
  db.prepare(`INSERT OR REPLACE INTO llm_cache (key, result) VALUES (?, ?)`).run(cacheKey, result);
}

```

When a request is made, QMD first invokes `getCachedResult`. If a hit occurs, the stored JSON payload is returned immediately, bypassing the remote LLM entirely. On a cache miss, the result from the LLM is persisted using `setCachedResult`, ensuring subsequent identical requests benefit from the cached value.

## Performance Impact in Real-World Workflows

### Query Expansion Caching

Query expansion generates semantic variants of user queries to improve retrieval recall. The implementation in **[src/store.ts#L2204-L2228]** demonstrates how the LLM cache in QMD eliminates redundant expansion calls:

```typescript
export async function expandQuery(query: string, model = DEFAULT_QUERY_MODEL, db: Database) {
  const cacheKey = getCacheKey("expandQuery", { query, model });
  const cached = getCachedResult(db, cacheKey);
  if (cached) return JSON.parse(cached) as ExpandedQuery[];

  const expanded = await llm.expandQuery(query);
  setCachedResult(db, cacheKey, JSON.stringify(expanded));
  return expanded;
}

```

Only the **first** invocation of a particular query-model pair contacts the LLM; subsequent calls are served from the cache. This turns a multi-second network round-trip into a sub-millisecond SQLite lookup, making repeated query expansions feel instantaneous.

### Reranking Caching

Semantic reranking evaluates document relevance using LLM-based scoring. The cache implementation in **[src/store.ts#L2245-L2264]** ensures that identical document chunks are scored only once:

```typescript
const cacheKey = getCacheKey("rerank", {
  query,
  file: result.file,
  model,
  chunk: textByFile.get(result.file) ?? ""
});
const cached = getCachedResult(db, cacheKey);
if (cached) {
  result.score = Number(cached);
} else {
  // … compute score via LLM …
  setCachedResult(db, cacheKey, result.score.toString());
}

```

Each unique combination of query, file path, model, and chunk text is cached separately. This is crucial when reranking large document sets where multiple queries might reference the same chunks, delivering **10-30× speed improvements** on subsequent runs.

## Cache Maintenance and Management

QMD provides explicit utilities to manage cache lifecycle and disk usage. The `clearCache` function in **[src/store.ts#L1124-L1127]** removes all cached entries:

```typescript
export function clearCache(db: Database): void {
  db.prepare(`DELETE FROM llm_cache`).run();
}

```

For more targeted cleanup, `deleteLLMCache` in **[src/store.ts#L1136-L1141]** drops the entire LLM cache table and returns the count of removed rows:

```typescript
export function deleteLLMCache(db: Database): number {
  const result = db.prepare(`DELETE FROM llm_cache`).run();
  return result.changes;
}

```

These functions are exposed through the CLI via the `qmd clear-cache` command, implemented in **[src/qmd.ts#L419]** and **[src/qmd.ts#L1407]**, allowing users to reclaim disk space or invalidate stale results after model upgrades.

## Summary

- The **LLM cache in QMD** stores raw model responses in a local SQLite database, eliminating redundant network calls to expensive LLM backends.
- **Deterministic SHA-256 hashing** of endpoint URLs and request payloads ensures that only exact duplicate queries hit the cache, preventing false matches.
- **Sub-millisecond SQLite lookups** replace multi-second HTTP round-trips, delivering 10-30× performance improvements for query expansion and reranking workflows.
- **Automatic cache management** via `clearCache` and CLI commands allows users to invalidate stale data and control disk usage without modifying application code.

## Frequently Asked Questions

### How does the LLM cache in QMD handle different model versions?

The cache key includes the model identifier as part of the request payload hash. When you change models, the SHA-256 hash changes, automatically isolating cached results from different model versions. If you upgrade a model and want to invalidate old caches, you can run `qmd clear-cache` to wipe the database.

### What is the performance difference between cached and non-cached LLM calls?

Non-cached LLM calls incur network latency ranging from 200 milliseconds to several seconds depending on the provider and payload size. Cached results are retrieved via SQLite indexed lookups that typically complete in under one millisecond on SSD storage. This difference makes repeated query expansions feel instantaneous and reduces reranking times by an order of magnitude.

### Where does QMD store the LLM cache data?

QMD persists the LLM cache in a SQLite database file located in the system's cache directory. The `llm_cache` table within this database stores key-value pairs where the key is a SHA-256 hex digest and the value is the raw JSON response from the LLM. You can access this database directly or use the provided `clearCache` and `deleteLLMCache` utilities for maintenance.

### Can I disable the LLM cache in QMD if needed?

While the source code shows that caching is integrated into core workflows like `expandQuery` and `rerank`, you can effectively bypass caching by initializing a fresh in-memory database or by calling `clearCache` before each operation. For persistent disabling, you would need to modify the `getCachedResult` calls in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts) to always return `null`, though this is not exposed as a configuration flag in the current implementation.