How the LLM Cache in QMD Delivers 10-30× Performance Improvements

The LLM cache in QMD stores raw model responses in a local SQLite database keyed by SHA-256 hashes of requests, eliminating redundant network calls and reducing query expansion and reranking latency from seconds to milliseconds.

QMD (Query Markup Documents) is an open-source framework that orchestrates large language model (LLM) operations—including query expansion, semantic reranking, and vector embedding—to enhance search and retrieval pipelines. Because every call to an LLM backend incurs significant network latency and consumes remote compute resources, the LLM cache in QMD provides a deterministic deduplication layer that transforms expensive HTTP round-trips into inexpensive local SQLite lookups.

How the LLM Cache in QMD Works

Cache Key Generation

At the heart of the caching mechanism is a deterministic hash function that generates unique cache keys based on the exact request parameters. Located in [src/store.ts#L1104-L1110], the getCacheKey function creates a SHA-256 hash of both the endpoint URL and the serialized request body:

export function getCacheKey(url: string, body: object): string {
  const hash = createHash("sha256");
  hash.update(url);                         // the endpoint (e.g. "expandQuery")
  hash.update(JSON.stringify(body));        // request payload (query, model, etc.)
  return hash.digest("hex");
}

By incorporating both the endpoint name and the full request payload into the hash, QMD guarantees that only exact duplicate calls hit the cache, preventing false positives when parameters differ slightly.

Reading and Writing Cached Results

The cache interacts with a local SQLite database through two core functions defined in [src/store.ts#L1111-L1116] and [src/store.ts#L1116-L1121]:

export function getCachedResult(db: Database, cacheKey: string): string | null {
  const row = db.prepare(`SELECT result FROM llm_cache WHERE key = ?`).get(cacheKey);
  return row?.result ?? null;
}

export function setCachedResult(db: Database, cacheKey: string, result: string): void {
  db.prepare(`INSERT OR REPLACE INTO llm_cache (key, result) VALUES (?, ?)`).run(cacheKey, result);
}

When a request is made, QMD first invokes getCachedResult. If a hit occurs, the stored JSON payload is returned immediately, bypassing the remote LLM entirely. On a cache miss, the result from the LLM is persisted using setCachedResult, ensuring subsequent identical requests benefit from the cached value.

Performance Impact in Real-World Workflows

Query Expansion Caching

Query expansion generates semantic variants of user queries to improve retrieval recall. The implementation in [src/store.ts#L2204-L2228] demonstrates how the LLM cache in QMD eliminates redundant expansion calls:

export async function expandQuery(query: string, model = DEFAULT_QUERY_MODEL, db: Database) {
  const cacheKey = getCacheKey("expandQuery", { query, model });
  const cached = getCachedResult(db, cacheKey);
  if (cached) return JSON.parse(cached) as ExpandedQuery[];

  const expanded = await llm.expandQuery(query);
  setCachedResult(db, cacheKey, JSON.stringify(expanded));
  return expanded;
}

Only the first invocation of a particular query-model pair contacts the LLM; subsequent calls are served from the cache. This turns a multi-second network round-trip into a sub-millisecond SQLite lookup, making repeated query expansions feel instantaneous.

Reranking Caching

Semantic reranking evaluates document relevance using LLM-based scoring. The cache implementation in [src/store.ts#L2245-L2264] ensures that identical document chunks are scored only once:

const cacheKey = getCacheKey("rerank", {
  query,
  file: result.file,
  model,
  chunk: textByFile.get(result.file) ?? ""
});
const cached = getCachedResult(db, cacheKey);
if (cached) {
  result.score = Number(cached);
} else {
  // … compute score via LLM …
  setCachedResult(db, cacheKey, result.score.toString());
}

Each unique combination of query, file path, model, and chunk text is cached separately. This is crucial when reranking large document sets where multiple queries might reference the same chunks, delivering 10-30× speed improvements on subsequent runs.

Cache Maintenance and Management

QMD provides explicit utilities to manage cache lifecycle and disk usage. The clearCache function in [src/store.ts#L1124-L1127] removes all cached entries:

export function clearCache(db: Database): void {
  db.prepare(`DELETE FROM llm_cache`).run();
}

For more targeted cleanup, deleteLLMCache in [src/store.ts#L1136-L1141] drops the entire LLM cache table and returns the count of removed rows:

export function deleteLLMCache(db: Database): number {
  const result = db.prepare(`DELETE FROM llm_cache`).run();
  return result.changes;
}

These functions are exposed through the CLI via the qmd clear-cache command, implemented in [src/qmd.ts#L419] and [src/qmd.ts#L1407], allowing users to reclaim disk space or invalidate stale results after model upgrades.

Summary

  • The LLM cache in QMD stores raw model responses in a local SQLite database, eliminating redundant network calls to expensive LLM backends.
  • Deterministic SHA-256 hashing of endpoint URLs and request payloads ensures that only exact duplicate queries hit the cache, preventing false matches.
  • Sub-millisecond SQLite lookups replace multi-second HTTP round-trips, delivering 10-30× performance improvements for query expansion and reranking workflows.
  • Automatic cache management via clearCache and CLI commands allows users to invalidate stale data and control disk usage without modifying application code.

Frequently Asked Questions

How does the LLM cache in QMD handle different model versions?

The cache key includes the model identifier as part of the request payload hash. When you change models, the SHA-256 hash changes, automatically isolating cached results from different model versions. If you upgrade a model and want to invalidate old caches, you can run qmd clear-cache to wipe the database.

What is the performance difference between cached and non-cached LLM calls?

Non-cached LLM calls incur network latency ranging from 200 milliseconds to several seconds depending on the provider and payload size. Cached results are retrieved via SQLite indexed lookups that typically complete in under one millisecond on SSD storage. This difference makes repeated query expansions feel instantaneous and reduces reranking times by an order of magnitude.

Where does QMD store the LLM cache data?

QMD persists the LLM cache in a SQLite database file located in the system's cache directory. The llm_cache table within this database stores key-value pairs where the key is a SHA-256 hex digest and the value is the raw JSON response from the LLM. You can access this database directly or use the provided clearCache and deleteLLMCache utilities for maintenance.

Can I disable the LLM cache in QMD if needed?

While the source code shows that caching is integrated into core workflows like expandQuery and rerank, you can effectively bypass caching by initializing a fresh in-memory database or by calling clearCache before each operation. For persistent disabling, you would need to modify the getCachedResult calls in src/store.ts to always return null, though this is not exposed as a configuration flag in the current implementation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →