# How QMD's MCP HTTP Daemon Manages Model Loading for Maximum Efficiency

> Discover how QMD's MCP HTTP daemon efficiently loads GGUF models, keeping them in memory across requests by default and only disposing lightweight contexts to maximize performance.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: internals
- Published: 2026-02-16

---

**QMD's MCP HTTP daemon maintains a singleton `LlamaCpp` instance that lazily loads GGUF models once and keeps them resident in memory across HTTP requests, using an inactivity timer that only disposes lightweight contexts—not the heavy model binaries—by default.**

The `tobi/qmd` repository implements a Model Context Protocol (MCP) server that exposes LLM capabilities over HTTP. Efficient model loading is critical because GGUF files can exceed several hundred megabytes, and reloading them for every request would introduce unacceptable latency. The daemon solves this through a sophisticated caching strategy that balances memory efficiency with request performance.

## The Singleton Pattern: One LlamaCpp Instance for All Requests

At the heart of QMD's model loading efficiency is a singleton pattern implemented in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts). The function `getDefaultLlamaCpp()` (lines 71‑78) creates a single `LlamaCpp` object the first time it is called and reuses it for all subsequent invocations.

All store-level operations—such as `store.searchFTS`, `store.searchVec`, `store.expandQuery`, and `store.rerank`—obtain the LLM through this function. Consequently, every HTTP request processed by the MCP daemon works with the same in-process model objects, eliminating the overhead of reinitializing the LLM infrastructure.

## Lazy Loading: Models Load Only When Needed

The QMD MCP HTTP daemon employs lazy loading to defer the expensive operation of reading GGUF files from disk until absolutely necessary. Within the `LlamaCpp` class in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts), three private methods handle this: `ensureEmbedModel`, `ensureGenerateModel`, and `ensureRerankModel` (lines 145‑168).

These methods check whether the respective model is already present in memory. If not, they invoke `llama.loadModel` to read the GGUF file from the cache directory. Because the `LlamaCpp` instance is shared across requests via the singleton pattern, the models remain in memory after the first request triggers their load, ensuring subsequent requests benefit from immediate availability.

## The Inactivity Timer: Keeping Models Resident

To balance memory usage with performance, QMD implements an inactivity timer that manages resource lifecycle without unloading the heavy model binaries by default. After each LLM operation, the method `touchActivity()` (lines 95‑108 in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts)) resets a timer (`inactivityTimer`).

When the timer fires—triggered after `DEFAULT_INACTIVITY_TIMEOUT_MS` (5 minutes, defined at lines 150‑154)—it calls `unloadIdleResources()`. By default, this routine **only disposes the lightweight contexts** (embedding and rerank contexts), not the underlying model binaries. The configuration flag `disposeModelsOnInactivity` defaults to `false` (lines 141‑146), meaning the heavy GGUF models stay resident in RAM as long as the daemon receives at least one request every five minutes.

If you require the models to be freed when the server becomes idle, you can set `disposeModelsOnInactivity: true` when constructing the `LlamaCpp` instance, either via an environment variable or a custom `LlamaCppConfig`.

## Session-Aware Safety Guards

The QMD MCP HTTP daemon includes safeguards to prevent models from being unloaded while actively processing requests. The unload routine checks `canUnloadLLM()` (exposed from [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts), lines 160‑166), which consults the `LLMSessionManager`.

This manager tracks active MCP tool calls and abort signals. If a request is still being processed, `canUnloadLLM()` returns false, preventing the inactivity timer from tearing down models mid-request. This guarantees that long-running operations like batch `deep_search` complete without interruption.

## Putting It All Together in the HTTP Daemon

When you start the QMD MCP HTTP server with:

```bash
qmd mcp --http --port 8181

```

The following sequence occurs in [`src/mcp.ts`](https://github.com/tobi/qmd/blob/main/src/mcp.ts):

1. **Create Store**: `const store = createStore();` (lines 31‑34) opens the SQLite index and prepares helper functions.
2. **Create MCP Server**: `const mcpServer = createMcpServer(store);` (lines 54‑55) registers QMD tools (`search`, `vector_search`, `get`, etc.) that internally call the LLM via the singleton `LlamaCpp`.
3. **Attach HTTP Transport**: `new WebStandardStreamableHTTPServerTransport({ enableJsonResponse: true })` (lines 55‑58) wraps the MCP server in a JSON-RPC-over-HTTP layer.
4. **Start HTTP Listener**: `httpServer.listen(port, "localhost", …)` (lines 50‑53) handles `/health` and `/mcp` POST requests. Each POST parses the JSON-RPC body and forwards it to `transport.handleRequest`, invoking the appropriate MCP tool.
5. **Model Reuse**: All tools call `store.searchVec` → `ensureDefaultLlamaCpp()` → lazy model load → `touchActivity()` (see [`store.ts`](https://github.com/tobi/qmd/blob/main/store.ts) → `searchVec` → [`llm.ts`](https://github.com/tobi/qmd/blob/main/llm.ts)). Because the same `LlamaCpp` instance persists for the daemon's lifetime, heavy model files load only once and stay resident.

## Summary

- **Singleton Architecture**: `getDefaultLlamaCpp()` in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) ensures one `LlamaCpp` instance serves all HTTP requests, eliminating per-request initialization overhead.
- **Lazy Loading**: Models (embedding, generation, rerank) load only when first needed via `ensureEmbedModel` and similar methods, reducing startup time.
- **Persistent Memory**: By default, `disposeModelsOnInactivity` is `false`, so GGUF models remain in RAM across the 5-minute inactivity window, with only lightweight contexts being freed.
- **Request Safety**: `canUnloadLLM()` and `LLMSessionManager` prevent model unloading during active MCP tool calls, ensuring long-running operations complete safely.

## Frequently Asked Questions

### How does QMD prevent reloading models for every HTTP request?

QMD uses a singleton `LlamaCpp` instance created by `getDefaultLlamaCpp()` in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts). All MCP tools obtain the LLM through this function, ensuring every HTTP request shares the same in-process model objects. The models are loaded lazily on first use and remain resident in memory for subsequent requests.

### What happens if the QMD MCP server is idle for several minutes?

After each LLM operation, `touchActivity()` resets a 5-minute inactivity timer (`DEFAULT_INACTIVITY_TIMEOUT_MS`). When the timer fires, `unloadIdleResources()` runs, but by default it only disposes lightweight embedding and rerank contexts—not the heavy GGUF model binaries. The models stay resident unless you explicitly set `disposeModelsOnInactivity: true`.

### Can QMD unload models while a request is still processing?

No. The `canUnloadLLM()` function in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) consults the `LLMSessionManager` to track active MCP tool calls and abort signals. If a request is in progress, the function returns false and prevents the inactivity timer from unloading models, ensuring long-running searches or batch operations complete without interruption.

### Which configuration options control model persistence in QMD?

The `LlamaCpp` constructor accepts a `disposeModelsOnInactivity` boolean (defaulting to `false` in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) lines 141‑146) and `DEFAULT_INACTIVITY_TIMEOUT_MS` (set to 5 minutes). You can override these via environment variables or a custom `LlamaCppConfig` to adjust how aggressively the daemon frees memory versus keeping models resident for fast request handling.