How QMD's MCP HTTP Daemon Manages Model Loading for Maximum Efficiency

QMD's MCP HTTP daemon maintains a singleton LlamaCpp instance that lazily loads GGUF models once and keeps them resident in memory across HTTP requests, using an inactivity timer that only disposes lightweight contexts—not the heavy model binaries—by default.

The tobi/qmd repository implements a Model Context Protocol (MCP) server that exposes LLM capabilities over HTTP. Efficient model loading is critical because GGUF files can exceed several hundred megabytes, and reloading them for every request would introduce unacceptable latency. The daemon solves this through a sophisticated caching strategy that balances memory efficiency with request performance.

The Singleton Pattern: One LlamaCpp Instance for All Requests

At the heart of QMD's model loading efficiency is a singleton pattern implemented in src/llm.ts. The function getDefaultLlamaCpp() (lines 71‑78) creates a single LlamaCpp object the first time it is called and reuses it for all subsequent invocations.

All store-level operations—such as store.searchFTS, store.searchVec, store.expandQuery, and store.rerank—obtain the LLM through this function. Consequently, every HTTP request processed by the MCP daemon works with the same in-process model objects, eliminating the overhead of reinitializing the LLM infrastructure.

Lazy Loading: Models Load Only When Needed

The QMD MCP HTTP daemon employs lazy loading to defer the expensive operation of reading GGUF files from disk until absolutely necessary. Within the LlamaCpp class in src/llm.ts, three private methods handle this: ensureEmbedModel, ensureGenerateModel, and ensureRerankModel (lines 145‑168).

These methods check whether the respective model is already present in memory. If not, they invoke llama.loadModel to read the GGUF file from the cache directory. Because the LlamaCpp instance is shared across requests via the singleton pattern, the models remain in memory after the first request triggers their load, ensuring subsequent requests benefit from immediate availability.

The Inactivity Timer: Keeping Models Resident

To balance memory usage with performance, QMD implements an inactivity timer that manages resource lifecycle without unloading the heavy model binaries by default. After each LLM operation, the method touchActivity() (lines 95‑108 in src/llm.ts) resets a timer (inactivityTimer).

When the timer fires—triggered after DEFAULT_INACTIVITY_TIMEOUT_MS (5 minutes, defined at lines 150‑154)—it calls unloadIdleResources(). By default, this routine only disposes the lightweight contexts (embedding and rerank contexts), not the underlying model binaries. The configuration flag disposeModelsOnInactivity defaults to false (lines 141‑146), meaning the heavy GGUF models stay resident in RAM as long as the daemon receives at least one request every five minutes.

If you require the models to be freed when the server becomes idle, you can set disposeModelsOnInactivity: true when constructing the LlamaCpp instance, either via an environment variable or a custom LlamaCppConfig.

Session-Aware Safety Guards

The QMD MCP HTTP daemon includes safeguards to prevent models from being unloaded while actively processing requests. The unload routine checks canUnloadLLM() (exposed from src/llm.ts, lines 160‑166), which consults the LLMSessionManager.

This manager tracks active MCP tool calls and abort signals. If a request is still being processed, canUnloadLLM() returns false, preventing the inactivity timer from tearing down models mid-request. This guarantees that long-running operations like batch deep_search complete without interruption.

Putting It All Together in the HTTP Daemon

When you start the QMD MCP HTTP server with:

qmd mcp --http --port 8181

The following sequence occurs in src/mcp.ts:

  1. Create Store: const store = createStore(); (lines 31‑34) opens the SQLite index and prepares helper functions.
  2. Create MCP Server: const mcpServer = createMcpServer(store); (lines 54‑55) registers QMD tools (search, vector_search, get, etc.) that internally call the LLM via the singleton LlamaCpp.
  3. Attach HTTP Transport: new WebStandardStreamableHTTPServerTransport({ enableJsonResponse: true }) (lines 55‑58) wraps the MCP server in a JSON-RPC-over-HTTP layer.
  4. Start HTTP Listener: httpServer.listen(port, "localhost", …) (lines 50‑53) handles /health and /mcp POST requests. Each POST parses the JSON-RPC body and forwards it to transport.handleRequest, invoking the appropriate MCP tool.
  5. Model Reuse: All tools call store.searchVec → ensureDefaultLlamaCpp() → lazy model load → touchActivity() (see store.ts → searchVec → llm.ts). Because the same LlamaCpp instance persists for the daemon's lifetime, heavy model files load only once and stay resident.

Summary

  • Singleton Architecture: getDefaultLlamaCpp() in src/llm.ts ensures one LlamaCpp instance serves all HTTP requests, eliminating per-request initialization overhead.
  • Lazy Loading: Models (embedding, generation, rerank) load only when first needed via ensureEmbedModel and similar methods, reducing startup time.
  • Persistent Memory: By default, disposeModelsOnInactivity is false, so GGUF models remain in RAM across the 5-minute inactivity window, with only lightweight contexts being freed.
  • Request Safety: canUnloadLLM() and LLMSessionManager prevent model unloading during active MCP tool calls, ensuring long-running operations complete safely.

Frequently Asked Questions

How does QMD prevent reloading models for every HTTP request?

QMD uses a singleton LlamaCpp instance created by getDefaultLlamaCpp() in src/llm.ts. All MCP tools obtain the LLM through this function, ensuring every HTTP request shares the same in-process model objects. The models are loaded lazily on first use and remain resident in memory for subsequent requests.

What happens if the QMD MCP server is idle for several minutes?

After each LLM operation, touchActivity() resets a 5-minute inactivity timer (DEFAULT_INACTIVITY_TIMEOUT_MS). When the timer fires, unloadIdleResources() runs, but by default it only disposes lightweight embedding and rerank contexts—not the heavy GGUF model binaries. The models stay resident unless you explicitly set disposeModelsOnInactivity: true.

Can QMD unload models while a request is still processing?

No. The canUnloadLLM() function in src/llm.ts consults the LLMSessionManager to track active MCP tool calls and abort signals. If a request is in progress, the function returns false and prevents the inactivity timer from unloading models, ensuring long-running searches or batch operations complete without interruption.

Which configuration options control model persistence in QMD?

The LlamaCpp constructor accepts a disposeModelsOnInactivity boolean (defaulting to false in src/llm.ts lines 141‑146) and DEFAULT_INACTIVITY_TIMEOUT_MS (set to 5 minutes). You can override these via environment variables or a custom LlamaCppConfig to adjust how aggressively the daemon frees memory versus keeping models resident for fast request handling.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →