How to Configure Custom GGUF Models for Embedding, Reranking, and Query Expansion in QMD

You can configure custom GGUF models in QMD by instantiating a new LlamaCpp class with your model URIs and registering it via setDefaultLlamaCpp, which globally overrides the default embedding, reranking, and query expansion models for all subsequent operations.

QMD (Query Markdown) is an open-source retrieval-augmented generation (RAG) tool by tobi that ships with three default GGUF models for vector search pipelines. While these defaults work out of the box, production deployments often require swapping in specialized models for domain-specific embedding, reranking, or query expansion tasks. This guide explains how to configure custom GGUF models in QMD using the programmatic API defined in src/llm.ts.

Understanding QMD's Default GGUF Model Configuration

QMD defines three default model constants in src/llm.ts (lines 177‑180) that are automatically downloaded from Hugging Face the first time they are needed:

Constant Default Model URI Purpose
DEFAULT_EMBED_MODEL hf:ggml-org/embeddinggemma-300M-GGUF/embeddinggemma-300M-Q8_0.gguf Vector embeddings
DEFAULT_RERANK_MODEL hf:ggml-org/Qwen3-Reranker-0.6B-Q8_0-GGUF/qwen3-reranker-0.6b-q8_0.gguf Re-ranking
DEFAULT_GENERATE_MODEL hf:tobil/qmd-query-expansion-1.7B-gguf/qmd-query-expansion-1.7B-q4_k_m.gguf Query expansion

These URIs use the hf: prefix, which node-llama-cpp resolves to Hugging Face Hub downloads. The parser for this format is implemented in src/llm.ts at lines 202‑209.

Preparing Your Custom GGUF Models

Hosting Models on Hugging Face

QMD expects models to follow the hf:<owner>/<repo>/<filename.gguf> URI format. Upload your quantized GGUF files to a Hugging Face repository, ensuring the filename matches exactly.

URI Format Requirements

The hf: scheme is mandatory. Local file paths are also supported if node-llama-cpp can resolve them, but Hugging Face URIs are recommended for reproducibility. The resolution logic in LlamaCpp.resolveModel() (lines 528‑531) delegates to node-llama-cpp's resolveModelFile function.

Downloading Custom Models with pullModels

Before using custom models, ensure they are cached locally. While qmd pull downloads the three defaults, you can programmatically fetch custom models using the pullModels function exported from src/llm.ts (lines 224‑227):

import { pullModels } from "./llm.js";

await pullModels(
  [
    "hf:myorg/myrepo/custom-embed.gguf",
    "hf:myorg/myrepo/custom-rerank.gguf",
    "hf:myorg/myrepo/custom-expand.gguf"
  ],
  { refresh: false }   // Set to true to force re-download
);

This downloads models to ~/.cache/qmd/models/ (or your configured modelCacheDir).

Configuring Custom Models in Code

Instantiating LlamaCpp with Custom URIs

The LlamaCpp class constructor accepts a LlamaCppConfig object (defined at lines 327‑332 in src/llm.ts) that overrides the default URIs:

import { LlamaCpp } from "./llm.js";

const customLLM = new LlamaCpp({
  embedModel:    "hf:myorg/myrepo/custom-embed.gguf",
  rerankModel:  "hf:myorg/myrepo/custom-rerank.gguf",
  generateModel:"hf:myorg/myrepo/custom-expand.gguf",
  // Optional: specify a custom cache directory
  // modelCacheDir: "/var/lib/qmd/models"
});

Registering the Default Instance

To make QMD use your custom instance globally, call setDefaultLlamaCpp (exported at lines 1384‑1386 in src/llm.ts):

import { setDefaultLlamaCpp } from "./llm.js";

setDefaultLlamaCpp(customLLM);

Once registered, all high-level QMD operations—including qmd query, qmd vsearch, and qmd embed—will use your custom models. The CLI entry point in src/qmd.ts (line 70) calls getDefaultLlamaCpp(), which returns your registered instance.

Verifying Your Configuration

After registration, verify that QMD uses your models by running a test query. The embedding flow routes through store.searchVec (line 2149‑2150 in src/store.ts), which calls llm.embed. Reranking occurs in llm.rerank (lines 1854‑1856), and query expansion uses llm.expandQuery (lines 3353‑3355). If you registered your custom LlamaCpp instance, all these calls will resolve to your specified GGUF files.

Summary

  • QMD defines default GGUF models in src/llm.ts as DEFAULT_EMBED_MODEL, DEFAULT_RERANK_MODEL, and DEFAULT_GENERATE_MODEL.
  • Custom models must use the hf:<owner>/<repo>/<file.gguf> URI format or valid local paths.
  • Download custom models programmatically using pullModels() before first use.
  • Instantiate a new LlamaCpp with your model URIs and register it via setDefaultLlamaCpp() to override defaults globally.
  • All QMD operations—embedding, reranking, and query expansion—will automatically use your custom configuration.

Frequently Asked Questions

How do I specify a local GGUF file instead of a Hugging Face URI?

You can pass an absolute file path directly to the embedModel, rerankModel, or generateModel fields in the LlamaCpp constructor. The node-llama-cpp library underlying QMD resolves local paths through its resolveModelFile function, provided the file exists and is readable.

Can I override only one model (e.g., just the embedder) while keeping the defaults for reranking and generation?

Yes. The LlamaCppConfig interface allows partial overrides. If you omit rerankModel or generateModel when constructing LlamaCpp, the instance will fall back to the built-in defaults defined in src/llm.ts. Only specify the URIs for the models you wish to replace.

Where does QMD cache downloaded models, and can I change the location?

By default, QMD caches models in ~/.cache/qmd/models/. You can override this by setting the modelCacheDir property in the LlamaCpp constructor configuration. This is useful for containerized deployments or when you need to store models on high-speed storage separate from the home directory.

Is there a CLI flag to change models, or must I use the programmatic API?

QMD does not expose CLI flags (e.g., --embed-model) to swap models. Configuration must be done through the programmatic API by importing LlamaCpp and setDefaultLlamaCpp from src/llm.ts, creating an instance with your custom URIs, and registering it before invoking any search or embedding commands.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →