# How to Configure Custom GGUF Models for Embedding, Reranking, and Query Expansion in QMD

> Configure custom GGUF models in QMD for embedding reranking and query expansion using the LlamaCpp class. Override defaults for powerful text analysis.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: how-to-guide
- Published: 2026-02-16

---

**You can configure custom GGUF models in QMD by instantiating a new `LlamaCpp` class with your model URIs and registering it via `setDefaultLlamaCpp`, which globally overrides the default embedding, reranking, and query expansion models for all subsequent operations.**

QMD (Query Markdown) is an open-source retrieval-augmented generation (RAG) tool by tobi that ships with three default GGUF models for vector search pipelines. While these defaults work out of the box, production deployments often require swapping in specialized models for domain-specific embedding, reranking, or query expansion tasks. This guide explains how to configure custom GGUF models in QMD using the programmatic API defined in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts).

## Understanding QMD's Default GGUF Model Configuration

QMD defines three default model constants in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) (lines 177‑180) that are automatically downloaded from Hugging Face the first time they are needed:

| Constant | Default Model URI | Purpose |
|----------|------------------|---------|
| `DEFAULT_EMBED_MODEL` | `hf:ggml-org/embeddinggemma-300M-GGUF/embeddinggemma-300M-Q8_0.gguf` | Vector embeddings |
| `DEFAULT_RERANK_MODEL` | `hf:ggml-org/Qwen3-Reranker-0.6B-Q8_0-GGUF/qwen3-reranker-0.6b-q8_0.gguf` | Re-ranking |
| `DEFAULT_GENERATE_MODEL` | `hf:tobil/qmd-query-expansion-1.7B-gguf/qmd-query-expansion-1.7B-q4_k_m.gguf` | Query expansion |

These URIs use the `hf:` prefix, which `node-llama-cpp` resolves to Hugging Face Hub downloads. The parser for this format is implemented in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) at lines 202‑209.

## Preparing Your Custom GGUF Models

### Hosting Models on Hugging Face

QMD expects models to follow the `hf:<owner>/<repo>/<filename.gguf>` URI format. Upload your quantized GGUF files to a Hugging Face repository, ensuring the filename matches exactly.

### URI Format Requirements

The `hf:` scheme is mandatory. Local file paths are also supported if `node-llama-cpp` can resolve them, but Hugging Face URIs are recommended for reproducibility. The resolution logic in `LlamaCpp.resolveModel()` (lines 528‑531) delegates to `node-llama-cpp`'s `resolveModelFile` function.

## Downloading Custom Models with pullModels

Before using custom models, ensure they are cached locally. While `qmd pull` downloads the three defaults, you can programmatically fetch custom models using the `pullModels` function exported from [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) (lines 224‑227):

```typescript
import { pullModels } from "./llm.js";

await pullModels(
  [
    "hf:myorg/myrepo/custom-embed.gguf",
    "hf:myorg/myrepo/custom-rerank.gguf",
    "hf:myorg/myrepo/custom-expand.gguf"
  ],
  { refresh: false }   // Set to true to force re-download
);

```

This downloads models to `~/.cache/qmd/models/` (or your configured `modelCacheDir`).

## Configuring Custom Models in Code

### Instantiating LlamaCpp with Custom URIs

The `LlamaCpp` class constructor accepts a `LlamaCppConfig` object (defined at lines 327‑332 in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts)) that overrides the default URIs:

```typescript
import { LlamaCpp } from "./llm.js";

const customLLM = new LlamaCpp({
  embedModel:    "hf:myorg/myrepo/custom-embed.gguf",
  rerankModel:  "hf:myorg/myrepo/custom-rerank.gguf",
  generateModel:"hf:myorg/myrepo/custom-expand.gguf",
  // Optional: specify a custom cache directory
  // modelCacheDir: "/var/lib/qmd/models"
});

```

### Registering the Default Instance

To make QMD use your custom instance globally, call `setDefaultLlamaCpp` (exported at lines 1384‑1386 in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts)):

```typescript
import { setDefaultLlamaCpp } from "./llm.js";

setDefaultLlamaCpp(customLLM);

```

Once registered, all high-level QMD operations—including `qmd query`, `qmd vsearch`, and `qmd embed`—will use your custom models. The CLI entry point in [`src/qmd.ts`](https://github.com/tobi/qmd/blob/main/src/qmd.ts) (line 70) calls `getDefaultLlamaCpp()`, which returns your registered instance.

## Verifying Your Configuration

After registration, verify that QMD uses your models by running a test query. The embedding flow routes through `store.searchVec` (line 2149‑2150 in [`src/store.ts`](https://github.com/tobi/qmd/blob/main/src/store.ts)), which calls `llm.embed`. Reranking occurs in `llm.rerank` (lines 1854‑1856), and query expansion uses `llm.expandQuery` (lines 3353‑3355). If you registered your custom `LlamaCpp` instance, all these calls will resolve to your specified GGUF files.

## Summary

- QMD defines default GGUF models in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) as `DEFAULT_EMBED_MODEL`, `DEFAULT_RERANK_MODEL`, and `DEFAULT_GENERATE_MODEL`.
- Custom models must use the `hf:<owner>/<repo>/<file.gguf>` URI format or valid local paths.
- Download custom models programmatically using `pullModels()` before first use.
- Instantiate a new `LlamaCpp` with your model URIs and register it via `setDefaultLlamaCpp()` to override defaults globally.
- All QMD operations—embedding, reranking, and query expansion—will automatically use your custom configuration.

## Frequently Asked Questions

### How do I specify a local GGUF file instead of a Hugging Face URI?

You can pass an absolute file path directly to the `embedModel`, `rerankModel`, or `generateModel` fields in the `LlamaCpp` constructor. The `node-llama-cpp` library underlying QMD resolves local paths through its `resolveModelFile` function, provided the file exists and is readable.

### Can I override only one model (e.g., just the embedder) while keeping the defaults for reranking and generation?

Yes. The `LlamaCppConfig` interface allows partial overrides. If you omit `rerankModel` or `generateModel` when constructing `LlamaCpp`, the instance will fall back to the built-in defaults defined in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts). Only specify the URIs for the models you wish to replace.

### Where does QMD cache downloaded models, and can I change the location?

By default, QMD caches models in `~/.cache/qmd/models/`. You can override this by setting the `modelCacheDir` property in the `LlamaCpp` constructor configuration. This is useful for containerized deployments or when you need to store models on high-speed storage separate from the home directory.

### Is there a CLI flag to change models, or must I use the programmatic API?

QMD does not expose CLI flags (e.g., `--embed-model`) to swap models. Configuration must be done through the programmatic API by importing `LlamaCpp` and `setDefaultLlamaCpp` from [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts), creating an instance with your custom URIs, and registering it before invoking any search or embedding commands.