# Which GGUF Models Does QMD Support and How to Configure Them

> Discover which GGUF models QMD supports for embeddings, reranking, and query expansion. Learn to configure custom HuggingFace GGUF models with LlamaCppConfig.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: how-to-guide
- Published: 2026-02-16

---

**QMD ships with three pre-configured GGUF models for embeddings, reranking, and query expansion, all of which can be overridden via the `LlamaCppConfig` interface to use custom HuggingFace GGUF URIs.**

QMD (Query Markdown Database) is an open-source semantic search tool that leverages local LLMs to power vector search, reranking, and query expansion. Understanding which GGUF models are supported and how to configure them is essential for customizing QMD's retrieval pipeline or swapping in specialized models for your use case.

## Default GGUF Models in QMD

QMD bundles three specific GGUF models by default, each serving a distinct stage in the search pipeline:

| Purpose | Model Name | HuggingFace URI | Size |
|---------|-----------|-----------------|------|
| Vector Embeddings | embeddinggemma-300M-Q8_0 | `hf:ggml-org/embeddinggemma-300M-GGUF/embeddinggemma-300M-Q8_0.gguf` | ~300 MiB |
| Reranking | qwen3-reranker-0.6b-q8_0 | `hf:ggml-org/Qwen3-Reranker-0.6B-Q8_0-GGUF/qwen3-reranker-0.6b-q8_0.gguf` | ~640 MiB |
| Query Expansion | qmd-query-expansion-1.7B-q4_k_m | `hf:tobil/qmd-query-expansion-1.7B-gguf/qmd-query-expansion-1.7B-q4_k_m.gguf` | ~1.1 GiB |

## Where Default Models Are Defined

These URIs are hard-coded as constants in the LLM abstraction layer. In [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts), the defaults are defined at lines 175-180:

- `DEFAULT_EMBED_MODEL` (line 177) for embeddings
- `DEFAULT_RERANK_MODEL` (line 178) for reranking
- `DEFAULT_GENERATE_MODEL` (line 180) for query expansion

When you instantiate a `LlamaCpp` class without configuration, it automatically falls back to these defaults.

## Configuring Custom GGUF Models

You can override any default model by passing a `LlamaCppConfig` object to the `LlamaCpp` constructor. The configuration interface accepts HuggingFace-style URIs for each model type:

```typescript
export type LlamaCppConfig = {
  embedModel?: string;      // e.g., "hf:org/model/gguf/file.gguf"
  generateModel?: string;   // e.g., "hf:org/model/gguf/file.gguf"
  rerankModel?: string;     // e.g., "hf:org/model/gguf/file.gguf"
  modelCacheDir?: string;   // Default: ~/.cache/qmd/models/
  inactivityTimeoutMs?: number;
  disposeModelsOnInactivity?: boolean;
};

```

Example instantiation with custom GGUF models:

```typescript
import { LlamaCpp } from "./llm.js";

const llm = new LlamaCpp({
  embedModel: "hf:myorg/custom-embed/gguf/model-Q4_0.gguf",
  rerankModel: "hf:myorg/custom-rerank/gguf/model-Q8_0.gguf",
  generateModel: "hf:myorg/custom-gen/gguf/model-Q4_K_M.gguf",
  modelCacheDir: "/var/cache/qmd"
});

```

## Downloading and Updating GGUF Models

QMD provides a CLI command to fetch the default GGUF models. The `pull` sub-command, implemented in [`src/qmd.ts`](https://github.com/tobi/qmd/blob/main/src/qmd.ts) at lines 2404-2409, enumerates the three default URIs and downloads them from HuggingFace if missing.

```bash

# Download all three default models to ~/.cache/qmd/models/

qmd pull

# Force re-download to update to latest versions

qmd pull --refresh

```

If you configure custom models programmatically, QMD automatically downloads them on first use during model initialization.

## How QMD Uses Each GGUF Model

Each model serves a specific function in the retrieval pipeline according to the source code implementation.

### Embedding Model

The embedding GGUF model powers vector search. When you call `store.searchVec()` or `store.embedDocument()`, QMD invokes `LlamaCpp.ensureEmbedModel()` at lines 548-550 in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) to load the model and generate text embeddings.

### Reranking Model

For result refinement, `store.rerank()` calls `LlamaCpp.ensureRerankModel()` to load the reranking GGUF model. This cross-encoder reorders initial retrieval results by relevance to improve precision.

### Query Expansion Model

The generation model handles query expansion via `store.expandQuery()`, which is used internally by `querySearch`. It calls `ensureGenerateModel()` and invokes `model.generate()` to rewrite or expand user queries for better recall.

## Summary

- QMD supports three specific GGUF models by default: `embeddinggemma-300M-Q8_0` for embeddings, `qwen3-reranker-0.6b-q8_0` for reranking, and `qmd-query-expansion-1.7B-q4_k_m` for query expansion.
- Default models are hard-coded in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) as `DEFAULT_EMBED_MODEL`, `DEFAULT_RERANK_MODEL`, and `DEFAULT_GENERATE_MODEL`.
- Override defaults by passing a `LlamaCppConfig` object with custom HuggingFace GGUF URIs to the `LlamaCpp` constructor.
- Use `qmd pull` to download default models and `qmd pull --refresh` to update them.

## Frequently Asked Questions

### Can I use my own custom GGUF models with QMD?

Yes. You can override any of the three default models by providing a `LlamaCppConfig` object when instantiating the `LlamaCpp` class. Pass HuggingFace-style URIs (e.g., `hf:org/model/gguf/file.gguf`) for the `embedModel`, `rerankModel`, or `generateModel` properties to use your own quantized models.

### Where does QMD store downloaded GGUF models?

By default, QMD caches models in `~/.cache/qmd/models/`. You can change this location by setting the `modelCacheDir` property in your `LlamaCppConfig` object, which is useful for CI environments or when running on systems with limited home directory space.

### How do I update the default GGUF models to the latest versions?

Run `qmd pull --refresh` from your terminal. This command forces QMD to re-download the three default models (embedding, reranking, and query expansion) from HuggingFace, ensuring you have the latest versions even if they were previously cached.

### What quantization levels do the default QMD models use?

The default models use different quantization levels optimized for their specific tasks: the embedding model uses Q8_0 (8-bit), the reranker uses Q8_0 (8-bit), and the query expansion model uses Q4_K_M (4-bit with medium K-quant). These balances provide good quality while keeping memory usage reasonable for local deployment.