Which GGUF Models Does QMD Support and How to Configure Them
QMD ships with three pre-configured GGUF models for embeddings, reranking, and query expansion, all of which can be overridden via the LlamaCppConfig interface to use custom HuggingFace GGUF URIs.
QMD (Query Markdown Database) is an open-source semantic search tool that leverages local LLMs to power vector search, reranking, and query expansion. Understanding which GGUF models are supported and how to configure them is essential for customizing QMD's retrieval pipeline or swapping in specialized models for your use case.
Default GGUF Models in QMD
QMD bundles three specific GGUF models by default, each serving a distinct stage in the search pipeline:
| Purpose | Model Name | HuggingFace URI | Size |
|---|---|---|---|
| Vector Embeddings | embeddinggemma-300M-Q8_0 | hf:ggml-org/embeddinggemma-300M-GGUF/embeddinggemma-300M-Q8_0.gguf |
~300 MiB |
| Reranking | qwen3-reranker-0.6b-q8_0 | hf:ggml-org/Qwen3-Reranker-0.6B-Q8_0-GGUF/qwen3-reranker-0.6b-q8_0.gguf |
~640 MiB |
| Query Expansion | qmd-query-expansion-1.7B-q4_k_m | hf:tobil/qmd-query-expansion-1.7B-gguf/qmd-query-expansion-1.7B-q4_k_m.gguf |
~1.1 GiB |
Where Default Models Are Defined
These URIs are hard-coded as constants in the LLM abstraction layer. In src/llm.ts, the defaults are defined at lines 175-180:
DEFAULT_EMBED_MODEL(line 177) for embeddingsDEFAULT_RERANK_MODEL(line 178) for rerankingDEFAULT_GENERATE_MODEL(line 180) for query expansion
When you instantiate a LlamaCpp class without configuration, it automatically falls back to these defaults.
Configuring Custom GGUF Models
You can override any default model by passing a LlamaCppConfig object to the LlamaCpp constructor. The configuration interface accepts HuggingFace-style URIs for each model type:
export type LlamaCppConfig = {
embedModel?: string; // e.g., "hf:org/model/gguf/file.gguf"
generateModel?: string; // e.g., "hf:org/model/gguf/file.gguf"
rerankModel?: string; // e.g., "hf:org/model/gguf/file.gguf"
modelCacheDir?: string; // Default: ~/.cache/qmd/models/
inactivityTimeoutMs?: number;
disposeModelsOnInactivity?: boolean;
};
Example instantiation with custom GGUF models:
import { LlamaCpp } from "./llm.js";
const llm = new LlamaCpp({
embedModel: "hf:myorg/custom-embed/gguf/model-Q4_0.gguf",
rerankModel: "hf:myorg/custom-rerank/gguf/model-Q8_0.gguf",
generateModel: "hf:myorg/custom-gen/gguf/model-Q4_K_M.gguf",
modelCacheDir: "/var/cache/qmd"
});
Downloading and Updating GGUF Models
QMD provides a CLI command to fetch the default GGUF models. The pull sub-command, implemented in src/qmd.ts at lines 2404-2409, enumerates the three default URIs and downloads them from HuggingFace if missing.
# Download all three default models to ~/.cache/qmd/models/
qmd pull
# Force re-download to update to latest versions
qmd pull --refresh
If you configure custom models programmatically, QMD automatically downloads them on first use during model initialization.
How QMD Uses Each GGUF Model
Each model serves a specific function in the retrieval pipeline according to the source code implementation.
Embedding Model
The embedding GGUF model powers vector search. When you call store.searchVec() or store.embedDocument(), QMD invokes LlamaCpp.ensureEmbedModel() at lines 548-550 in src/llm.ts to load the model and generate text embeddings.
Reranking Model
For result refinement, store.rerank() calls LlamaCpp.ensureRerankModel() to load the reranking GGUF model. This cross-encoder reorders initial retrieval results by relevance to improve precision.
Query Expansion Model
The generation model handles query expansion via store.expandQuery(), which is used internally by querySearch. It calls ensureGenerateModel() and invokes model.generate() to rewrite or expand user queries for better recall.
Summary
- QMD supports three specific GGUF models by default:
embeddinggemma-300M-Q8_0for embeddings,qwen3-reranker-0.6b-q8_0for reranking, andqmd-query-expansion-1.7B-q4_k_mfor query expansion. - Default models are hard-coded in
src/llm.tsasDEFAULT_EMBED_MODEL,DEFAULT_RERANK_MODEL, andDEFAULT_GENERATE_MODEL. - Override defaults by passing a
LlamaCppConfigobject with custom HuggingFace GGUF URIs to theLlamaCppconstructor. - Use
qmd pullto download default models andqmd pull --refreshto update them.
Frequently Asked Questions
Can I use my own custom GGUF models with QMD?
Yes. You can override any of the three default models by providing a LlamaCppConfig object when instantiating the LlamaCpp class. Pass HuggingFace-style URIs (e.g., hf:org/model/gguf/file.gguf) for the embedModel, rerankModel, or generateModel properties to use your own quantized models.
Where does QMD store downloaded GGUF models?
By default, QMD caches models in ~/.cache/qmd/models/. You can change this location by setting the modelCacheDir property in your LlamaCppConfig object, which is useful for CI environments or when running on systems with limited home directory space.
How do I update the default GGUF models to the latest versions?
Run qmd pull --refresh from your terminal. This command forces QMD to re-download the three default models (embedding, reranking, and query expansion) from HuggingFace, ensuring you have the latest versions even if they were previously cached.
What quantization levels do the default QMD models use?
The default models use different quantization levels optimized for their specific tasks: the embedding model uses Q8_0 (8-bit), the reranker uses Q8_0 (8-bit), and the query expansion model uses Q4_K_M (4-bit with medium K-quant). These balances provide good quality while keeping memory usage reasonable for local deployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →