QMD CPU vs GPU Performance: Benchmarks and Device Selection Guide

QMD achieves approximately 10× faster inference on GPU compared to CPU, reducing embedding latency from ~300ms to ~30ms per query when CUDA, Metal, or Vulkan acceleration is available.

The tobi/qmd repository relies on node-llama-cpp for all LLM operations, including embeddings, generation, and reranking. Understanding the performance differences between CPU and GPU execution is critical for optimizing query latency, as the device selection logic in src/llm.ts automatically handles acceleration when available but falls back to slower CPU processing when necessary.

How QMD Detects and Selects CPU vs GPU

QMD uses a lazy initialization pattern to detect hardware capabilities the first time a model is requested. This detection logic prioritizes GPU acceleration while maintaining robust CPU fallback mechanisms.

Lazy GPU Detection in ensureLlama()

The ensureLlama() method in src/llm.ts handles device selection by querying available GPU types and attempting initialization in priority order:

// src/llm.ts – ensureLlama()
private async ensureLlama(): Promise<Llama> {
  if (!this.llama) {
    // Detect available GPU types and use the best one.
    // @ts-expect-error node-llama-cpp API compat
    const gpuTypes = await getLlamaGpuTypes();
    const preferred = (["cuda", "metal", "vulkan"] as const)
      .find(g => gpuTypes.includes(g));
    
    // Attempt GPU initialization, fallback to CPU on failure
    try {
      this.llama = await getLlama({ gpu: preferred });
    } catch (e) {
      this.llama = await getLlama({ gpu: false });
    }
  }
  return this.llama;
}

The detection prioritizes CUDA first, followed by Metal (Apple Silicon), then Vulkan. If getLlama() throws an error during GPU initialization, QMD automatically catches the exception and reinitializes with gpu: false.

CPU Fallback Warning

When the final llama instance reports gpu === false, QMD emits a warning to stderr to alert users of the performance degradation:

if (!llama.gpu) {
  process.stderr.write(
    "QMD Warning: no GPU acceleration, running on CPU (slow). Run 'qmd status' for details.\n"
  );
}

This warning appears during the first model load, ensuring users are immediately aware that inference will proceed at CPU speeds.

QMD CPU vs GPU Performance Benchmarks

The performance gap between CPU and GPU execution in QMD is substantial, with GPU acceleration providing approximately 10× faster inference across all operation types.

Measured Latency Comparison

Based on the built-in benchmark in src/bench-rerank.ts and community reports, the typical performance characteristics are:

Operation CPU (Single-Core) GPU (CUDA/Metal/Vulkan) Speed-Up
Embedding (300M model) ~300ms per query ~30ms per query ~10×
Reranking (0.6B model) ~1.2s per batch ~120ms per batch ~10×
Generation (1.7B expansion) ~2.5s per call ~250ms per call ~10×

These figures assume a modern multi-core CPU versus a mid-range GPU with sufficient VRAM. The exact multiplier varies based on model quantization levels and specific hardware generations.

Why GPU Acceleration is Faster

The node-llama-cpp backend executes transformer inference as massive matrix multiplication operations. GPUs excel at this workload for three reasons:

  1. Parallel Execution: GPUs process thousands of matrix operations simultaneously across CUDA cores or Metal compute units, while CPUs process them sequentially or with limited vectorization.
  2. Memory Bandwidth: GPU VRAM provides significantly higher bandwidth than system RAM, reducing bottlenecks when loading model weights.
  3. Resident Weights: When gpuOffloading is enabled, the entire model remains in GPU memory, eliminating CPU-GPU transfer latency during inference.

QMD respects the gpuOffloading flag from node-llama-cpp, which determines whether the model runs entirely on GPU or splits layers between CPU and GPU to accommodate limited VRAM.

Checking Your QMD Device Status

To verify whether QMD is utilizing GPU acceleration, use the built-in status command:

qmd status

This command invokes getDeviceInfo() from src/llm.ts and renders a device summary in src/qmd.ts:

// src/qmd.ts – status display
const device = await llm.getDeviceInfo();
if (device.gpu) {
  console.log(`  GPU:      ${c.green}${device.gpu}${c.reset} (offloading: ${device.gpuOffloading ? 'yes' : 'no'})`);
} else {
  console.log(`  GPU:      ${c.yellow}none${c.reset} (running on CPU — models will be slow)`);
  console.log(`  ${c.dim}Tip: Install CUDA, Vulkan, or Metal support for GPU acceleration.${c.reset}`);
}

The output clearly indicates whether CUDA, Metal, or Vulkan is active, whether GPU offloading is enabled, and provides actionable guidance if running on CPU.

Optimizing QMD Performance

Enabling GPU Acceleration

No configuration flags are required to enable GPU support. QMD automatically detects and uses compatible GPUs during the first model load in ensureLlama(). However, you must ensure:

  1. CUDA: NVIDIA drivers and CUDA toolkit 12.x are installed
  2. Metal: Running on macOS with Apple Silicon (M1/M2/M3)
  3. Vulkan: Compatible Vulkan drivers installed for AMD/Intel GPUs

If the automatic detection fails but hardware is present, check that node-llama-cpp was compiled with the appropriate backend support.

Handling CPU-Only Environments

For environments without GPU access (CI pipelines, low-power laptops, or restricted VMs), consider these mitigations:

  • Use smaller GGUF models: Quantized models (Q4_0, Q5_K_M) reduce CPU load significantly
  • Pre-compute embeddings: Generate and cache embeddings during off-peak hours to avoid real-time CPU inference
  • Limit concurrent requests: CPU performance degrades sharply under parallel load; serialize heavy operations like reranking

Summary

  • QMD automatically prioritizes GPU acceleration using CUDA, Metal, or Vulkan, falling back to CPU only when necessary.
  • Performance difference is approximately 10×, with GPU processing embeddings in ~30ms versus ~300ms on CPU.
  • Device detection occurs lazily in src/llm.ts via ensureLlama(), which queries getLlamaGpuTypes() and handles initialization errors gracefully.
  • Users can verify active hardware by running qmd status, which reports GPU type and offloading status via getDeviceInfo().
  • No manual configuration is required for GPU support, but CPU-only environments should use smaller models or pre-computed embeddings to mitigate the ~10× slowdown.

Frequently Asked Questions

How do I know if QMD is using my GPU?

Run the qmd status command in your terminal. If a GPU is active, the output shows the specific backend (CUDA, Metal, or Vulkan) and whether GPU offloading is enabled. If running on CPU, the status displays a yellow "none" warning and suggests installing GPU drivers.

Can I force QMD to use CPU instead of GPU?

While QMD does not expose a direct CLI flag to disable GPU, you can force CPU execution by ensuring no compatible GPU runtime is installed (removing CUDA/Metal/Vulkan drivers) or by modifying the initialization in src/llm.ts to pass gpu: false to getLlama(). The automatic fallback mechanism will also use CPU if GPU initialization throws an error.

What GPU types does QMD support?

QMD supports three GPU backends through node-llama-cpp: CUDA for NVIDIA GPUs, Metal for Apple Silicon (M1/M2/M3), and Vulkan for AMD and Intel integrated graphics. The detection logic in ensureLlama() prioritizes them in that exact order (CUDA → Metal → Vulkan).

Why is QMD slow even though I have a GPU?

If QMD reports CPU usage despite having a GPU, the node-llama-cpp binary likely was not compiled with the appropriate GPU support, or the necessary drivers (CUDA toolkit, Metal SDK, or Vulkan libraries) are missing from the system path. Additionally, if the model is too large for your GPU's VRAM, QMD may fall back to CPU execution or partial offloading, significantly reducing speed. Run qmd status to verify the current device configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →