# QMD CPU vs GPU Performance: Benchmarks and Device Selection Guide

> Discover QMD CPU vs GPU performance differences. See benchmarks showing 10x faster inference on GPU, reducing latency from 300ms to 30ms. Choose the right device for your needs.

- Repository: [Tobias Lütke/qmd](https://github.com/tobi/qmd)
- Tags: performance
- Published: 2026-02-16

---

**QMD achieves approximately 10× faster inference on GPU compared to CPU, reducing embedding latency from ~300ms to ~30ms per query when CUDA, Metal, or Vulkan acceleration is available.**

The `tobi/qmd` repository relies on **node-llama-cpp** for all LLM operations, including embeddings, generation, and reranking. Understanding the performance differences between CPU and GPU execution is critical for optimizing query latency, as the device selection logic in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) automatically handles acceleration when available but falls back to slower CPU processing when necessary.

## How QMD Detects and Selects CPU vs GPU

QMD uses a lazy initialization pattern to detect hardware capabilities the first time a model is requested. This detection logic prioritizes GPU acceleration while maintaining robust CPU fallback mechanisms.

### Lazy GPU Detection in ensureLlama()

The `ensureLlama()` method in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) handles device selection by querying available GPU types and attempting initialization in priority order:

```typescript
// src/llm.ts – ensureLlama()
private async ensureLlama(): Promise<Llama> {
  if (!this.llama) {
    // Detect available GPU types and use the best one.
    // @ts-expect-error node-llama-cpp API compat
    const gpuTypes = await getLlamaGpuTypes();
    const preferred = (["cuda", "metal", "vulkan"] as const)
      .find(g => gpuTypes.includes(g));
    
    // Attempt GPU initialization, fallback to CPU on failure
    try {
      this.llama = await getLlama({ gpu: preferred });
    } catch (e) {
      this.llama = await getLlama({ gpu: false });
    }
  }
  return this.llama;
}

```

The detection prioritizes **CUDA** first, followed by **Metal** (Apple Silicon), then **Vulkan**. If `getLlama()` throws an error during GPU initialization, QMD automatically catches the exception and reinitializes with `gpu: false`.

### CPU Fallback Warning

When the final `llama` instance reports `gpu === false`, QMD emits a warning to stderr to alert users of the performance degradation:

```typescript
if (!llama.gpu) {
  process.stderr.write(
    "QMD Warning: no GPU acceleration, running on CPU (slow). Run 'qmd status' for details.\n"
  );
}

```

This warning appears during the first model load, ensuring users are immediately aware that inference will proceed at CPU speeds.

## QMD CPU vs GPU Performance Benchmarks

The performance gap between CPU and GPU execution in QMD is substantial, with GPU acceleration providing approximately **10× faster** inference across all operation types.

### Measured Latency Comparison

Based on the built-in benchmark in [`src/bench-rerank.ts`](https://github.com/tobi/qmd/blob/main/src/bench-rerank.ts) and community reports, the typical performance characteristics are:

| Operation | CPU (Single-Core) | GPU (CUDA/Metal/Vulkan) | Speed-Up |
|-----------|-------------------|-------------------------|----------|
| **Embedding** (300M model) | ~300ms per query | ~30ms per query | **~10×** |
| **Reranking** (0.6B model) | ~1.2s per batch | ~120ms per batch | **~10×** |
| **Generation** (1.7B expansion) | ~2.5s per call | ~250ms per call | **~10×** |

These figures assume a modern multi-core CPU versus a mid-range GPU with sufficient VRAM. The exact multiplier varies based on model quantization levels and specific hardware generations.

### Why GPU Acceleration is Faster

The **node-llama-cpp** backend executes transformer inference as massive matrix multiplication operations. GPUs excel at this workload for three reasons:

1. **Parallel Execution**: GPUs process thousands of matrix operations simultaneously across CUDA cores or Metal compute units, while CPUs process them sequentially or with limited vectorization.
2. **Memory Bandwidth**: GPU VRAM provides significantly higher bandwidth than system RAM, reducing bottlenecks when loading model weights.
3. **Resident Weights**: When `gpuOffloading` is enabled, the entire model remains in GPU memory, eliminating CPU-GPU transfer latency during inference.

QMD respects the `gpuOffloading` flag from `node-llama-cpp`, which determines whether the model runs entirely on GPU or splits layers between CPU and GPU to accommodate limited VRAM.

## Checking Your QMD Device Status

To verify whether QMD is utilizing GPU acceleration, use the built-in status command:

```bash
qmd status

```

This command invokes `getDeviceInfo()` from [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) and renders a device summary in [`src/qmd.ts`](https://github.com/tobi/qmd/blob/main/src/qmd.ts):

```typescript
// src/qmd.ts – status display
const device = await llm.getDeviceInfo();
if (device.gpu) {
  console.log(`  GPU:      ${c.green}${device.gpu}${c.reset} (offloading: ${device.gpuOffloading ? 'yes' : 'no'})`);
} else {
  console.log(`  GPU:      ${c.yellow}none${c.reset} (running on CPU — models will be slow)`);
  console.log(`  ${c.dim}Tip: Install CUDA, Vulkan, or Metal support for GPU acceleration.${c.reset}`);
}

```

The output clearly indicates whether CUDA, Metal, or Vulkan is active, whether GPU offloading is enabled, and provides actionable guidance if running on CPU.

## Optimizing QMD Performance

### Enabling GPU Acceleration

No configuration flags are required to enable GPU support. QMD automatically detects and uses compatible GPUs during the first model load in `ensureLlama()`. However, you must ensure:

1. **CUDA**: NVIDIA drivers and CUDA toolkit 12.x are installed
2. **Metal**: Running on macOS with Apple Silicon (M1/M2/M3)
3. **Vulkan**: Compatible Vulkan drivers installed for AMD/Intel GPUs

If the automatic detection fails but hardware is present, check that `node-llama-cpp` was compiled with the appropriate backend support.

### Handling CPU-Only Environments

For environments without GPU access (CI pipelines, low-power laptops, or restricted VMs), consider these mitigations:

- **Use smaller GGUF models**: Quantized models (Q4_0, Q5_K_M) reduce CPU load significantly
- **Pre-compute embeddings**: Generate and cache embeddings during off-peak hours to avoid real-time CPU inference
- **Limit concurrent requests**: CPU performance degrades sharply under parallel load; serialize heavy operations like reranking

## Summary

- **QMD automatically prioritizes GPU acceleration** using CUDA, Metal, or Vulkan, falling back to CPU only when necessary.
- **Performance difference is approximately 10×**, with GPU processing embeddings in ~30ms versus ~300ms on CPU.
- **Device detection occurs lazily** in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) via `ensureLlama()`, which queries `getLlamaGpuTypes()` and handles initialization errors gracefully.
- **Users can verify active hardware** by running `qmd status`, which reports GPU type and offloading status via `getDeviceInfo()`.
- **No manual configuration is required** for GPU support, but CPU-only environments should use smaller models or pre-computed embeddings to mitigate the ~10× slowdown.

## Frequently Asked Questions

### How do I know if QMD is using my GPU?

Run the `qmd status` command in your terminal. If a GPU is active, the output shows the specific backend (CUDA, Metal, or Vulkan) and whether GPU offloading is enabled. If running on CPU, the status displays a yellow "none" warning and suggests installing GPU drivers.

### Can I force QMD to use CPU instead of GPU?

While QMD does not expose a direct CLI flag to disable GPU, you can force CPU execution by ensuring no compatible GPU runtime is installed (removing CUDA/Metal/Vulkan drivers) or by modifying the initialization in [`src/llm.ts`](https://github.com/tobi/qmd/blob/main/src/llm.ts) to pass `gpu: false` to `getLlama()`. The automatic fallback mechanism will also use CPU if GPU initialization throws an error.

### What GPU types does QMD support?

QMD supports three GPU backends through **node-llama-cpp**: **CUDA** for NVIDIA GPUs, **Metal** for Apple Silicon (M1/M2/M3), and **Vulkan** for AMD and Intel integrated graphics. The detection logic in `ensureLlama()` prioritizes them in that exact order (CUDA → Metal → Vulkan).

### Why is QMD slow even though I have a GPU?

If QMD reports CPU usage despite having a GPU, the **node-llama-cpp** binary likely was not compiled with the appropriate GPU support, or the necessary drivers (CUDA toolkit, Metal SDK, or Vulkan libraries) are missing from the system path. Additionally, if the model is too large for your GPU's VRAM, QMD may fall back to CPU execution or partial offloading, significantly reducing speed. Run `qmd status` to verify the current device configuration.