# What GPU Backends Does llmfit Support? A Complete Hardware Abstraction Guide

> Discover the extensive GPU backends llmfit supports including CUDA Metal ROCm Vulkan SYCL Ascend and CPU variants Learn about hardware abstraction in this comprehensive guide

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: api-reference
- Published: 2026-08-20

---

**llmfit supports eight GPU and CPU backends: CUDA, Metal, ROCm, Vulkan, SYCL, Ascend, and two CPU variants (ARM and x86), all defined in the `GpuBackend` enum in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs).**

The llmfit library provides a unified hardware abstraction layer that automatically detects and categorizes compute devices across diverse platforms. Whether you're running inference on an NVIDIA data center GPU, an Apple Silicon MacBook, or a Huawei Ascend NPU, llmfit's backend system identifies the hardware and selects appropriate performance profiles and execution paths.

## Complete List of Supported GPU Backends

The `GpuBackend` enum in [[`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs#L6-L14) defines every backend variant that llmfit can recognize:

| Backend | Variant | Description |
|---------|---------|-------------|
| **CUDA** | `GpuBackend::Cuda` | NVIDIA GPUs accessed via `nvidia-smi` or CUDA driver APIs |
| **Metal** | `GpuBackend::Metal` | Apple Silicon GPUs with unified memory architecture |
| **ROCm** | `GpuBackend::Rocm` | AMD GPUs on Linux with ROCm stack installed (`rocm-smi`) |
| **Vulkan** | `GpuBackend::Vulkan` | Cross-vendor GPU fallback using Vulkan API queries |
| **SYCL** | `GpuBackend::Sycl` | Intel oneAPI GPUs including Arc/Xe graphics |
| **CPU (ARM)** | `GpuBackend::CpuArm` | ARM-based CPUs including Apple Silicon in CPU-only mode |
| **CPU (x86)** | `GpuBackend::CpuX86` | Standard x86 desktop and server processors |
| **Ascend** | `GpuBackend::Ascend` | Huawei Ascend NPUs detected via `npu-smi` |

Each variant implements the `label()` method for human-readable display names and supports equality comparisons for backend-specific conditional logic.

## Automatic Backend Detection

The `SystemSpecs::detect()` function in [[`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs#L72-L79) performs comprehensive hardware discovery. It queries multiple platform-specific tools to populate GPU information:

- **`nvidia-smi`** for NVIDIA CUDA devices
- **`rocm-smi`** for AMD ROCm GPUs
- **`lspci`** for hardware enumeration on Linux
- **Windows WMI/PowerShell** for GPU queries on Windows
- **`system_profiler`** on macOS
- **`npu-smi`** for Huawei Ascend detection

The detection routine produces a `SystemSpecs` struct containing the primary `backend` field and a vector of `GpuInfo` structs, each recording its own backend classification.

## Using GPU Backends in Practice

The detected backend drives core functionality throughout llmfit. In [[`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs), the `default_gpu_backend` function maps `SystemSpecs` to a `GpuBackend` and applies backend-specific throughput coefficients for inference planning.

Here's how to work with backends programmatically:

```rust
use llmfit_core::hardware::{SystemSpecs, GpuBackend};

fn main() {
    // Detect system hardware configuration
    let specs = SystemSpecs::detect();

    // Access the primary GPU backend
    println!("Primary backend: {}", specs.backend.label());

    // Iterate through all detected GPUs
    for gpu in &specs.gpus {
        println!(
            "Device: {} | VRAM: {:.1} GB | Backend: {}",
            gpu.name,
            gpu.vram_gb.unwrap_or(0.0),
            gpu.backend.label()
        );
    }

    // Execute backend-specific code paths
    match specs.backend {
        GpuBackend::Cuda => enable_cuda_optimizations(),
        GpuBackend::Metal => enable_metal_memory_pools(),
        GpuBackend::Vulkan => enable_vulkan_fallback(),
        _ => println!("Using generic backend: {}", specs.backend.label()),
    }
}

fn enable_cuda_optimizations() {
    println!("CUDA path: enabling TensorRT and cuBLAS");
}

fn enable_metal_memory_pools() {
    println!("Metal path: configuring unified memory");
}

fn enable_vulkan_fallback() {
    println!("Vulkan path: using generic compute shaders");
}

```

The backend information also surfaces through llmfit's external interfaces. The HTTP API in [[`llmfit-tui/src/serve_api.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs) and the MCP server in [[`llmfit-tui/src/mcp_server.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs) both expose detected backends to clients and dashboards.

## Platform Coverage and Fallback Strategy

**Vendor-specific tools take priority** in the detection order. CUDA, ROCm, Metal, SYCL, and Ascend each have dedicated detection paths using their native utilities. 

**Vulkan serves as the universal fallback** when no vendor-specific tool responds. This ensures llmfit can identify AMD, Intel, and other GPUs even without proprietary driver stacks installed.

**CPU backends activate automatically** when no GPU is detected or when users explicitly force CPU-only execution. The distinction between `CpuArm` and `CpuX86` enables architecture-specific optimization paths.

## Backend Selection in Inference Planning

The `RunMode` selection logic uses `SystemSpecs.backend` to determine execution strategy. GPU-capable backends (CUDA, Metal, ROCm, Vulkan, SYCL, Ascend) trigger `RunMode::Gpu`, while CPU backends select `RunMode::Cpu`. Each backend carries characteristic performance coefficients that affect batch size recommendations, memory allocation strategies, and throughput estimates in the planning pipeline.

## Summary

- **Eight backends supported**: CUDA, Metal, ROCm, Vulkan, SYCL, Ascend, CpuArm, and CpuX86
- **Core definition**: `GpuBackend` enum in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs)
- **Automatic detection**: `SystemSpecs::detect()` queries platform tools and populates backend information
- **Usage pattern**: Access `specs.backend` for the primary GPU, iterate `specs.gpus` for all devices, match on variants for conditional logic
- **Integration points**: [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) for performance coefficients, [`serve_api.rs`](https://github.com/AlexsJones/llmfit/blob/main/serve_api.rs) and [`mcp_server.rs`](https://github.com/AlexsJones/llmfit/blob/main/mcp_server.rs) for external exposure

## Frequently Asked Questions

### How does llmfit handle systems with multiple GPUs?

`SystemSpecs::detect()` returns a vector of `GpuInfo` structs in `specs.gpus`, each with its own backend classification. The `specs.backend` field identifies the primary GPU, typically the one with the most VRAM or highest compute capability. The planning pipeline in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) evaluates all available GPUs when generating execution recommendations.

### Can I force llmfit to use a specific backend?

The detection system prioritizes vendor-specific tools, but you can influence backend selection by controlling which GPU drivers and utilities are available in your environment. Removing `nvidia-smi` from PATH, for example, would prevent CUDA detection and allow Vulkan fallback on compatible NVIDIA hardware. Explicit backend forcing requires modifying the `SystemSpecs` detection logic in [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs).

### What happens if no GPU is detected?

`SystemSpecs::detect()` falls back to `GpuBackend::CpuX86` or `GpuBackend::CpuArm` based on the host architecture. The inference planner adjusts its recommendations accordingly, suggesting smaller batch sizes and model quantization strategies appropriate for CPU execution. Both CPU backends support full llmfit functionality, albeit with reduced throughput compared to GPU acceleration.