How GpuBackends Affect LLM Fit Scoring in llmfit: A Technical Deep-Dive

The GpuBackend selected in llmfit directly determines your LLM fit score through raw-throughput multipliers (ranging from 70.0 for CpuX86 to 390.0 for Ascend), feature eligibility like TurboQuant, and runtime path selection.

The llmfit crate evaluates how well large language models fit a host system by estimating tokens per second (TPS) for each possible run mode. This estimation hinges on the detected or selected GPU backend, making backend choice one of the most significant levers in fit scoring. Each backend carries distinct performance characteristics, quantization support, and hardware-specific behaviors that flow directly into the final score.

GpuBackend Throughput Multipliers in plan.rs

The core scoring logic resides in llmfit-core/src/plan.rs, where eight backend variants define raw-throughput multipliers used during TPS estimation. These constants scale baseline model performance into run-mode-specific predictions:

Backend Multiplier Key Behavior
Cuda 220.0 Enables TurboQuant fast-path quantization
Metal 160.0 Apple-silicon only; unified memory skips CpuOffload
Rocm 180.0 AMD GPUs via generic GPU throughput formula
Vulkan 150.0 Generic GPU interface, no quantization shortcuts
Sycl 100.0 Conservative TPS estimates for lower-end backends
Ascend 390.0 Huawei Ascend chips; highest multiplier
CpuX86 70.0 x86 CPU fallback when no GPU detected
CpuArm 90.0 ARM CPU fallback for non-Apple platforms

These values are defined at lines 251-258 in plan.rs according to the llmfit source code.

Backend Selection and System Detection

The pipeline begins with default_gpu_backend(system), called at lines 347-351 in plan.rs. This function inspects detected hardware and assigns the most appropriate GpuBackend variant. The result is stored in SystemSpecs.backend, a field typed to the GpuBackend enum declared in llmfit-core/src/hardware.rs (lines 6-14).

// From hardware.rs - the GpuBackend enum definition
pub enum GpuBackend {
    Cuda,
    Metal,
    Rocm,
    Vulkan,
    Sycl,
    Ascend,
    CpuX86,
    CpuArm,
}

Once assigned, this backend propagates through the entire fit analysis.

How GpuBackend Impacts the Fit Score Calculation

During ModelFit::analyze_with_forced_runtime in fit.rs, the selected backend feeds into Plan::estimate_tps. Here the backend multiplier directly scales the model's baseline throughput:

  1. Raw performance scaling — baseline TPS × backend multiplier = estimated run-mode TPS
  2. Quantization eligibility — TurboQuant applies only when system.backend == GpuBackend::Cuda (see the conditional at lines 653-658 in plan.rs)
  3. Path selection logic — Metal on Apple-silicon eliminates CPU-offload consideration because RAM equals VRAM

Higher estimated TPS yields better fit scores. The same model can therefore receive dramatically different scores on CUDA versus Metal versus CPU-only systems.

Practical Examples: Controlling GpuBackend in llmfit

Rust API: Manually Specify a Backend

use llmfit_core::{
    hardware::{GpuBackend, SystemSpecs},
    models::ModelDatabase,
    analysis::build_model_fits,
};

fn main() -> anyhow::Result<()> {
    // Simulate a CUDA-capable system
    let specs = SystemSpecs {
        backend: GpuBackend::Cuda,  // Forces CUDA scoring path
        unified_memory: false,
        ..Default::default()
    };

    let db = ModelDatabase::new()?;
    let fits = build_model_fits(&specs, &db, None)?;

    let best = fits.iter().max_by_key(|f| f.score).unwrap();
    println!("Best model for CUDA backend: {}", best.model.name);
    Ok(())
}

This pattern is useful for benchmarking hypothetical hardware or debugging scoring discrepancies.

CLI Override: Force Backend Detection


# Simulate Metal scoring on any platform

cargo run -- --gpu-backend metal --cli fit --top 5

# Force CPU-only evaluation

cargo run -- --gpu-backend cpux86 --cli fit --top 5

# Evaluate Ascend-style scoring

cargo run -- --gpu-backend ascend --cli fit --top 5

The --gpu-backend argument is parsed in llmfit-tui/src/main.rs and injected into SystemSpecs before analysis begins.

Key Source Files Controlling GpuBackend Scoring

File Purpose
llmfit-core/src/hardware.rs Defines GpuBackend enum and hardware detection structures
llmfit-core/src/plan.rs Backend selection logic, throughput multipliers, and TPS estimation
llmfit-core/src/fit.rs Orchestrates fit analysis using backend-derived TPS values
llmfit-tui/src/main.rs CLI entry point; handles --gpu-backend override parsing
llmfit-tui/src/tui_ui.rs / serve_api.rs Display layers presenting backend-influenced scores

Summary

  • Throughput multipliers range 70.0-390.0 across eight GpuBackend variants, directly scaling fit scores
  • CUDA uniquely enables TurboQuant optimization via conditional check in plan.rs
  • Metal on Apple-silicon alters runtime path selection due to unified memory architecture
  • Backend selection occurs via default_gpu_backend() but can be overridden via CLI or Rust API
  • The same model configuration produces different fit scores depending on detected or selected backend

Frequently Asked Questions

Why does the same model get different fit scores on different machines?

The GpuBackend multiplier varies by detected hardware. A CUDA system applies 220.0× baseline TPS while a CPU-x86 system applies only 70.0×. Additionally, CUDA systems may qualify for TurboQuant acceleration, further widening the gap. As implemented in llmfit, these differences are intentional to reflect real-world performance variation.

Can I force llmfit to use a specific backend for testing?

Yes. Pass --gpu-backend <variant> to the CLI, or construct SystemSpecs with an explicit backend field in Rust. Both methods bypass automatic detection. See main.rs for CLI parsing and hardware.rs for valid enum variants.

Why is Ascend's multiplier (390.0) higher than CUDA's (220.0)?

The llmfit source code treats Huawei Ascend chips as aggressively optimized for LLM inference in its scoring model. This reflects the hardware's dedicated AI acceleration architecture. Actual deployed performance depends on model compatibility and driver maturity.

Does Metal skip quantization entirely?

No—Metal simply skips the CPU-offload path because Apple-silicon uses unified memory. Quantization still occurs, but without the TurboQuant fast-path available only to CUDA backends as gated at lines 653-658 in plan.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →