How llmfit Handles Multi-GPU VRAM Detection and Aggregation for Model Fitting

llmfit detects all GPUs on the host, aggregates VRAM from identical-model cards into a combined memory pool, and uses this total for fit-scoring while preserving primary GPU details for display.

Multi-GPU machines are common in LLM inference and training workflows, but accurately reporting available VRAM requires more than simple enumeration. According to the llmfit source code, the hardware.rs module implements a sophisticated detection pipeline that treats multi-GPU setups as a single logical memory pool for model compatibility checks.

Multi-GPU VRAM Detection Pipeline

The detection workflow in llmfit-core/src/hardware.rs follows four distinct phases: vendor-specific parsing, cross-vendor merging, primary GPU selection, and VRAM aggregation.

Vendor-Specific GPU Discovery

llmfit queries each GPU vendor through native tooling and normalizes results into a common GpuInfo struct.

NVIDIA GPUs are parsed by parse_nvidia_smi_list (and its extended variant), which extracts:

  • model — the GPU product name
  • count — number of identical cards
  • vram_gb — per-card memory capacity
// From llmfit-core/src/hardware.rs lines 37-40
pub struct GpuInfo {
    pub model: String,
    pub count: u32,      // Number of identical cards
    pub vram_gb: f32,    // Per-card VRAM
    // ...
}

AMD GPUs follow the same pattern through parse_rocm_smi_output and detect_amd_gpu_sysfs_info, also returning count and per-card vram_gb (lines 55-58).

Merging Results and Filtering Integrated GPUs

The detect_all_gpus function (lines 98-105) combines NVIDIA, AMD, Intel, Apple, Ascend, and other vendor detections into a single Vec<GpuInfo>. It then prefers discrete GPUs and removes integrated graphics when a discrete card is present—ensuring that hybrid laptop configurations don't skew fit calculations with misleadingly low iGPU memory.

Primary GPU Selection

After merging, llmfit sorts GPUs by VRAM capacity and designates the first entry as the primary GPU:

// Conceptual flow from lines 1000-1005
gpus.sort_by(|a, b| b.vram_gb.partial_cmp(&a.vram_gb).unwrap());
let primary = gpus.first();  // Highest VRAM card for display

The primary GPU's VRAM appears in user-facing output, while all GPU counts contribute to backend calculations.

Aggregating Total VRAM for Fit Scoring

The critical aggregation happens in total_gpu_vram_gb calculation (lines 1010-1019):

// Summation across all detected GPUs
total_gpu_vram_gb = gpus.iter()
    .map(|g| g.vram_gb * g.count as f32)
    .sum();

This per-card VRAM × card count multiplication ensures that two 24 GB RTX 4090 cards correctly report 48 GB of aggregate memory. The fit scorer in llmfit-core/src/fit.rs uses this total to evaluate model compatibility—treating the multi-GPU machine as one logical pool rather than forcing models onto a single card.

SystemSpecs Data Structure

The SystemSpecs struct (lines 48-54) exposes both values for different use cases:

Field Purpose
gpu_vram_gb Primary GPU VRAM for display
total_gpu_vram_gb Aggregated pool for fit decisions

Practical Usage Examples

CLI System Information

Display detected hardware with automatic VRAM aggregation:

$ cargo run --quiet -- system
System:
  CPU: 12‑core Intel(R) Xeon
  RAM: 64.0 GB
  GPUs:
    • NVIDIA RTX 4090 (2 × 24 GB)  ← primary GPU, 24 GB shown
    • total GPU VRAM: 48 GB        ← aggregated VRAM for fit scoring

Override for Testing

Force a specific aggregated VRAM value without physical hardware:

$ cargo run --quiet -- system --vram_gb 48

# Treats system as having 48 GB total pool, equivalent to

# 2×24 GB or 4×12 GB configurations

Programmatic Access

Use the core library directly for custom tooling:

use llmfit_core::hardware::SystemSpecs;

let specs = SystemSpecs::detect();
println!("Primary VRAM: {:.1} GB", specs.gpu_vram_gb.unwrap_or(0.0));
println!("Aggregated VRAM: {:.1} GB", specs.total_gpu_vram_gb.unwrap_or(0.0));

Key Implementation Files

Summary

  • llmfit detects GPUs per-vendor through native tools (nvidia-smi, rocm-smi, sysfs)
  • Identical-model cards are counted and aggregated, not listed individually
  • The primary GPU (highest VRAM) drives display output while total aggregated VRAM drives fit decisions
  • Integrated GPUs are filtered out when discrete cards are present
  • Both values are available programmatically via SystemSpecs

Frequently Asked Questions

How does llmfit distinguish between identical GPU models?

parse_nvidia_smi_list and AMD equivalents group cards by model name and return a count field. Two RTX 4090 cards become one GpuInfo { model: "RTX 4090", count: 2, vram_gb: 24.0 } entry, which total_gpu_vram_gb then expands to 48 GB.

Can llmfit handle mixed GPU configurations?

Yes. detect_all_gpus merges NVIDIA, AMD, Intel, and other vendors into one list. However, total_gpu_vram_gb sums all detected cards regardless of vendor, so mixed setups (e.g., RTX 4090 + RX 7900 XTX) report their combined memory. The primary GPU selection still prefers the highest-VRAM single card.

Why does llmfit show different VRAM values in display versus fit calculations?

The gpu_vram_gb field shows the primary card's capacity for user clarity—"you have an RTX 4090." The total_gpu_vram_gb field reports the operational pool for the fit scorer, which must account for all available memory when evaluating model compatibility across multiple cards.

Is there a way to disable multi-GPU aggregation?

The --vram_gb CLI flag overrides detection entirely, treating the supplied value as the total pool. There's no toggle to list cards individually; the aggregation design is fundamental to llmfit's approach of treating multi-GPU hosts as unified memory systems.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →