# How llmfit Handles Multi-GPU VRAM Detection and Aggregation for Model Fitting

> Learn how llmfit detects and aggregates multi-GPU VRAM for efficient model fitting. Discover its intelligent handling of identical GPU models for enhanced performance.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-08-20

---

**llmfit detects all GPUs on the host, aggregates VRAM from identical-model cards into a combined memory pool, and uses this total for fit-scoring while preserving primary GPU details for display.**

Multi-GPU machines are common in LLM inference and training workflows, but accurately reporting available VRAM requires more than simple enumeration. According to the llmfit source code, the [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) module implements a sophisticated detection pipeline that treats multi-GPU setups as a single logical memory pool for model compatibility checks.

## Multi-GPU VRAM Detection Pipeline

The detection workflow in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) follows four distinct phases: vendor-specific parsing, cross-vendor merging, primary GPU selection, and VRAM aggregation.

### Vendor-Specific GPU Discovery

llmfit queries each GPU vendor through native tooling and normalizes results into a common `GpuInfo` struct.

**NVIDIA GPUs** are parsed by `parse_nvidia_smi_list` (and its extended variant), which extracts:
- `model` — the GPU product name
- `count` — number of identical cards
- `vram_gb` — per-card memory capacity

```rust
// From llmfit-core/src/hardware.rs lines 37-40
pub struct GpuInfo {
    pub model: String,
    pub count: u32,      // Number of identical cards
    pub vram_gb: f32,    // Per-card VRAM
    // ...
}

```

**AMD GPUs** follow the same pattern through `parse_rocm_smi_output` and `detect_amd_gpu_sysfs_info`, also returning `count` and per-card `vram_gb` (lines 55-58).

### Merging Results and Filtering Integrated GPUs

The `detect_all_gpus` function (lines 98-105) combines NVIDIA, AMD, Intel, Apple, Ascend, and other vendor detections into a single `Vec<GpuInfo>`. It then **prefers discrete GPUs** and removes integrated graphics when a discrete card is present—ensuring that hybrid laptop configurations don't skew fit calculations with misleadingly low iGPU memory.

### Primary GPU Selection

After merging, llmfit sorts GPUs by VRAM capacity and designates the first entry as the **primary GPU**:

```rust
// Conceptual flow from lines 1000-1005
gpus.sort_by(|a, b| b.vram_gb.partial_cmp(&a.vram_gb).unwrap());
let primary = gpus.first();  // Highest VRAM card for display

```

The primary GPU's VRAM appears in user-facing output, while all GPU counts contribute to backend calculations.

## Aggregating Total VRAM for Fit Scoring

The critical aggregation happens in `total_gpu_vram_gb` calculation (lines 1010-1019):

```rust
// Summation across all detected GPUs
total_gpu_vram_gb = gpus.iter()
    .map(|g| g.vram_gb * g.count as f32)
    .sum();

```

This **per-card VRAM × card count** multiplication ensures that two 24 GB RTX 4090 cards correctly report 48 GB of aggregate memory. The fit scorer in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) uses this total to evaluate model compatibility—treating the multi-GPU machine as one logical pool rather than forcing models onto a single card.

### SystemSpecs Data Structure

The `SystemSpecs` struct (lines 48-54) exposes both values for different use cases:

| Field | Purpose |
|-------|---------|
| `gpu_vram_gb` | Primary GPU VRAM for display |
| `total_gpu_vram_gb` | Aggregated pool for fit decisions |

## Practical Usage Examples

### CLI System Information

Display detected hardware with automatic VRAM aggregation:

```bash
$ cargo run --quiet -- system
System:
  CPU: 12‑core Intel(R) Xeon
  RAM: 64.0 GB
  GPUs:
    • NVIDIA RTX 4090 (2 × 24 GB)  ← primary GPU, 24 GB shown
    • total GPU VRAM: 48 GB        ← aggregated VRAM for fit scoring

```

### Override for Testing

Force a specific aggregated VRAM value without physical hardware:

```bash
$ cargo run --quiet -- system --vram_gb 48

# Treats system as having 48 GB total pool, equivalent to

# 2×24 GB or 4×12 GB configurations

```

### Programmatic Access

Use the core library directly for custom tooling:

```rust
use llmfit_core::hardware::SystemSpecs;

let specs = SystemSpecs::detect();
println!("Primary VRAM: {:.1} GB", specs.gpu_vram_gb.unwrap_or(0.0));
println!("Aggregated VRAM: {:.1} GB", specs.total_gpu_vram_gb.unwrap_or(0.0));

```

## Key Implementation Files

- **[`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs)** — Core detection, parsing, and aggregation logic; defines `SystemSpecs` with dual VRAM fields
- **[`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)** — Consumes `total_gpu_vram_gb` for model compatibility evaluation
- **[`llmfit-tui/src/display.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs)** — Renders primary GPU information in terminal output
- **[`llmfit-tui/src/serve_api.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs)** — Exposes both VRAM values via HTTP API endpoints

## Summary

- llmfit detects GPUs per-vendor through native tools (`nvidia-smi`, `rocm-smi`, sysfs)
- Identical-model cards are **counted and aggregated**, not listed individually
- The **primary GPU** (highest VRAM) drives display output while **total aggregated VRAM** drives fit decisions
- Integrated GPUs are **filtered out** when discrete cards are present
- Both values are available programmatically via `SystemSpecs`

## Frequently Asked Questions

### How does llmfit distinguish between identical GPU models?

`parse_nvidia_smi_list` and AMD equivalents group cards by model name and return a `count` field. Two RTX 4090 cards become one `GpuInfo { model: "RTX 4090", count: 2, vram_gb: 24.0 }` entry, which `total_gpu_vram_gb` then expands to 48 GB.

### Can llmfit handle mixed GPU configurations?

Yes. `detect_all_gpus` merges NVIDIA, AMD, Intel, and other vendors into one list. However, `total_gpu_vram_gb` sums all detected cards regardless of vendor, so mixed setups (e.g., RTX 4090 + RX 7900 XTX) report their combined memory. The primary GPU selection still prefers the highest-VRAM single card.

### Why does llmfit show different VRAM values in display versus fit calculations?

The `gpu_vram_gb` field shows the **primary card's** capacity for user clarity—"you have an RTX 4090." The `total_gpu_vram_gb` field reports the **operational pool** for the fit scorer, which must account for all available memory when evaluating model compatibility across multiple cards.

### Is there a way to disable multi-GPU aggregation?

The `--vram_gb` CLI flag overrides detection entirely, treating the supplied value as the total pool. There's no toggle to list cards individually; the aggregation design is fundamental to llmfit's approach of treating multi-GPU hosts as unified memory systems.