# How LLMFIT Uses Multi-GPU Tensor Splitting to Aggregate VRAM for Model Fit Scoring

> Discover how LLMFIT aggregates VRAM across identical GPUs for model fit scoring. Learn to run large models across multiple cards with multi-GPU tensor splitting for efficient tensor-parallel inference.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-08-21

---

**LLMFIT sums the VRAM of all identical GPUs into `total_gpu_vram_gb` during hardware detection, then uses this aggregated memory pool to score model fit for tensor-parallel inference, enabling large models to run across multiple cards.**

The LLMFIT project by AlexsJones provides intelligent hardware-aware model fitting for large language models. When multiple identical GPUs are present, the system leverages **multi-GPU tensor splitting to aggregate VRAM**, treating the combined memory as a single resource for fit scoring rather than evaluating each card in isolation.

## Hardware Detection and VRAM Aggregation

LLMFIT begins by detecting available hardware through the `SystemSpecs` struct in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs). During this phase, the system identifies each GPU by name and capacity, specifically grouping cards by model to ensure compatibility.

For identical GPUs, the code aggregates individual VRAM capacities into a single field:

```rust
// In llmfit-core/src/hardware.rs (lines 2198-2210)
// Per-GPU VRAM values are summed into total_gpu_vram_gb

```

The `total_gpu_vram_gb` field represents the **sum of all VRAM across same-model GPUs**. For example, two RTX 4090 cards (24 GB each) result in a `total_gpu_vram_gb` value of 48 GB. This aggregation occurs automatically during system initialization and forms the foundation for all subsequent fit calculations.

## Execution Path Selection in ModelFit::analyze_inner

Once hardware detection completes, the `ModelFit::analyze_inner` function in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) determines whether to use CPU inference, single-GPU execution, or tensor-parallel distribution. The function evaluates two primary paths depending on whether cluster mode is enabled.

When `system.cluster_mode` is active, LLFMIT treats the aggregated VRAM as a distributed pool across multiple nodes:

```rust
// In llmfit-core/src/fit.rs (lines 399-410)
if system.cluster_mode {
    // Cluster mode: vLLM with tensor parallelism across multiple nodes.
    // Total VRAM is the sum across all nodes (NCCL handles distribution).
    let pool = system.total_gpu_vram_gb.unwrap_or(0.0);
    let tp_size = system.cluster_node_count;
    // ...
}

```

For standard multi-GPU configurations without cluster mode, the code uses the same aggregated value for local tensor splitting:

```rust
// In llmfit-core/src/fit.rs (lines 442-452)
} else if let Some(system_vram) = system.total_gpu_vram_gb {
    // Use total VRAM across all same-model GPUs for fit scoring.
    // Multi-GPU inference (tensor splitting) is supported by llama.cpp, vLLM, etc.
    // ...
}

```

In both cases, the `system_vram` or `pool` variable contains the **aggregated VRAM total**, allowing the `choose_quant` helper to select appropriate quantization levels that fit within the combined memory budget. If the model fits, the system sets `RunMode::Gpu` for local multi-GPU or `RunMode::TensorParallel` for cluster configurations.

## Fit Scoring with Aggregated Memory

The actual fit scoring occurs in the `score_fit` function, which compares the model's memory requirements against the available aggregated VRAM. Because the VRAM pool represents the sum of all identical GPUs, models that would exceed a single card's capacity can still achieve favorable scores.

The scoring algorithm evaluates:

- **Perfect**: Ample headroom within the aggregated pool
- **Good**: Comfortable fit with reasonable headroom
- **Marginal**: Tight fit with minimal overhead
- **TooTight**: Insufficient aggregated memory

When `cluster_mode` is enabled, NCCL (NVIDIA Collective Communications Library) handles the actual tensor distribution across nodes, while the fit scoring logic validates that the total required memory remains within the summed `total_gpu_vram_gb` limit.

## Command-Line Usage for Multi-GPU Inference

You can trigger VRAM aggregation and tensor-parallel scoring through the LLMFIT CLI:

```bash

# Example 1 – Let LLMFIT automatically use all detected GPUs.

$ llmfit fit llama-2-70b

```

In this mode, LLMFIT automatically detects all compatible GPUs, sums their VRAM, and determines if the model fits within the combined pool.

```bash

# Example 2 – Force tensor-parallel (multi-GPU) mode on a single machine.

# The flag --cluster enables cluster_mode; the node count defaults to the

# number of GPUs of the same model.

$ llmfit fit --cluster llama-2-70b

```

The `--cluster` flag explicitly activates the `TensorParallel` code path, which validates the model against the aggregated VRAM pool and prepares for distributed inference.

```bash

# Example 3 – Manually override the detected VRAM (useful for testing).

$ llmfit fit --gpu-memory-override 24.0 llama-2-13b

```

This override value is also summed across detected GPUs when calculating aggregate capacity, allowing you to simulate different hardware configurations.

## Summary

- **`SystemSpecs`** detects individual GPUs and aggregates same-model VRAM into `total_gpu_vram_gb` in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) (lines 2198-2210).
- **`ModelFit::analyze_inner`** routes execution to tensor-parallel paths when multiple GPUs are present, using the pooled VRAM value for capacity checks in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs).
- **Aggregated VRAM** allows models exceeding single-GPU limits to score as Perfect or Good fits when split across multiple cards via tensor parallelism.
- **Cluster mode** extends this aggregation across physical nodes using NCCL, while the `--cluster` CLI flag activates this behavior in [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs).

## Frequently Asked Questions

### How does LLMFIT handle GPUs with different VRAM sizes?

LLMFIT aggregates VRAM specifically across **same-model cards**. If you have heterogeneous GPUs (different models with varying VRAM), the system groups them by model name and aggregates within those groups. The fit scoring logic evaluates each homogeneous group separately, as tensor parallelism requires identical GPU capabilities for optimal performance.

### What is the difference between regular multi-GPU mode and cluster mode?

In **regular multi-GPU mode**, LLMFIT uses the aggregated `total_gpu_vram_gb` for local tensor splitting across GPUs in a single machine, typically using frameworks like llama.cpp or vLLM. In **cluster mode** (enabled via `--cluster`), the same aggregated VRAM calculation applies, but the system treats the pool as distributed across multiple nodes using NCCL for inter-node communication, with `cluster_node_count` determining the tensor-parallel size.

### Which quantization formats benefit most from aggregated VRAM aggregation?

Large parameter models (such as 70B or 405B variants) that cannot fit in a single GPU's VRAM benefit most. By aggregating VRAM across two or more cards, LLMFIT can score these models as fitting at higher precision levels (Q4 or Q5) rather than requiring extreme quantization or CPU offloading. The `choose_quant` helper automatically selects the highest viable precision that fits within the summed memory pool.

### How can I override VRAM detection if LLMFIT misdetects my hardware?

Use the `--gpu-memory-override` flag followed by the per-GPU VRAM value in gigabytes. LLMFIT will multiply this value by the number of detected same-model GPUs to calculate the aggregate pool. For example, `--gpu-memory-override 24.0` with two GPUs results in an aggregate 48 GB pool for fit scoring, useful for testing compatibility before hardware acquisition or working around detection edge cases.