How LLMFIT Uses Multi-GPU Tensor Splitting to Aggregate VRAM for Model Fit Scoring
LLMFIT sums the VRAM of all identical GPUs into total_gpu_vram_gb during hardware detection, then uses this aggregated memory pool to score model fit for tensor-parallel inference, enabling large models to run across multiple cards.
The LLMFIT project by AlexsJones provides intelligent hardware-aware model fitting for large language models. When multiple identical GPUs are present, the system leverages multi-GPU tensor splitting to aggregate VRAM, treating the combined memory as a single resource for fit scoring rather than evaluating each card in isolation.
Hardware Detection and VRAM Aggregation
LLMFIT begins by detecting available hardware through the SystemSpecs struct in llmfit-core/src/hardware.rs. During this phase, the system identifies each GPU by name and capacity, specifically grouping cards by model to ensure compatibility.
For identical GPUs, the code aggregates individual VRAM capacities into a single field:
// In llmfit-core/src/hardware.rs (lines 2198-2210)
// Per-GPU VRAM values are summed into total_gpu_vram_gb
The total_gpu_vram_gb field represents the sum of all VRAM across same-model GPUs. For example, two RTX 4090 cards (24 GB each) result in a total_gpu_vram_gb value of 48 GB. This aggregation occurs automatically during system initialization and forms the foundation for all subsequent fit calculations.
Execution Path Selection in ModelFit::analyze_inner
Once hardware detection completes, the ModelFit::analyze_inner function in llmfit-core/src/fit.rs determines whether to use CPU inference, single-GPU execution, or tensor-parallel distribution. The function evaluates two primary paths depending on whether cluster mode is enabled.
When system.cluster_mode is active, LLFMIT treats the aggregated VRAM as a distributed pool across multiple nodes:
// In llmfit-core/src/fit.rs (lines 399-410)
if system.cluster_mode {
// Cluster mode: vLLM with tensor parallelism across multiple nodes.
// Total VRAM is the sum across all nodes (NCCL handles distribution).
let pool = system.total_gpu_vram_gb.unwrap_or(0.0);
let tp_size = system.cluster_node_count;
// ...
}
For standard multi-GPU configurations without cluster mode, the code uses the same aggregated value for local tensor splitting:
// In llmfit-core/src/fit.rs (lines 442-452)
} else if let Some(system_vram) = system.total_gpu_vram_gb {
// Use total VRAM across all same-model GPUs for fit scoring.
// Multi-GPU inference (tensor splitting) is supported by llama.cpp, vLLM, etc.
// ...
}
In both cases, the system_vram or pool variable contains the aggregated VRAM total, allowing the choose_quant helper to select appropriate quantization levels that fit within the combined memory budget. If the model fits, the system sets RunMode::Gpu for local multi-GPU or RunMode::TensorParallel for cluster configurations.
Fit Scoring with Aggregated Memory
The actual fit scoring occurs in the score_fit function, which compares the model's memory requirements against the available aggregated VRAM. Because the VRAM pool represents the sum of all identical GPUs, models that would exceed a single card's capacity can still achieve favorable scores.
The scoring algorithm evaluates:
- Perfect: Ample headroom within the aggregated pool
- Good: Comfortable fit with reasonable headroom
- Marginal: Tight fit with minimal overhead
- TooTight: Insufficient aggregated memory
When cluster_mode is enabled, NCCL (NVIDIA Collective Communications Library) handles the actual tensor distribution across nodes, while the fit scoring logic validates that the total required memory remains within the summed total_gpu_vram_gb limit.
Command-Line Usage for Multi-GPU Inference
You can trigger VRAM aggregation and tensor-parallel scoring through the LLMFIT CLI:
# Example 1 – Let LLMFIT automatically use all detected GPUs.
$ llmfit fit llama-2-70b
In this mode, LLMFIT automatically detects all compatible GPUs, sums their VRAM, and determines if the model fits within the combined pool.
# Example 2 – Force tensor-parallel (multi-GPU) mode on a single machine.
# The flag --cluster enables cluster_mode; the node count defaults to the
# number of GPUs of the same model.
$ llmfit fit --cluster llama-2-70b
The --cluster flag explicitly activates the TensorParallel code path, which validates the model against the aggregated VRAM pool and prepares for distributed inference.
# Example 3 – Manually override the detected VRAM (useful for testing).
$ llmfit fit --gpu-memory-override 24.0 llama-2-13b
This override value is also summed across detected GPUs when calculating aggregate capacity, allowing you to simulate different hardware configurations.
Summary
SystemSpecsdetects individual GPUs and aggregates same-model VRAM intototal_gpu_vram_gbinllmfit-core/src/hardware.rs(lines 2198-2210).ModelFit::analyze_innerroutes execution to tensor-parallel paths when multiple GPUs are present, using the pooled VRAM value for capacity checks inllmfit-core/src/fit.rs.- Aggregated VRAM allows models exceeding single-GPU limits to score as Perfect or Good fits when split across multiple cards via tensor parallelism.
- Cluster mode extends this aggregation across physical nodes using NCCL, while the
--clusterCLI flag activates this behavior inllmfit-tui/src/main.rs.
Frequently Asked Questions
How does LLMFIT handle GPUs with different VRAM sizes?
LLMFIT aggregates VRAM specifically across same-model cards. If you have heterogeneous GPUs (different models with varying VRAM), the system groups them by model name and aggregates within those groups. The fit scoring logic evaluates each homogeneous group separately, as tensor parallelism requires identical GPU capabilities for optimal performance.
What is the difference between regular multi-GPU mode and cluster mode?
In regular multi-GPU mode, LLMFIT uses the aggregated total_gpu_vram_gb for local tensor splitting across GPUs in a single machine, typically using frameworks like llama.cpp or vLLM. In cluster mode (enabled via --cluster), the same aggregated VRAM calculation applies, but the system treats the pool as distributed across multiple nodes using NCCL for inter-node communication, with cluster_node_count determining the tensor-parallel size.
Which quantization formats benefit most from aggregated VRAM aggregation?
Large parameter models (such as 70B or 405B variants) that cannot fit in a single GPU's VRAM benefit most. By aggregating VRAM across two or more cards, LLMFIT can score these models as fitting at higher precision levels (Q4 or Q5) rather than requiring extreme quantization or CPU offloading. The choose_quant helper automatically selects the highest viable precision that fits within the summed memory pool.
How can I override VRAM detection if LLMFIT misdetects my hardware?
Use the --gpu-memory-override flag followed by the per-GPU VRAM value in gigabytes. LLMFIT will multiply this value by the number of detected same-model GPUs to calculate the aggregate pool. For example, --gpu-memory-override 24.0 with two GPUs results in an aggregate 48 GB pool for fit scoring, useful for testing compatibility before hardware acquisition or working around detection edge cases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →