How llmfit Prioritizes VRAM Over System RAM for GPU Systems: A Deep Dive into the Architecture
llmfit explicitly checks GPU VRAM capacity before system RAM when selecting execution modes, preferring RunMode::Gpu whenever the model fits entirely in video memory.
The llmfit repository implements a VRAM-first memory management strategy that determines how large language models execute on hardware with discrete GPUs. This article examines the source code architecture that enforces this priority, from hardware detection through run-mode selection to user-facing output.
System Detection: Capturing GPU and RAM Specifications
Before any model can be analyzed, llmfit queries the host machine's capabilities. In llmfit-core/src/hardware.rs, the SystemSpecs struct captures both memory pools:
pub struct SystemSpecs {
pub total_ram_gb: f32,
pub gpu_vram_gb: Option<f32>,
pub has_gpu: bool,
// ... additional fields
}
The has_gpu boolean flag drives all subsequent VRAM-prioritized decisions. When detect() populates this struct, it sets gpu_vram_gb to Some(vram) only when a compatible GPU is present, making the distinction between GPU-capable and CPU-only systems explicit at the type level.
Memory Requirement Calculation: VRAM as the Primary Constraint
The core analysis logic resides in llmfit-core/src/fit.rs within ModelFit::analyze_inner. For any given model, the code first computes min_vram_gb—the minimum GPU memory required for efficient inference based on parameter count and quantization level.
When system.has_gpu is true, this value is immediately compared against system.gpu_vram_gb. The comparison order matters: VRAM feasibility is evaluated before any system RAM calculations, establishing the priority hierarchy in code.
// Simplified from fit.rs analysis logic
let gpu_fit = system.gpu_vram_gb.map(|vram| min_vram_gb <= vram);
This early VRAM check feeds directly into the run-mode selection that follows.
Run-Mode Selection: The VRAM-First Decision Tree
The definitive prioritization occurs in llmfit-core/src/plan.rs. The execution path selection follows a strict VRAM-first ordering:
-
RunMode::Gpu— Selected whensystem.has_gpuis true andmin_vram_gb <= system.gpu_vram_gb. The entire model resides in VRAM; this is the optimal path. -
RunMode::MoeOffload— For MoE (Mixture-of-Experts) models where active experts fit in VRAM. The GPU accelerates inference while inactive experts remain in system RAM. -
RunMode::CpuOffload— When the model exceeds VRAM but GPU acceleration remains partially viable. Weights spill to system RAM as needed. -
RunMode::CpuOnly— Fallback when no GPU is present or the model cannot leverage VRAM.
This branch ordering guarantees VRAM utilization is exhausted before any CPU-only path is considered. The logic explicitly prefers partial GPU utilization (CpuOffload) over abandoning the GPU entirely.
Fit Level Scoring: Reflecting VRAM Priority in Recommendations
The FitLevel enum in llmfit-core/src/fit.rs codifies the quality of hardware-model matching:
FitLevel::Perfect— "Recommended memory met on GPU"FitLevel::Good— Fits with minor constraintsFitLevel::Fair— Functional but suboptimalFitLevel::Poor— Not recommended
Crucially, Perfect requires GPU residency. The fit_level field derives from both run_mode and the memory utilization ratio (memory_required_gb / memory_available_gb). Because RunMode::Gpu is evaluated first in the selection logic, models fitting in VRAM automatically receive the highest fit level—regardless of abundant system RAM.
// From fit.rs - fit_level determination
let fit_level = match run_mode {
RunMode::Gpu if utilization <= 0.8 => FitLevel::Perfect,
RunMode::Gpu => FitLevel::Good,
RunMode::MoeOffload => FitLevel::Good,
RunMode::CpuOffload => FitLevel::Fair,
RunMode::CpuOnly => FitLevel::Poor,
};
User-Facing Output: Surfacing VRAM Decisions
Both interfaces make the VRAM-first policy visible. In llmfit-tui/src/display.rs, the results table includes a "Fits in VRAM" column computed from fit.run_mode == RunMode::Gpu:
// From display.rs rendering logic
let vram_badge = if fit.run_mode == RunMode::Gpu {
"✅"
} else {
"❌"
};
The CLI (llmfit-tui/src/main.rs) similarly highlights GPU execution paths in its formatted output, reinforcing that VRAM fit status is the primary compatibility indicator.
Practical Verification
Query your system's model compatibility to see VRAM prioritization in action:
llmfit fit
Example output demonstrating the decision logic:
| Model | Min VRAM (GB) | GPU VRAM (GB) | Run Mode | Fits in VRAM |
|---|---|---|---|---|
| Qwen2-7B-Q4_K_M | 4.2 | 12.0 | Gpu | ✅ |
| Qwen2-72B-Q4_K_M | 39.0 | 12.0 | CpuOffload | ❌ |
| DeepSeek-V2-Lite | 5.8 | 12.0 | MoeOffload | Partial |
Programmatically verify the selected path:
use llmfit_core::{hardware::SystemSpecs, fit::ModelFit, models::LlmModel};
let system = SystemSpecs::detect();
let model = LlmModel::from_name("Qwen2-7B-Q4_K_M")?;
let fit = ModelFit::analyze(&model, &system);
match fit.run_mode {
RunMode::Gpu => println!("VRAM-only execution: {:.1} GB", fit.memory_required_gb),
RunMode::CpuOffload => println!("Spilling to system RAM"),
_ => println!("Alternative execution path"),
}
Summary
- Hardware detection in
hardware.rscaptures GPU VRAM as a distinct capability from system RAM - Memory analysis in
fit.rscomputes GPU requirements before considering RAM alternatives - Run-mode selection in
plan.rsfollows a strict VRAM-first hierarchy: Gpu → MoeOffload → CpuOffload → CpuOnly - Fit scoring reserves
Perfectratings exclusively for GPU-resident models - User interfaces surface VRAM fit status as the primary compatibility metric
Frequently Asked Questions
How does llmfit handle systems with multiple GPUs?
The current SystemSpecs aggregates VRAM from all detected GPUs into gpu_vram_gb as a single value. Multi-GPU configurations are treated as a unified memory pool for fit determination. Per-GPU granularity is not exposed in the public API as of the analyzed version.
What happens when a model barely exceeds available VRAM?
llmfit selects RunMode::CpuOffload, keeping as many weights as possible in VRAM while spilling the remainder to system RAM. This preserves partial GPU acceleration rather than falling back to CpuOnly. The threshold is strict: any VRAM overrun triggers offload, with no attempt at aggressive memory compression.
Can llmfit be forced to ignore GPU VRAM and use system RAM instead?
No direct flag exists in the analyzed codebase. The RunMode selection is deterministic based on hardware detection and model requirements. Users would need to artificially modify SystemSpecs::detect() output or use environment variables to hide GPU availability from the detection logic.
Does llmfit consider quantization effects on VRAM requirements?
Yes. The min_vram_gb calculation in fit.rs incorporates quantization level (Q4, Q5, Q8, etc.) when determining memory needs. Lower-precision quantizations reduce the VRAM threshold, enabling larger models to qualify for RunMode::Gpu that would otherwise require offload.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →