# How llmfit Prioritizes VRAM Over System RAM for GPU Systems: A Deep Dive into the Architecture

> Discover how llmfit prioritizes VRAM over system RAM for GPU systems. Learn about its architecture and execution modes for optimal performance.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: architecture
- Published: 2026-08-20

---

**llmfit explicitly checks GPU VRAM capacity before system RAM when selecting execution modes, preferring `RunMode::Gpu` whenever the model fits entirely in video memory.**

The `llmfit` repository implements a VRAM-first memory management strategy that determines how large language models execute on hardware with discrete GPUs. This article examines the source code architecture that enforces this priority, from hardware detection through run-mode selection to user-facing output.

## System Detection: Capturing GPU and RAM Specifications

Before any model can be analyzed, `llmfit` queries the host machine's capabilities. In [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs), the `SystemSpecs` struct captures both memory pools:

```rust
pub struct SystemSpecs {
    pub total_ram_gb: f32,
    pub gpu_vram_gb: Option<f32>,
    pub has_gpu: bool,
    // ... additional fields
}

```

The `has_gpu` boolean flag drives all subsequent VRAM-prioritized decisions. When `detect()` populates this struct, it sets `gpu_vram_gb` to `Some(vram)` only when a compatible GPU is present, making the distinction between GPU-capable and CPU-only systems explicit at the type level.

## Memory Requirement Calculation: VRAM as the Primary Constraint

The core analysis logic resides in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) within `ModelFit::analyze_inner`. For any given model, the code first computes `min_vram_gb`—the minimum GPU memory required for efficient inference based on parameter count and quantization level.

When `system.has_gpu` is true, this value is immediately compared against `system.gpu_vram_gb`. The comparison order matters: VRAM feasibility is evaluated before any system RAM calculations, establishing the priority hierarchy in code.

```rust
// Simplified from fit.rs analysis logic
let gpu_fit = system.gpu_vram_gb.map(|vram| min_vram_gb <= vram);

```

This early VRAM check feeds directly into the run-mode selection that follows.

## Run-Mode Selection: The VRAM-First Decision Tree

The definitive prioritization occurs in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs). The execution path selection follows a strict VRAM-first ordering:

1. **`RunMode::Gpu`** — Selected when `system.has_gpu` is true and `min_vram_gb <= system.gpu_vram_gb`. The entire model resides in VRAM; this is the optimal path.

2. **`RunMode::MoeOffload`** — For MoE (Mixture-of-Experts) models where active experts fit in VRAM. The GPU accelerates inference while inactive experts remain in system RAM.

3. **`RunMode::CpuOffload`** — When the model exceeds VRAM but GPU acceleration remains partially viable. Weights spill to system RAM as needed.

4. **`RunMode::CpuOnly`** — Fallback when no GPU is present or the model cannot leverage VRAM.

This branch ordering guarantees VRAM utilization is exhausted before any CPU-only path is considered. The logic explicitly prefers partial GPU utilization (`CpuOffload`) over abandoning the GPU entirely.

## Fit Level Scoring: Reflecting VRAM Priority in Recommendations

The `FitLevel` enum in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) codifies the quality of hardware-model matching:

- `FitLevel::Perfect` — "Recommended memory met on GPU"
- `FitLevel::Good` — Fits with minor constraints
- `FitLevel::Fair` — Functional but suboptimal
- `FitLevel::Poor` — Not recommended

Crucially, `Perfect` requires GPU residency. The `fit_level` field derives from both `run_mode` and the memory utilization ratio (`memory_required_gb / memory_available_gb`). Because `RunMode::Gpu` is evaluated first in the selection logic, models fitting in VRAM automatically receive the highest fit level—regardless of abundant system RAM.

```rust
// From fit.rs - fit_level determination
let fit_level = match run_mode {
    RunMode::Gpu if utilization <= 0.8 => FitLevel::Perfect,
    RunMode::Gpu => FitLevel::Good,
    RunMode::MoeOffload => FitLevel::Good,
    RunMode::CpuOffload => FitLevel::Fair,
    RunMode::CpuOnly => FitLevel::Poor,
};

```

## User-Facing Output: Surfacing VRAM Decisions

Both interfaces make the VRAM-first policy visible. In [`llmfit-tui/src/display.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs), the results table includes a "Fits in VRAM" column computed from `fit.run_mode == RunMode::Gpu`:

```rust
// From display.rs rendering logic
let vram_badge = if fit.run_mode == RunMode::Gpu {
    "✅"
} else {
    "❌"
};

```

The CLI ([`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs)) similarly highlights GPU execution paths in its formatted output, reinforcing that VRAM fit status is the primary compatibility indicator.

## Practical Verification

Query your system's model compatibility to see VRAM prioritization in action:

```bash
llmfit fit

```

Example output demonstrating the decision logic:

| Model | Min VRAM (GB) | GPU VRAM (GB) | Run Mode | Fits in VRAM |
|-------|---------------|---------------|----------|--------------|
| Qwen2-7B-Q4_K_M | 4.2 | 12.0 | Gpu | ✅ |
| Qwen2-72B-Q4_K_M | 39.0 | 12.0 | CpuOffload | ❌ |
| DeepSeek-V2-Lite | 5.8 | 12.0 | MoeOffload | Partial |

Programmatically verify the selected path:

```rust
use llmfit_core::{hardware::SystemSpecs, fit::ModelFit, models::LlmModel};

let system = SystemSpecs::detect();
let model = LlmModel::from_name("Qwen2-7B-Q4_K_M")?;
let fit = ModelFit::analyze(&model, &system);

match fit.run_mode {
    RunMode::Gpu => println!("VRAM-only execution: {:.1} GB", fit.memory_required_gb),
    RunMode::CpuOffload => println!("Spilling to system RAM"),
    _ => println!("Alternative execution path"),
}

```

## Summary

- **Hardware detection** in [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) captures GPU VRAM as a distinct capability from system RAM
- **Memory analysis** in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) computes GPU requirements before considering RAM alternatives
- **Run-mode selection** in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) follows a strict VRAM-first hierarchy: Gpu → MoeOffload → CpuOffload → CpuOnly
- **Fit scoring** reserves `Perfect` ratings exclusively for GPU-resident models
- **User interfaces** surface VRAM fit status as the primary compatibility metric

## Frequently Asked Questions

### How does llmfit handle systems with multiple GPUs?

The current `SystemSpecs` aggregates VRAM from all detected GPUs into `gpu_vram_gb` as a single value. Multi-GPU configurations are treated as a unified memory pool for fit determination. Per-GPU granularity is not exposed in the public API as of the analyzed version.

### What happens when a model barely exceeds available VRAM?

`llmfit` selects `RunMode::CpuOffload`, keeping as many weights as possible in VRAM while spilling the remainder to system RAM. This preserves partial GPU acceleration rather than falling back to `CpuOnly`. The threshold is strict: any VRAM overrun triggers offload, with no attempt at aggressive memory compression.

### Can llmfit be forced to ignore GPU VRAM and use system RAM instead?

No direct flag exists in the analyzed codebase. The `RunMode` selection is deterministic based on hardware detection and model requirements. Users would need to artificially modify `SystemSpecs::detect()` output or use environment variables to hide GPU availability from the detection logic.

### Does llmfit consider quantization effects on VRAM requirements?

Yes. The `min_vram_gb` calculation in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) incorporates quantization level (Q4, Q5, Q8, etc.) when determining memory needs. Lower-precision quantizations reduce the VRAM threshold, enabling larger models to qualify for `RunMode::Gpu` that would otherwise require offload.