# How LLMFIT plan.rs Calculates Minimum and Recommended VRAM and RAM Requirements

> Discover how LLMFIT plan.rs calculates VRAM and RAM needs for GPU, CPU offload, and CPU-only execution. Learn memory thresholds based on model size and context.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-08-21

---

**LLMFIT determines hardware requirements by evaluating three execution paths—GPU, CPU offload, and CPU-only—using the `build_path_estimate` function in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) to compute minimum and recommended memory thresholds based on model size, quantization, and context length.**

The hardware estimation logic in [AlexsJones/llmfit](https://github.com/AlexsJones/llmfit) centers on [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs), which orchestrates the planning feature. This module calculates precise VRAM and RAM requirements by analyzing the model memory footprint and applying path-specific multipliers to ensure stable inference across different hardware configurations.

## The Three Execution Paths

[`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) evaluates every model against three distinct execution strategies defined in `build_path_estimate` (lines 7028-7044 and 7062-7102). Each path uses different memory pools:

- **GPU**: Uses VRAM as the primary memory pool with minimal system RAM for overhead
- **CPU offload**: Uses system RAM as the primary pool with a small VRAM spill buffer  
- **CPU-only**: Uses system RAM exclusively with no GPU memory requirements

The planner invokes `model.estimate_memory_gb_with_kv(quant, context, kv_quant)` to establish the base **model memory footprint**—the total bytes required for weights, KV cache, and overhead—before applying path-specific calculations.

## GPU Path Calculations

For the GPU execution path, `build_path_estimate` computes requirements assuming the model resides entirely in VRAM:

**Minimum VRAM** equals the full model memory footprint:

```rust
min_vram = model_mem

```

**Recommended VRAM** takes the larger of the model's catalog recommendation or 20% headroom:

```rust
rec_vram = model.recommended_ram_gb.max(model_mem * 1.2);

```

**Minimum RAM** reserves headroom for KV cache and system overhead:

```rust
min_ram = (model_mem * 0.2).max(8.0)  // GB

```

**Recommended RAM** adds a 25% buffer with a 12 GB floor:

```rust
rec_ram = (min_ram * 1.25).max(12.0)  // GB

```

These values populate a `HardwareEstimate` struct (lines 7025-7041) that the planner later grades against host resources.

## CPU Offload Path Calculations

The CPU offload path activates when the machine lacks unified memory. This configuration uses system RAM as the decisive resource:

**VRAM requirements** shrink to fixed spill buffers:

```rust
min_vram = 2.0  // GB
rec_vram = 4.0  // GB

```

**RAM requirements** absorb the full model weight:

```rust
min_ram = model_mem
rec_ram = model_mem * 1.2

```

In this mode, VRAM functions merely as an auxiliary buffer while the RAM pool determines feasibility.

## CPU-Only Path Calculations

For CPU-only execution, `build_path_estimate` eliminates VRAM requirements entirely:

```rust
vram_gb: None
min_ram = model_mem  
rec_ram = model_mem * 1.2

```

This path provides the lowest hardware barrier but caps performance potential, always returning a **Marginal** fit level regardless of available RAM.

## Grading Hardware Fit with fit_level_for

After calculating requirements, [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) grades feasibility using `fit_level_for` (lines 3000-3028). This function compares available resources against calculated minimums and recommendations:

- **TooTight**: Returned when required memory exceeds available capacity
- **Perfect**: Available resources meet or exceed recommended thresholds (GPU path only)
- **Good**: Available resources exceed minimum requirements by at least 20%
- **Marginal**: Resources meet minimums but lack the 20% safety margin

For CPU-only paths, the heuristic caps the grade at **Marginal** regardless of surplus RAM. When resources prove insufficient, the planner generates `UpgradeDelta` suggestions indicating specific GB additions needed for the target resource.

## Practical Implementation Examples

### Generating a Complete Hardware Plan

```rust
use llmfit_core::plan::{estimate_model_plan, PlanRequest};
use llmfit_core::hardware::{SystemSpecs, GpuBackend};
use llmfit_core::models::LlmModel;

let system = SystemSpecs {
    total_ram_gb: 16.0,
    available_ram_gb: 12.0,
    total_cpu_cores: 8,
    cpu_name: "Intel i7".into(),
    has_gpu: true,
    gpu_vram_gb: Some(6.0),
    total_gpu_vram_gb: Some(6.0),
    gpu_name: Some("GeForce GTX 1650".into()),
    unified_memory: false,
    backend: GpuBackend::Cuda,
    ..Default::default()
};

let request = PlanRequest {
    context: 4096,
    quant: None,
    target_tps: None,
    kv_quant: None,
};

let model = LlmModel::load_from_name("Qwen-Test-7B")?;
let plan = estimate_model_plan(&model, &request, &system)?;

// Access GPU path estimates (index 0 typically)
if let Some(gpu) = &plan.run_paths[0].minimum {
    println!("GPU min VRAM: {:.1} GB", gpu.vram_gb.unwrap_or(0.0));
    println!("GPU min RAM: {:.1} GB", gpu.ram_gb);
}

```

### Inspecting Upgrade Recommendations

```rust
for delta in plan.upgrade_deltas {
    if let (Some(gb), Some(resource)) = (delta.add_gb, &delta.resource) {
        println!("Add {:.1}GB {} → {}", gb, resource, delta.description);
    }
}

```

### Evaluating CPU-Only Feasibility

```rust
let request = PlanRequest { 
    context: 8192, 
    quant: None, 
    target_tps: None, 
    kv_quant: None 
};

let plan = estimate_model_plan(&model, &request, &system)?;
let cpu_path = plan.run_paths
    .iter()
    .find(|p| p.path == PlanRunPath::CpuOnly)
    .expect("CPU path not found");

if let Some(est) = &cpu_path.minimum {
    println!("CPU-only requires {:.1} GB RAM", est.ram_gb);
}

```

## Summary

- [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) calculates hardware requirements through `build_path_estimate`, evaluating GPU, CPU offload, and CPU-only execution paths.
- **GPU path** requires full model memory in VRAM (`min_vram = model_mem`) with recommended VRAM adding 20% headroom or using catalog values.
- **CPU offload** fixes VRAM at 2-4 GB spill buffers while requiring full model memory in RAM plus 20% buffer.
- **CPU-only** eliminates VRAM requirements entirely, using only system RAM at 1.0x minimum and 1.2x recommended multipliers.
- `fit_level_for` (lines 3000-3028) grades feasibility as Perfect, Good, Marginal, or TooTight based on 20% safety margins.
- All calculations derive from `model.estimate_memory_gb_with_kv`, which accounts for quantization, context length, and KV cache quantization.

## Frequently Asked Questions

### What determines whether LLMFIT chooses GPU or CPU offload paths?

The planner automatically evaluates all three paths but selects execution strategies based on `SystemSpecs` detection in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs). If `has_gpu` is true and `gpu_vram_gb` meets minimum thresholds, the GPU path receives priority grading. CPU offload activates when the GPU lacks sufficient VRAM but the system has adequate RAM, provided `unified_memory` is false.

### Why does the GPU path calculate minimum RAM as 20% of model memory?

The 20% RAM allocation (`model_mem * 0.2`) reserves space for the KV cache, activation buffers, and system overhead that cannot reside in VRAM during inference. The `.max(8.0)` floor ensures laptops and edge devices maintain minimum operational headroom regardless of model size, as implemented in `build_path_estimate` lines 7028-7044.

### How does recommended VRAM differ from minimum VRAM in the calculations?

Minimum VRAM represents the absolute floor where the model fits without crashing (`model_mem`), while recommended VRAM incorporates performance headroom. The code takes the larger value between the model's catalog recommendation (`recommended_ram_gb`) or 20% overhead (`model_mem * 1.2`), ensuring users receive guidance that accounts for both empirical testing data and safety margins.

### What happens when available memory falls between minimum and recommended values?

When resources exceed minimums but fail to meet recommendations, `fit_level_for` assigns a **Marginal** or **Good** grade depending on the 1.2x threshold. For GPU paths, achieving at least 120% of minimum requirements yields **Good**; otherwise **Marginal**. CPU-only paths cap at **Marginal** regardless of surplus memory, reflecting the inherent performance limitations of CPU inference.