How LLMFIT plan.rs Calculates Minimum and Recommended VRAM and RAM Requirements

LLMFIT determines hardware requirements by evaluating three execution paths—GPU, CPU offload, and CPU-only—using the build_path_estimate function in plan.rs to compute minimum and recommended memory thresholds based on model size, quantization, and context length.

The hardware estimation logic in AlexsJones/llmfit centers on llmfit-core/src/plan.rs, which orchestrates the planning feature. This module calculates precise VRAM and RAM requirements by analyzing the model memory footprint and applying path-specific multipliers to ensure stable inference across different hardware configurations.

The Three Execution Paths

plan.rs evaluates every model against three distinct execution strategies defined in build_path_estimate (lines 7028-7044 and 7062-7102). Each path uses different memory pools:

  • GPU: Uses VRAM as the primary memory pool with minimal system RAM for overhead
  • CPU offload: Uses system RAM as the primary pool with a small VRAM spill buffer
  • CPU-only: Uses system RAM exclusively with no GPU memory requirements

The planner invokes model.estimate_memory_gb_with_kv(quant, context, kv_quant) to establish the base model memory footprint—the total bytes required for weights, KV cache, and overhead—before applying path-specific calculations.

GPU Path Calculations

For the GPU execution path, build_path_estimate computes requirements assuming the model resides entirely in VRAM:

Minimum VRAM equals the full model memory footprint:

min_vram = model_mem

Recommended VRAM takes the larger of the model's catalog recommendation or 20% headroom:

rec_vram = model.recommended_ram_gb.max(model_mem * 1.2);

Minimum RAM reserves headroom for KV cache and system overhead:

min_ram = (model_mem * 0.2).max(8.0)  // GB

Recommended RAM adds a 25% buffer with a 12 GB floor:

rec_ram = (min_ram * 1.25).max(12.0)  // GB

These values populate a HardwareEstimate struct (lines 7025-7041) that the planner later grades against host resources.

CPU Offload Path Calculations

The CPU offload path activates when the machine lacks unified memory. This configuration uses system RAM as the decisive resource:

VRAM requirements shrink to fixed spill buffers:

min_vram = 2.0  // GB
rec_vram = 4.0  // GB

RAM requirements absorb the full model weight:

min_ram = model_mem
rec_ram = model_mem * 1.2

In this mode, VRAM functions merely as an auxiliary buffer while the RAM pool determines feasibility.

CPU-Only Path Calculations

For CPU-only execution, build_path_estimate eliminates VRAM requirements entirely:

vram_gb: None
min_ram = model_mem  
rec_ram = model_mem * 1.2

This path provides the lowest hardware barrier but caps performance potential, always returning a Marginal fit level regardless of available RAM.

Grading Hardware Fit with fit_level_for

After calculating requirements, plan.rs grades feasibility using fit_level_for (lines 3000-3028). This function compares available resources against calculated minimums and recommendations:

  • TooTight: Returned when required memory exceeds available capacity
  • Perfect: Available resources meet or exceed recommended thresholds (GPU path only)
  • Good: Available resources exceed minimum requirements by at least 20%
  • Marginal: Resources meet minimums but lack the 20% safety margin

For CPU-only paths, the heuristic caps the grade at Marginal regardless of surplus RAM. When resources prove insufficient, the planner generates UpgradeDelta suggestions indicating specific GB additions needed for the target resource.

Practical Implementation Examples

Generating a Complete Hardware Plan

use llmfit_core::plan::{estimate_model_plan, PlanRequest};
use llmfit_core::hardware::{SystemSpecs, GpuBackend};
use llmfit_core::models::LlmModel;

let system = SystemSpecs {
    total_ram_gb: 16.0,
    available_ram_gb: 12.0,
    total_cpu_cores: 8,
    cpu_name: "Intel i7".into(),
    has_gpu: true,
    gpu_vram_gb: Some(6.0),
    total_gpu_vram_gb: Some(6.0),
    gpu_name: Some("GeForce GTX 1650".into()),
    unified_memory: false,
    backend: GpuBackend::Cuda,
    ..Default::default()
};

let request = PlanRequest {
    context: 4096,
    quant: None,
    target_tps: None,
    kv_quant: None,
};

let model = LlmModel::load_from_name("Qwen-Test-7B")?;
let plan = estimate_model_plan(&model, &request, &system)?;

// Access GPU path estimates (index 0 typically)
if let Some(gpu) = &plan.run_paths[0].minimum {
    println!("GPU min VRAM: {:.1} GB", gpu.vram_gb.unwrap_or(0.0));
    println!("GPU min RAM: {:.1} GB", gpu.ram_gb);
}

Inspecting Upgrade Recommendations

for delta in plan.upgrade_deltas {
    if let (Some(gb), Some(resource)) = (delta.add_gb, &delta.resource) {
        println!("Add {:.1}GB {} → {}", gb, resource, delta.description);
    }
}

Evaluating CPU-Only Feasibility

let request = PlanRequest { 
    context: 8192, 
    quant: None, 
    target_tps: None, 
    kv_quant: None 
};

let plan = estimate_model_plan(&model, &request, &system)?;
let cpu_path = plan.run_paths
    .iter()
    .find(|p| p.path == PlanRunPath::CpuOnly)
    .expect("CPU path not found");

if let Some(est) = &cpu_path.minimum {
    println!("CPU-only requires {:.1} GB RAM", est.ram_gb);
}

Summary

  • plan.rs calculates hardware requirements through build_path_estimate, evaluating GPU, CPU offload, and CPU-only execution paths.
  • GPU path requires full model memory in VRAM (min_vram = model_mem) with recommended VRAM adding 20% headroom or using catalog values.
  • CPU offload fixes VRAM at 2-4 GB spill buffers while requiring full model memory in RAM plus 20% buffer.
  • CPU-only eliminates VRAM requirements entirely, using only system RAM at 1.0x minimum and 1.2x recommended multipliers.
  • fit_level_for (lines 3000-3028) grades feasibility as Perfect, Good, Marginal, or TooTight based on 20% safety margins.
  • All calculations derive from model.estimate_memory_gb_with_kv, which accounts for quantization, context length, and KV cache quantization.

Frequently Asked Questions

What determines whether LLMFIT chooses GPU or CPU offload paths?

The planner automatically evaluates all three paths but selects execution strategies based on SystemSpecs detection in llmfit-core/src/hardware.rs. If has_gpu is true and gpu_vram_gb meets minimum thresholds, the GPU path receives priority grading. CPU offload activates when the GPU lacks sufficient VRAM but the system has adequate RAM, provided unified_memory is false.

Why does the GPU path calculate minimum RAM as 20% of model memory?

The 20% RAM allocation (model_mem * 0.2) reserves space for the KV cache, activation buffers, and system overhead that cannot reside in VRAM during inference. The .max(8.0) floor ensures laptops and edge devices maintain minimum operational headroom regardless of model size, as implemented in build_path_estimate lines 7028-7044.

Minimum VRAM represents the absolute floor where the model fits without crashing (model_mem), while recommended VRAM incorporates performance headroom. The code takes the larger value between the model's catalog recommendation (recommended_ram_gb) or 20% overhead (model_mem * 1.2), ensuring users receive guidance that accounts for both empirical testing data and safety margins.

When resources exceed minimums but fail to meet recommendations, fit_level_for assigns a Marginal or Good grade depending on the 1.2x threshold. For GPU paths, achieving at least 120% of minimum requirements yields Good; otherwise Marginal. CPU-only paths cap at Marginal regardless of surplus memory, reflecting the inherent performance limitations of CPU inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →