min_vram_gb vs min_ram_gb in llmfit: GPU VRAM vs CPU RAM Requirements Explained

min_vram_gb specifies the video RAM required for GPU inference while min_ram_gb specifies the system memory required for CPU-only inference, with both values calculated from the model's parameter count using different overhead multipliers.

The llmfit crate by AlexsJones/llmfit tracks two distinct hardware requirements for every model in its catalog to determine whether a system can run inference on GPU, CPU, or both. Understanding the difference between these fields is essential for selecting compatible models and avoiding out-of-memory errors during execution.

What min_vram_gb and min_ram_gb Represent

min_vram_gb (GPU Inference)

The min_vram_gb field indicates the minimum video RAM (VRAM) required to load and run a model on a GPU. This value represents the GPU-inference threshold, accounting for the model weights stored in graphics memory plus a 1.1× activation overhead factor. When llmfit evaluates the Gpu or CpuOffload execution paths, it compares the system's available VRAM against this figure.

min_ram_gb (CPU Inference)

The min_ram_gb field indicates the minimum system RAM required to load a model entirely into main memory for CPU-only inference. This value includes a higher 1.2× overhead factor to account for the full model storage in system memory plus CPU inference buffers. The CpuOnly execution path validates system RAM against this threshold.

How the Requirements Are Calculated

Both metrics derive from the model's parameter count using Q4_K_M quantization (0.5 bytes per parameter), but apply different overhead coefficients:

  • VRAM Formula: params × 0.5 bytes / 1024³ × 1.1 (activation overhead)
  • RAM Formula: params × 0.5 bytes / 1024³ × 1.2 (memory overhead)

According to the project documentation in AGENTS.md (lines 136–138), these formulas estimate practical memory consumption rather than raw model size, ensuring users account for runtime overhead during inference.

Implementation in the llmfit Codebase

The hardware requirements are defined across several key files:

  • llmfit-core/src/models.rs: Contains the ModelEntry struct that declares both fields as Option<f64>, allowing legacy entries to omit values while requiring them for new models.
  • scripts/scrape_hf_models.py: Python scraper that computes min_vram_gb and min_ram_gb from raw Hugging Face parameter counts using the formulas above.
  • llmfit-core/data/hf_models.json: Embedded JSON database storing the pre-calculated values for runtime lookup.
  • AGENTS.md: Documents the schema and calculation methodology for both fields.

Practical Usage Examples

Access these requirements programmatically through the ModelDatabase API:

use llmfit_core::models::{ModelDatabase, ModelEntry};

fn check_hardware_requirements(model_name: &str) {
    let db = ModelDatabase::new().expect("failed to load model DB");
    
    let entry: &ModelEntry = db
        .entries()
        .iter()
        .find(|e| e.name.eq_ignore_ascii_case(model_name))
        .expect("model not found");

    let vram = entry.min_vram_gb.unwrap_or_default();
    let ram = entry.min_ram_gb.unwrap_or_default();

    println!("Model: {}", entry.name);
    println!("  GPU VRAM required: {:.2} GB", vram);
    println!("  CPU RAM required:  {:.2} GB", ram);
}

For command-line inspection, query the JSON output:

cargo run -- --json fit --model "Llama-2-7b-Chat" | jq '.models[0] | {name, min_vram_gb, min_ram_gb}'

Execution Path Selection

llmfit uses these values to determine feasible execution strategies:

  • GPU Path: Checks available_vram >= min_vram_gb before attempting GPU inference or partial offloading.
  • CPU Path: Checks available_ram >= min_ram_gb before falling back to CPU-only execution.

On Apple Silicon systems with unified memory architecture, VRAM and system RAM share the same physical pool. The codebase marks this in SystemSpecs::unified_memory, making min_vram_gb and min_ram_gb effectively reference the same memory limit despite their semantic differences.

Summary

  • min_vram_gb targets GPU inference with 1.1× overhead for activation memory.
  • min_ram_gb targets CPU inference with 1.2× overhead for full model storage.
  • Both fields reside in ModelEntry structs defined in llmfit-core/src/models.rs.
  • The Python scraper in scripts/scrape_hf_models.py populates these values using the Q4_K_M quantization formula.
  • Apple Silicon devices treat both limits as unified memory constraints.

Frequently Asked Questions

Why does CPU inference require more memory than GPU inference?

CPU inference stores the entire dequantized model in system memory with a 1.2× overhead factor to accommodate activation buffers and operating system overhead. GPU inference uses a lower 1.1× overhead because the GPU manages weight storage more efficiently and shares memory architecture differently than CPU RAM.

What happens if my system meets min_ram_gb but not min_vram_gb?

llmfit automatically routes execution to the CPU-only path when VRAM is insufficient but system RAM meets the min_ram_gb threshold. The framework checks SystemSpecs against both values to determine the optimal execution strategy without manual intervention.

Are these values exact or conservative estimates?

The values are conservative estimates based on Q4_K_M quantization formulas documented in AGENTS.md. The 0.5 bytes per parameter calculation assumes 4-bit quantization, while the 1.1× and 1.2× multipliers provide safety margins for activation memory and system overhead during inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →