# min_vram_gb vs min_ram_gb in llmfit: GPU VRAM vs CPU RAM Requirements Explained

> Understand min_vram_gb vs min_ram_gb in llmfit. Learn the difference between GPU VRAM and CPU RAM requirements for LLM inference based on model size. Optimize your setup now.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-09-13

---

**`min_vram_gb` specifies the video RAM required for GPU inference while `min_ram_gb` specifies the system memory required for CPU-only inference, with both values calculated from the model's parameter count using different overhead multipliers.**

The `llmfit` crate by **AlexsJones/llmfit** tracks two distinct hardware requirements for every model in its catalog to determine whether a system can run inference on GPU, CPU, or both. Understanding the difference between these fields is essential for selecting compatible models and avoiding out-of-memory errors during execution.

## What min_vram_gb and min_ram_gb Represent

### min_vram_gb (GPU Inference)

The **`min_vram_gb`** field indicates the minimum video RAM (VRAM) required to load and run a model on a GPU. This value represents the **GPU-inference** threshold, accounting for the model weights stored in graphics memory plus a 1.1× activation overhead factor. When `llmfit` evaluates the `Gpu` or `CpuOffload` execution paths, it compares the system's available VRAM against this figure.

### min_ram_gb (CPU Inference)

The **`min_ram_gb`** field indicates the minimum system RAM required to load a model entirely into main memory for **CPU-only** inference. This value includes a higher 1.2× overhead factor to account for the full model storage in system memory plus CPU inference buffers. The `CpuOnly` execution path validates system RAM against this threshold.

## How the Requirements Are Calculated

Both metrics derive from the model's parameter count using Q4_K_M quantization (0.5 bytes per parameter), but apply different overhead coefficients:

- **VRAM Formula**: `params × 0.5 bytes / 1024³ × 1.1` (activation overhead)
- **RAM Formula**: `params × 0.5 bytes / 1024³ × 1.2` (memory overhead)

According to the project documentation in [`AGENTS.md`](https://github.com/AlexsJones/llmfit/blob/main/AGENTS.md) (lines 136–138), these formulas estimate practical memory consumption rather than raw model size, ensuring users account for runtime overhead during inference.

## Implementation in the llmfit Codebase

The hardware requirements are defined across several key files:

- **[`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs)**: Contains the `ModelEntry` struct that declares both fields as `Option<f64>`, allowing legacy entries to omit values while requiring them for new models.
- **[`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py)**: Python scraper that computes `min_vram_gb` and `min_ram_gb` from raw Hugging Face parameter counts using the formulas above.
- **[`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json)**: Embedded JSON database storing the pre-calculated values for runtime lookup.
- **[`AGENTS.md`](https://github.com/AlexsJones/llmfit/blob/main/AGENTS.md)**: Documents the schema and calculation methodology for both fields.

## Practical Usage Examples

Access these requirements programmatically through the `ModelDatabase` API:

```rust
use llmfit_core::models::{ModelDatabase, ModelEntry};

fn check_hardware_requirements(model_name: &str) {
    let db = ModelDatabase::new().expect("failed to load model DB");
    
    let entry: &ModelEntry = db
        .entries()
        .iter()
        .find(|e| e.name.eq_ignore_ascii_case(model_name))
        .expect("model not found");

    let vram = entry.min_vram_gb.unwrap_or_default();
    let ram = entry.min_ram_gb.unwrap_or_default();

    println!("Model: {}", entry.name);
    println!("  GPU VRAM required: {:.2} GB", vram);
    println!("  CPU RAM required:  {:.2} GB", ram);
}

```

For command-line inspection, query the JSON output:

```bash
cargo run -- --json fit --model "Llama-2-7b-Chat" | jq '.models[0] | {name, min_vram_gb, min_ram_gb}'

```

## Execution Path Selection

`llmfit` uses these values to determine feasible execution strategies:

- **GPU Path**: Checks `available_vram >= min_vram_gb` before attempting GPU inference or partial offloading.
- **CPU Path**: Checks `available_ram >= min_ram_gb` before falling back to CPU-only execution.

On Apple Silicon systems with unified memory architecture, VRAM and system RAM share the same physical pool. The codebase marks this in `SystemSpecs::unified_memory`, making `min_vram_gb` and `min_ram_gb` effectively reference the same memory limit despite their semantic differences.

## Summary

- **`min_vram_gb`** targets GPU inference with 1.1× overhead for activation memory.
- **`min_ram_gb`** targets CPU inference with 1.2× overhead for full model storage.
- Both fields reside in `ModelEntry` structs defined in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs).
- The Python scraper in [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py) populates these values using the Q4_K_M quantization formula.
- Apple Silicon devices treat both limits as unified memory constraints.

## Frequently Asked Questions

### Why does CPU inference require more memory than GPU inference?

CPU inference stores the entire dequantized model in system memory with a 1.2× overhead factor to accommodate activation buffers and operating system overhead. GPU inference uses a lower 1.1× overhead because the GPU manages weight storage more efficiently and shares memory architecture differently than CPU RAM.

### What happens if my system meets min_ram_gb but not min_vram_gb?

`llmfit` automatically routes execution to the **CPU-only** path when VRAM is insufficient but system RAM meets the `min_ram_gb` threshold. The framework checks `SystemSpecs` against both values to determine the optimal execution strategy without manual intervention.

### Are these values exact or conservative estimates?

The values are **conservative estimates** based on Q4_K_M quantization formulas documented in [`AGENTS.md`](https://github.com/AlexsJones/llmfit/blob/main/AGENTS.md). The 0.5 bytes per parameter calculation assumes 4-bit quantization, while the 1.1× and 1.2× multipliers provide safety margins for activation memory and system overhead during inference.