Why llmfit Multiplies RAM by 1.2 and VRAM by 1.1 in scripts/scrape_hf_models.py

The llmfit project applies a 1.2× multiplier to raw model size for RAM estimates to cover KV-cache, activation buffers, and operating system overhead on CPU, while using a tighter 1.1× multiplier for VRAM because GPU inference benefits from unified memory architectures and requires only a 10% buffer for temporary tensors.

The llmfit repository calculates hardware requirements for large language models by deriving memory consumption from parameter counts and quantization bit-rates. In scripts/scrape_hf_models.py, the estimate_ram and estimate_vram functions convert raw parameter counts into gigabyte estimates using quantization-aware bytes-per-parameter values, then scale the results by 1.2 for CPU RAM and 1.1 for GPU VRAM to reflect realistic runtime conditions.

The Memory Estimation Functions

The core logic resides in scripts/scrape_hf_models.py, where two distinct functions handle CPU and GPU memory projections.

estimate_ram (Lines 44-52)

The estimate_ram function calculates both minimum and recommended RAM requirements:

def estimate_ram(total_params: int, quant: str) -> tuple[float, float]:
    bpp = QUANT_BPP.get(quant, 0.5)                     # bytes-per-parameter

    model_size_gb = (total_params * bpp) / (1024**3)    # raw size in GiB

    min_ram_gb = model_size_gb * RUNTIME_OVERHEAD        # 1.2 × size

    recommended_ram_gb = model_size_gb * 2.0             # generous headroom

    return round(min_ram_gb, 1), round(recommended_ram_gb, 1)

This implementation uses the RUNTIME_OVERHEAD constant defined at line 12 as 1.2, yielding approximately 20% extra capacity beyond the raw model weights.

estimate_vram (Lines 62-66)

For GPU inference, estimate_vram applies a leaner calculation:

def estimate_vram(total_params: int, quant: str) -> float:
    bpp = QUANT_BPP.get(quant, 0.5)
    model_size_gb = (total_params * bpp) / (1024**3)
    vram_gb = model_size_gb * 1.1                        # 10% overhead

    return round(max(vram_gb, 0.5), 1)

The 1.1 multiplier here represents a 10% safety margin rather than the 20% allocated for CPU runtimes.

Why RAM Uses a 1.2× Multiplier

The 1.2 factor (defined as RUNTIME_OVERHEAD = 1.2 at line 12) accounts for three primary sources of memory consumption beyond static model weights:

  • KV-cache: Transformer inference requires storing key-value pairs for attention mechanisms, which scales with sequence length and batch size.
  • Activation buffers: Intermediate computational tensors allocated during forward passes.
  • Operating system overhead: Memory reserved by the OS for process management, dynamic libraries, and paging.

Multiplying the raw model size by 1.2 produces a realistic lower bound for CPU inference, ensuring the system does not immediately exhaust available RAM upon loading the model.

Why VRAM Uses a 1.1× Multiplier

GPU memory estimation relies on a smaller 1.1 multiplier because:

  • Unified memory architecture: Many modern GPU platforms manage memory more efficiently than CPU-based systems, reducing duplication.
  • Kernel optimization: GPU inference engines typically pre-allocate memory pools and reuse activation tensors more aggressively.
  • Minimal OS interference: The 10% buffer covers only temporary activation spikes and minor driver overhead without the significant OS background consumption seen in CPU environments.

This tighter estimate prevents over-provisioning while maintaining a safety margin against out-of-memory errors during generation.

Quantization-Aware Size Calculation

Both formulas derive from QUANT_BPP, a dictionary mapping quantization schemes to bytes-per-parameter (e.g., Q4_K_M → 0.5). By applying the multiplier to the quantization-compressed size rather than the full-precision parameter count, the estimates reflect actual disk-to-memory footprints.


# Example mapping from QUANT_BPP

QUANT_BPP = {
    "Q4_K_M": 0.5,
    "Q5_K_M": 0.625,
    "Q8_0": 1.0,
    # ...

}

Practical Usage Examples

Estimate memory for a 7-billion-parameter model using Q4_K_M quantization:

from scripts.scrape_hf_models import estimate_ram, estimate_vram

params = 7_000_000_000
quant = "Q4_K_M"  # 0.5 bytes per parameter

min_ram, rec_ram = estimate_ram(params, quant)
min_vram = estimate_vram(params, quant)

print(f"CPU RAM:  {min_ram} GB (min) / {rec_ram} GB (recommended)")
print(f"GPU VRAM: {min_vram} GB")

# Output:

# CPU RAM:  3.5 GB (min) / 7.0 GB (recommended)

# GPU VRAM: 3.9 GB

Run the full scraping pipeline to generate JSON with memory estimates:

python3 scripts/scrape_hf_models.py -n 10

This outputs fields including min_ram_gb, recommended_ram_gb, and min_vram_gb for each model processed.

Integration with llmfit-core

The Rust codebase in llmfit-core/src/models.rs consumes the JSON generated by scrape_hf_models.py and applies identical memory calculations when validating hardware compatibility. This ensures consistent estimations across the Python scraping utilities and the core Rust inference engine.

Summary

  • estimate_ram multiplies raw model size by 1.2 (RUNTIME_OVERHEAD) to cover KV-cache, activations, and OS overhead on CPU.
  • estimate_vram uses 1.1 to provide a 10% safety margin for GPU activation tensors and driver overhead.
  • Both calculations start from quantization-aware sizes via QUANT_BPP to ensure accuracy across formats like Q4_K_M and Q8_0.
  • Recommended RAM doubles the raw size (2.0×) to prevent OOM errors during long sequences.
  • The formulas align Python scraping logic with the Rust core in llmfit-core/src/models.rs.

Frequently Asked Questions

Why is the VRAM multiplier lower than the RAM multiplier?

GPU inference benefits from unified memory architectures and more efficient kernel memory management, requiring only a 10% buffer for activations versus the 20% needed for CPU runtime overhead, KV-cache, and operating system consumption.

Where is the 1.2 constant defined in the codebase?

The RUNTIME_OVERHEAD = 1.2 constant is defined at line 12 of scripts/scrape_hf_models.py and applied in the estimate_ram function between lines 44 and 52.

Does the 1.2× RAM estimate include the KV-cache?

Yes, the 20% overhead factor empirically covers KV-cache growth, activation buffers, and system overhead observed during CPU inference of quantized models.

How does quantization affect these calculations?

Both functions look up bytes-per-parameter in the QUANT_BPP dictionary (e.g., 0.5 for Q4_K_M) before applying the multipliers, ensuring the 1.2× and 1.1× factors scale against the compressed model size rather than full-precision parameters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →