# Why llmfit Multiplies RAM by 1.2 and VRAM by 1.1 in scripts/scrape_hf_models.py

> Discover why llmfit uses 1.2x RAM and 1.1x VRAM multipliers in scrape_hf_models.py to accurately estimate resource needs for CPU and GPU inference.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-12

---

**The `llmfit` project applies a 1.2× multiplier to raw model size for RAM estimates to cover KV-cache, activation buffers, and operating system overhead on CPU, while using a tighter 1.1× multiplier for VRAM because GPU inference benefits from unified memory architectures and requires only a 10% buffer for temporary tensors.**

The `llmfit` repository calculates hardware requirements for large language models by deriving memory consumption from parameter counts and quantization bit-rates. In [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py), the `estimate_ram` and `estimate_vram` functions convert raw parameter counts into gigabyte estimates using quantization-aware bytes-per-parameter values, then scale the results by **1.2** for CPU RAM and **1.1** for GPU VRAM to reflect realistic runtime conditions.

## The Memory Estimation Functions

The core logic resides in [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py), where two distinct functions handle CPU and GPU memory projections.

### estimate_ram (Lines 44-52)

The `estimate_ram` function calculates both minimum and recommended RAM requirements:

```python
def estimate_ram(total_params: int, quant: str) -> tuple[float, float]:
    bpp = QUANT_BPP.get(quant, 0.5)                     # bytes-per-parameter

    model_size_gb = (total_params * bpp) / (1024**3)    # raw size in GiB

    min_ram_gb = model_size_gb * RUNTIME_OVERHEAD        # 1.2 × size

    recommended_ram_gb = model_size_gb * 2.0             # generous headroom

    return round(min_ram_gb, 1), round(recommended_ram_gb, 1)

```

This implementation uses the **`RUNTIME_OVERHEAD`** constant defined at line 12 as `1.2`, yielding approximately 20% extra capacity beyond the raw model weights.

### estimate_vram (Lines 62-66)

For GPU inference, `estimate_vram` applies a leaner calculation:

```python
def estimate_vram(total_params: int, quant: str) -> float:
    bpp = QUANT_BPP.get(quant, 0.5)
    model_size_gb = (total_params * bpp) / (1024**3)
    vram_gb = model_size_gb * 1.1                        # 10% overhead

    return round(max(vram_gb, 0.5), 1)

```

The 1.1 multiplier here represents a 10% safety margin rather than the 20% allocated for CPU runtimes.

## Why RAM Uses a 1.2× Multiplier

The **1.2** factor (defined as `RUNTIME_OVERHEAD = 1.2` at line 12) accounts for three primary sources of memory consumption beyond static model weights:

- **KV-cache**: Transformer inference requires storing key-value pairs for attention mechanisms, which scales with sequence length and batch size.
- **Activation buffers**: Intermediate computational tensors allocated during forward passes.
- **Operating system overhead**: Memory reserved by the OS for process management, dynamic libraries, and paging.

Multiplying the raw model size by 1.2 produces a realistic lower bound for CPU inference, ensuring the system does not immediately exhaust available RAM upon loading the model.

## Why VRAM Uses a 1.1× Multiplier

GPU memory estimation relies on a smaller **1.1** multiplier because:

- **Unified memory architecture**: Many modern GPU platforms manage memory more efficiently than CPU-based systems, reducing duplication.
- **Kernel optimization**: GPU inference engines typically pre-allocate memory pools and reuse activation tensors more aggressively.
- **Minimal OS interference**: The 10% buffer covers only temporary activation spikes and minor driver overhead without the significant OS background consumption seen in CPU environments.

This tighter estimate prevents over-provisioning while maintaining a safety margin against out-of-memory errors during generation.

## Quantization-Aware Size Calculation

Both formulas derive from **`QUANT_BPP`**, a dictionary mapping quantization schemes to bytes-per-parameter (e.g., `Q4_K_M` → `0.5`). By applying the multiplier to the quantization-compressed size rather than the full-precision parameter count, the estimates reflect actual disk-to-memory footprints.

```python

# Example mapping from QUANT_BPP

QUANT_BPP = {
    "Q4_K_M": 0.5,
    "Q5_K_M": 0.625,
    "Q8_0": 1.0,
    # ...

}

```

## Practical Usage Examples

Estimate memory for a 7-billion-parameter model using Q4_K_M quantization:

```python
from scripts.scrape_hf_models import estimate_ram, estimate_vram

params = 7_000_000_000
quant = "Q4_K_M"  # 0.5 bytes per parameter

min_ram, rec_ram = estimate_ram(params, quant)
min_vram = estimate_vram(params, quant)

print(f"CPU RAM:  {min_ram} GB (min) / {rec_ram} GB (recommended)")
print(f"GPU VRAM: {min_vram} GB")

# Output:

# CPU RAM:  3.5 GB (min) / 7.0 GB (recommended)

# GPU VRAM: 3.9 GB

```

Run the full scraping pipeline to generate JSON with memory estimates:

```bash
python3 scripts/scrape_hf_models.py -n 10

```

This outputs fields including `min_ram_gb`, `recommended_ram_gb`, and `min_vram_gb` for each model processed.

## Integration with llmfit-core

The Rust codebase in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) consumes the JSON generated by [`scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scrape_hf_models.py) and applies identical memory calculations when validating hardware compatibility. This ensures consistent estimations across the Python scraping utilities and the core Rust inference engine.

## Summary

- **`estimate_ram`** multiplies raw model size by **1.2** (`RUNTIME_OVERHEAD`) to cover KV-cache, activations, and OS overhead on CPU.
- **`estimate_vram`** uses **1.1** to provide a 10% safety margin for GPU activation tensors and driver overhead.
- Both calculations start from quantization-aware sizes via **`QUANT_BPP`** to ensure accuracy across formats like Q4_K_M and Q8_0.
- **Recommended RAM** doubles the raw size (2.0×) to prevent OOM errors during long sequences.
- The formulas align Python scraping logic with the Rust core in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs).

## Frequently Asked Questions

### Why is the VRAM multiplier lower than the RAM multiplier?

GPU inference benefits from unified memory architectures and more efficient kernel memory management, requiring only a 10% buffer for activations versus the 20% needed for CPU runtime overhead, KV-cache, and operating system consumption.

### Where is the 1.2 constant defined in the codebase?

The `RUNTIME_OVERHEAD = 1.2` constant is defined at line 12 of [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py) and applied in the `estimate_ram` function between lines 44 and 52.

### Does the 1.2× RAM estimate include the KV-cache?

Yes, the 20% overhead factor empirically covers KV-cache growth, activation buffers, and system overhead observed during CPU inference of quantized models.

### How does quantization affect these calculations?

Both functions look up bytes-per-parameter in the `QUANT_BPP` dictionary (e.g., 0.5 for Q4_K_M) before applying the multipliers, ensuring the 1.2× and 1.1× factors scale against the compressed model size rather than full-precision parameters.