# How the llmfit Python Scraper Derives RAM and VRAM Requirements from Parameter Count

> Discover how the llmfit Python scraper calculates RAM and VRAM needs from model parameter counts. Learn about quantization formats, byte-per-parameter values, and overhead multipliers for accurate estimations.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-12

---

**The [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py) scraper estimates memory requirements by mapping quantization formats to bytes-per-parameter values, converting the result to GiB, then applying runtime-specific overhead multipliers of 1.2× for minimum RAM, 2.0× for recommended RAM, and 1.1× for minimum VRAM.**

The **AlexsJones/llmfit** repository provides hardware compatibility tooling for Large Language Models. Its Python scraper processes Hugging Face model metadata to derive memory requirements from raw parameter counts, enabling the core Rust engine in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) to determine if a given GPU or CPU configuration can host specific quantized models.

## Quantization-Aware Bytes-Per-Parameter Mapping

The scraper begins by determining how many bytes each parameter consumes based on the quantization format. In [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py), the `QUANT_BPP` dictionary (lines 94-103) maps quantization strings to their respective bytes-per-parameter values.

For the default **Q4_K_M** quantization used by the scraper, the mapping assigns **`0.5`** bytes per parameter. This value represents the compressed size of each weight after 4-bit quantization with K-means clustering and medium mixed precision.

## Converting Parameters to Raw Model Size

Once the bytes-per-parameter value is selected, the scraper calculates the raw model size in GiB. Both the `estimate_ram()` and `estimate_vram()` functions implement the same base conversion:

```python
model_size_gb = (total_params * bpp) / (1024**3)

```

This formula appears in `estimate_ram()` (lines 443-450) and `estimate_vram()` (lines 61-65), converting the total parameter count multiplied by the quantization-specific bytes-per-parameter into gibibytes.

## Runtime Overhead Calculations

After determining the raw model size, the scraper applies distinct scaling factors to account for runtime memory consumption patterns. Each estimate targets a specific hardware constraint and uses a different safety margin.

### Minimum RAM Estimation (min_ram_gb)

The **minimum RAM** calculation applies a 20% overhead to accommodate the KV-cache, activation layers, and operating system requirements. The `estimate_ram()` function (lines 50-51) multiplies the raw model size by **`RUNTIME_OVERHEAD = 1.2`**, then enforces a floor of 1 GiB (lines 55-57):

```python
min_ram_gb = max(1.0, round(model_size_gb * 1.2, 1))

```

### Recommended RAM Estimation (recommended_ram_gb)

For **recommended RAM**, the scraper assumes users require headroom for concurrent processes and larger context windows. The same `estimate_ram()` function (lines 52-53) applies a generous **2.0× multiplier** to the raw model size:

```python
recommended_ram_gb = round(model_size_gb * 2.0, 1)

```

This doubling of the base model size ensures comfortable operation without memory pressure.

### Minimum VRAM Estimation (min_vram_gb)

GPU estimates use a smaller overhead factor since VRAM primarily stores weights and a modest activation buffer. The `estimate_vram()` function (lines 66-67) multiplies the raw size by **1.1** and enforces a 0.5 GiB minimum:

```python
min_vram_gb = max(0.5, round(model_size_gb * 1.1, 1)

```

This 10% buffer accounts for CUDA context overhead and intermediate activations during inference.

## Practical Implementation Example

The following example demonstrates how the scraper processes a 7-billion-parameter model using the default Q4_K_M quantization:

```python
from scripts.scrape_hf_models import estimate_ram, estimate_vram

# 7B parameter model with Q4_K_M quantization

total_params = 7_000_000_000
quantization = "Q4_K_M"  # 0.5 bytes per parameter

# Calculate memory requirements

min_ram, rec_ram = estimate_ram(total_params, quantization)
min_vram = estimate_vram(total_params, quantization)

print(f"Min RAM: {min_ram} GB")
print(f"Recommended RAM: {rec_ram} GB") 
print(f"Min VRAM: {min_vram} GB")

```

Executing this code produces the values written to [`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json) for a 7B Q4 model: approximately 3.9 GB minimum RAM, 7.0 GB recommended RAM, and 3.5 GB minimum VRAM.

## Summary

- The scraper uses the `QUANT_BPP` mapping in [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py) (lines 94-103) to determine bytes-per-parameter based on quantization format.
- Raw model size in GiB is calculated as `(total_params * bpp) / (1024**3)` in both `estimate_ram()` and `estimate_vram()`.
- **Minimum RAM** applies a 1.2× multiplier (`RUNTIME_OVERHEAD`) with a 1 GiB floor to accommodate system overhead.
- **Recommended RAM** doubles (2.0×) the raw model size to ensure operational headroom.
- **Minimum VRAM** uses a 1.1× multiplier with a 0.5 GiB floor for GPU weight storage plus activation buffers.
- These values are stored in [`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json) and consumed by [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) for hardware compatibility checks.

## Frequently Asked Questions

### What quantization format does the scraper use by default?

The scraper defaults to **Q4_K_M** quantization, which assigns 0.5 bytes per parameter according to the `QUANT_BPP` dictionary in [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py). This format provides a balance between model compression and inference quality.

### Why is the recommended RAM twice the raw model size?

The 2.0× multiplier accounts for real-world usage scenarios including large context windows, concurrent system processes, and activation caching. This headroom prevents out-of-memory errors during peak usage, unlike the minimum RAM calculation which only covers basic runtime overhead.

### How does the scraper handle models with unusual parameter counts?

Both `estimate_ram()` and `estimate_vram()` enforce minimum floors of 1.0 GiB and 0.5 GiB respectively (lines 55-57 and 66-67). This ensures that even tiny models or edge cases report usable memory values rather than rounding down to zero.

### Where does the scraper store the derived memory values?

After processing, the scraper writes `min_ram_gb`, `recommended_ram_gb`, and `min_vram_gb` to [`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json). The Rust core in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) deserializes this data to perform hardware compatibility matching at runtime.