# The Role of scripts/scrape_hf_models.py in Maintaining the 33-Model Schema

> Discover how scripts scrape_hf_models.py maintains the 33-model schema for llmfit. This script is the central source for discovering, validating, and refreshing models for hardware-fit calculations.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-11

---

**The [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py) script serves as the single source of truth that discovers, validates, and continuously refreshes the 33-model catalog powering llmfit's hardware-fit calculations.**

In the llmfit repository, the [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py) file functions as the central engine responsible for building and maintaining the 33-model schema. This Python script generates the [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) catalog that embedded Rust binaries use to determine hardware compatibility for large language models.

## How the Scraper Builds the 33-Model Catalog

### Discovering Models from Curated Lists and Hugging Face

The script begins with a **curated list** defined in the `TARGET_MODELS` constant, then optionally downloads the top‑N most‑downloaded models from the Hugging Face API. This dual approach produces the set of repository IDs that populate the 33‑entry [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) file. By default, the scraper combines the curated list with the top 1000 models, though this can be constrained via command‑line arguments.

### Fetching Raw Metadata via the Hugging Face API

For each repository ID, the script invokes `fetch_model_info()` (lines 81‑87) to retrieve the full API record. It then extracts the **parameter count** from the `safetensors` metadata, which serves as the foundation for all downstream hardware calculations.

## Calculating Hardware Requirements for Each Model

### Estimating RAM and VRAM Consumption

Using the extracted parameter count and the model’s `default_quant` setting, the script derives precise hardware constraints:

- **`estimate_ram()`** (lines 443‑458) computes the **minimum RAM** and **recommended RAM** required to load and run the model.
- **`estimate_vram()`** (lines 61‑67) calculates the **minimum VRAM** necessary for GPU inference.

These estimates ensure the 33‑model schema accurately reflects real‑world deployment constraints.

## Enriching Schema Data and Preventing Information Loss

### Architecture Detection and MoE Identification

Beyond basic parameters, the script enriches entries with architectural metadata:

- **`extract_arch_metadata`** (lines 25‑91) parses configuration files to determine precise **KV‑cache formulas** required for memory prediction.
- **`detect_moe`** (lines 93‑134) identifies Mixture of Experts (MoE) architectures and estimates active parameter counts, which significantly affects memory consumption calculations.

### Merging Existing Values During Weekly Scrapes

To prevent data erosion during automated updates, **`preserve_existing_metadata()`** (lines 94‑123) copies non‑null values from the previous catalog version into the newly scraped entries. This ensures that manually verified fields—such as license information, language tags, and context‑length overrides—persist across weekly scrapes even if the Hugging Face API returns incomplete data.

### Integrating GGUF Quantization Sources

After the initial scrape, **`enrich_gguf_sources()`** (lines 90‑165) queries known GGUF providers for pre‑quantized versions of each model. The script caches these results to avoid unnecessary API calls, adding download URLs and quantization variants to the schema.

## Generating the Embedded JSON Output

The final dictionary structure (lines 54‑84) is written as a JSON array to [`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json). This file is embedded directly into the Rust binary via `include_str!`, making the 33‑model schema available at compile time. The catalog drives all downstream components: hardware‑fit analysis, CLI table rendering, TUI views, and the web dashboard. For convenience, the [`scripts/update_models.sh`](https://github.com/AlexsJones/llmfit/blob/main/scripts/update_models.sh) wrapper automates the full pipeline from scraping to binary rebuild.

## Running the Scraper to Update the Catalog

Execute the script from the repository root to regenerate the catalog:

1. **Basic execution** (curated list + top 1000):

```bash
python3 scripts/scrape_hf_models.py

```

2. **Limit to top 500 models only**:

```bash
python3 scripts/scrape_hf_models.py -n 500

```

3. **Accelerate with parallel threads**:

```bash
python3 scripts/scrape_hf_models.py --threads 8

```

4. **Rebuild the Rust binary** to embed the updated schema:

```bash
cargo build

```

## Summary

- **[`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py)** creates and maintains the 33‑model schema by combining curated lists with live Hugging Face API data.
- Key functions like `estimate_ram()`, `estimate_vram()`, and `extract_arch_metadata` derive hardware requirements and architectural details for accurate deployment predictions.
- **`preserve_existing_metadata()`** safeguards against data loss during weekly automated updates by merging previous catalog values.
- The script outputs to [`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json), which is embedded into the Rust binary and powers llmfit’s hardware compatibility engine.

## Frequently Asked Questions

### What file does scripts/scrape_hf_models.py generate?

The script generates [`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json), a JSON array containing the 33‑model schema. This file is embedded into the Rust binary via `include_str!` and contains hardware requirements, architectural metadata, and GGUF source links for each model.

### How does the script prevent data loss during weekly updates?

The **`preserve_existing_metadata()`** function (lines 94‑123) copies non‑null values from the previous catalog version into the new entries. This ensures that fields like license information, context‑length overrides, and language tags are retained even if the Hugging Face API returns incomplete records during automated scrapes.

### What hardware metrics are included in the 33-model schema?

The schema stores **minimum RAM**, **recommended RAM**, and **minimum VRAM** calculated by `estimate_ram()` (lines 443‑458) and `estimate_vram()` (lines 61‑67). These values are derived from parameter counts, default quantization settings, and architecture‑specific KV‑cache formulas.

### How do I update the embedded catalog in the Rust binary?

After running `python3 scripts/scrape_hf_models.py`, execute `cargo build` to recompile the binary with the updated [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) file. Alternatively, run the convenience wrapper [`scripts/update_models.sh`](https://github.com/AlexsJones/llmfit/blob/main/scripts/update_models.sh), which performs both the Python scrape and the Rust rebuild in sequence.