The Role of scripts/scrape_hf_models.py in Maintaining the 33-Model Schema

The scripts/scrape_hf_models.py script serves as the single source of truth that discovers, validates, and continuously refreshes the 33-model catalog powering llmfit's hardware-fit calculations.

In the llmfit repository, the scripts/scrape_hf_models.py file functions as the central engine responsible for building and maintaining the 33-model schema. This Python script generates the hf_models.json catalog that embedded Rust binaries use to determine hardware compatibility for large language models.

How the Scraper Builds the 33-Model Catalog

Discovering Models from Curated Lists and Hugging Face

The script begins with a curated list defined in the TARGET_MODELS constant, then optionally downloads the top‑N most‑downloaded models from the Hugging Face API. This dual approach produces the set of repository IDs that populate the 33‑entry hf_models.json file. By default, the scraper combines the curated list with the top 1000 models, though this can be constrained via command‑line arguments.

Fetching Raw Metadata via the Hugging Face API

For each repository ID, the script invokes fetch_model_info() (lines 81‑87) to retrieve the full API record. It then extracts the parameter count from the safetensors metadata, which serves as the foundation for all downstream hardware calculations.

Calculating Hardware Requirements for Each Model

Estimating RAM and VRAM Consumption

Using the extracted parameter count and the model’s default_quant setting, the script derives precise hardware constraints:

  • estimate_ram() (lines 443‑458) computes the minimum RAM and recommended RAM required to load and run the model.
  • estimate_vram() (lines 61‑67) calculates the minimum VRAM necessary for GPU inference.

These estimates ensure the 33‑model schema accurately reflects real‑world deployment constraints.

Enriching Schema Data and Preventing Information Loss

Architecture Detection and MoE Identification

Beyond basic parameters, the script enriches entries with architectural metadata:

  • extract_arch_metadata (lines 25‑91) parses configuration files to determine precise KV‑cache formulas required for memory prediction.
  • detect_moe (lines 93‑134) identifies Mixture of Experts (MoE) architectures and estimates active parameter counts, which significantly affects memory consumption calculations.

Merging Existing Values During Weekly Scrapes

To prevent data erosion during automated updates, preserve_existing_metadata() (lines 94‑123) copies non‑null values from the previous catalog version into the newly scraped entries. This ensures that manually verified fields—such as license information, language tags, and context‑length overrides—persist across weekly scrapes even if the Hugging Face API returns incomplete data.

Integrating GGUF Quantization Sources

After the initial scrape, enrich_gguf_sources() (lines 90‑165) queries known GGUF providers for pre‑quantized versions of each model. The script caches these results to avoid unnecessary API calls, adding download URLs and quantization variants to the schema.

Generating the Embedded JSON Output

The final dictionary structure (lines 54‑84) is written as a JSON array to llmfit-core/data/hf_models.json. This file is embedded directly into the Rust binary via include_str!, making the 33‑model schema available at compile time. The catalog drives all downstream components: hardware‑fit analysis, CLI table rendering, TUI views, and the web dashboard. For convenience, the scripts/update_models.sh wrapper automates the full pipeline from scraping to binary rebuild.

Running the Scraper to Update the Catalog

Execute the script from the repository root to regenerate the catalog:

  1. Basic execution (curated list + top 1000):
python3 scripts/scrape_hf_models.py
  1. Limit to top 500 models only:
python3 scripts/scrape_hf_models.py -n 500
  1. Accelerate with parallel threads:
python3 scripts/scrape_hf_models.py --threads 8
  1. Rebuild the Rust binary to embed the updated schema:
cargo build

Summary

  • scripts/scrape_hf_models.py creates and maintains the 33‑model schema by combining curated lists with live Hugging Face API data.
  • Key functions like estimate_ram(), estimate_vram(), and extract_arch_metadata derive hardware requirements and architectural details for accurate deployment predictions.
  • preserve_existing_metadata() safeguards against data loss during weekly automated updates by merging previous catalog values.
  • The script outputs to llmfit-core/data/hf_models.json, which is embedded into the Rust binary and powers llmfit’s hardware compatibility engine.

Frequently Asked Questions

What file does scripts/scrape_hf_models.py generate?

The script generates llmfit-core/data/hf_models.json, a JSON array containing the 33‑model schema. This file is embedded into the Rust binary via include_str! and contains hardware requirements, architectural metadata, and GGUF source links for each model.

How does the script prevent data loss during weekly updates?

The preserve_existing_metadata() function (lines 94‑123) copies non‑null values from the previous catalog version into the new entries. This ensures that fields like license information, context‑length overrides, and language tags are retained even if the Hugging Face API returns incomplete records during automated scrapes.

What hardware metrics are included in the 33-model schema?

The schema stores minimum RAM, recommended RAM, and minimum VRAM calculated by estimate_ram() (lines 443‑458) and estimate_vram() (lines 61‑67). These values are derived from parameter counts, default quantization settings, and architecture‑specific KV‑cache formulas.

How do I update the embedded catalog in the Rust binary?

After running python3 scripts/scrape_hf_models.py, execute cargo build to recompile the binary with the updated hf_models.json file. Alternatively, run the convenience wrapper scripts/update_models.sh, which performs both the Python scrape and the Rust rebuild in sequence.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →