How to Regenerate the Embedded hf_models.json Catalog in llmfit

The embedded hf_models.json catalog is regenerated automatically by the scripts/scrape_hf_models.py scraper, which fetches model metadata from the Hugging Face API and writes it to llmfit-core/data/hf_models.json, after which Rust embeds the file at compile time using include_str!.

The llmfit project maintains an authoritative catalog of Large Language Models by embedding a JSON file directly into the Rust core library. Understanding how the embedded hf_models.json catalog is regenerated ensures developers can safely update model definitions without manual interventions that would be overwritten.

The Regeneration Pipeline

The regeneration process follows a strict three-phase pipeline: data acquisition, file generation, and binary embedding.

Step 1: Scraping Hugging Face Metadata

The regeneration starts in scripts/scrape_hf_models.py, a Python scraper that contacts the Hugging Face API. The script targets a curated list of models defined in the TARGET_MODELS constant and optionally discovers the most-downloaded models dynamically.

For each model, the scraper extracts comprehensive metadata including parameter counts, RAM and VRAM estimates, architecture details, quantization formats, use-case tags, capabilities, supported languages, and licensing information. This logic resides in the main scraper body, ensuring the catalog captures quantitative hardware requirements alongside qualitative model characteristics.

Step 2: Writing the Catalog File

Once metadata collection completes, the script serializes the gathered list of model dictionaries to disk. The output is written to llmfit-core/data/hf_models.json, as defined by the output_paths variable at lines 3021–3024 in the scraper source.

The resulting JSON serves as the single source of truth for the application. It is not checked into version control as a static asset to be hand-edited; rather, it is treated as a build artifact that the scraper regenerates on demand.

Step 3: Embedding at Compile Time

The Rust core library ingests this JSON at compile time using the include_str! macro. In llmfit-core/src/models.rs at lines 1301–1302, the code declares:

const HF_MODELS_JSON: &str = include_str!("../data/hf_models.json");

This directive embeds the file contents directly into the compiled binary as a string literal. Consequently, any changes to hf_models.json require a recompilation of the Rust code to update the embedded catalog.

Automation and Developer Workflows

The project supports both automated scheduled updates and manual developer-driven regeneration.

Weekly Automated Refresh

A GitHub Actions workflow in .github/workflows/weekly-model-update.yml automates the regeneration process. Running on a weekly schedule, the workflow executes the scraper and validates the output using python -m json.tool to ensure syntactic correctness before committing.

The workflow at lines 54–55 handles the commit and push operations, ensuring the repository stays synchronized with the latest Hugging Face model releases without manual intervention.

Manual Local Regeneration

When working locally, developers must regenerate the catalog and rebuild the binary to see changes. Execute the following from the repository root:


# Regenerate the JSON catalog

python3 scripts/scrape_hf_models.py

# Rebuild to embed the updated catalog

cargo build

After running these commands, the new hf_models.json is baked into the binary produced by cargo build.

To inspect the generated JSON without rebuilding:

jq '.' llmfit-core/data/hf_models.json | less

Why Manual Edits Are Prohibited

The project documentation in docs/how-it-works.md (lines 125–130) explicitly warns against editing hf_models.json by hand. Because the file is regenerated automatically by the scraper, manual modifications would be lost during the next CI run or local refresh. The supported workflow always requires running the scraper to update model definitions, ensuring data consistency and traceability through version-controlled Python code rather than opaque JSON diffs.

Summary

Frequently Asked Questions

How often is the hf_models.json catalog updated automatically?

The catalog updates automatically once per week via the GitHub Actions workflow defined in .github/workflows/weekly-model-update.yml. The workflow runs the scraper, validates the JSON structure with python -m json.tool, and commits any changes back to the repository.

Can I add a custom model to the catalog without using the scraper?

No. The documentation in docs/how-it-works.md explicitly prohibits manual edits to llmfit-core/data/hf_models.json. To add a custom model, you must modify the TARGET_MODELS constant or discovery logic in scripts/scrape_hf_models.py and run the scraper to regenerate the file, ensuring the change persists through subsequent automated updates.

Why does the Rust binary need to be rebuilt after regenerating the JSON?

The Rust code uses include_str! to embed hf_models.json at compile time (see llmfit-core/src/models.rs, lines 1301–1302). This macro reads the file during compilation and inlines its contents as a string constant. Therefore, changes to the JSON file only take effect after running cargo build to produce a new binary with the updated catalog embedded.

What metadata does the scraper extract from Hugging Face?

The scraper extracts parameter counts, estimated RAM and VRAM requirements, architecture details, supported quantization formats, use-case classifications, model capabilities, supported languages, and licensing information. This comprehensive metadata enables llmfit to match users with appropriate models based on hardware constraints and task requirements.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →