# How to Regenerate the Embedded hf_models.json Catalog in llmfit

> Learn how to regenerate the embedded hf_models.json catalog in llmfit using the scrape_hf_models.py script. Understand the process of fetching and embedding Hugging Face model metadata.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: how-to-guide
- Published: 2026-09-11

---

**The embedded [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) catalog is regenerated automatically by the [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py) scraper, which fetches model metadata from the Hugging Face API and writes it to [`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json), after which Rust embeds the file at compile time using `include_str!`.**

The `llmfit` project maintains an authoritative catalog of Large Language Models by embedding a JSON file directly into the Rust core library. Understanding how the embedded [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) catalog is regenerated ensures developers can safely update model definitions without manual interventions that would be overwritten.

## The Regeneration Pipeline

The regeneration process follows a strict three-phase pipeline: data acquisition, file generation, and binary embedding.

### Step 1: Scraping Hugging Face Metadata

The regeneration starts in [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py), a Python scraper that contacts the Hugging Face API. The script targets a curated list of models defined in the `TARGET_MODELS` constant and optionally discovers the most-downloaded models dynamically.

For each model, the scraper extracts comprehensive metadata including parameter counts, RAM and VRAM estimates, architecture details, quantization formats, use-case tags, capabilities, supported languages, and licensing information. This logic resides in the main scraper body, ensuring the catalog captures quantitative hardware requirements alongside qualitative model characteristics.

### Step 2: Writing the Catalog File

Once metadata collection completes, the script serializes the gathered list of model dictionaries to disk. The output is written to [`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json), as defined by the `output_paths` variable at lines 3021–3024 in the scraper source.

The resulting JSON serves as the single source of truth for the application. It is **not** checked into version control as a static asset to be hand-edited; rather, it is treated as a build artifact that the scraper regenerates on demand.

### Step 3: Embedding at Compile Time

The Rust core library ingests this JSON at compile time using the `include_str!` macro. In [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) at lines 1301–1302, the code declares:

```rust
const HF_MODELS_JSON: &str = include_str!("../data/hf_models.json");

```

This directive embeds the file contents directly into the compiled binary as a string literal. Consequently, any changes to [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) require a recompilation of the Rust code to update the embedded catalog.

## Automation and Developer Workflows

The project supports both automated scheduled updates and manual developer-driven regeneration.

### Weekly Automated Refresh

A GitHub Actions workflow in [`.github/workflows/weekly-model-update.yml`](https://github.com/AlexsJones/llmfit/blob/main/.github/workflows/weekly-model-update.yml) automates the regeneration process. Running on a weekly schedule, the workflow executes the scraper and validates the output using `python -m json.tool` to ensure syntactic correctness before committing.

The workflow at lines 54–55 handles the commit and push operations, ensuring the repository stays synchronized with the latest Hugging Face model releases without manual intervention.

### Manual Local Regeneration

When working locally, developers must regenerate the catalog and rebuild the binary to see changes. Execute the following from the repository root:

```bash

# Regenerate the JSON catalog

python3 scripts/scrape_hf_models.py

# Rebuild to embed the updated catalog

cargo build

```

After running these commands, the new [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) is baked into the binary produced by `cargo build`.

To inspect the generated JSON without rebuilding:

```bash
jq '.' llmfit-core/data/hf_models.json | less

```

## Why Manual Edits Are Prohibited

The project documentation in [`docs/how-it-works.md`](https://github.com/AlexsJones/llmfit/blob/main/docs/how-it-works.md) (lines 125–130) explicitly warns against editing [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) by hand. Because the file is regenerated automatically by the scraper, manual modifications would be lost during the next CI run or local refresh. The supported workflow always requires running the scraper to update model definitions, ensuring data consistency and traceability through version-controlled Python code rather than opaque JSON diffs.

## Summary

- The [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py) scraper fetches metadata from the Hugging Face API for models listed in `TARGET_MODELS`.
- It writes the processed data to [`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json) via the `output_paths` variable (lines 3021–3024).
- Rust embeds the file at compile time using `include_str!` in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) (lines 1301–1302).
- The [`.github/workflows/weekly-model-update.yml`](https://github.com/AlexsJones/llmfit/blob/main/.github/workflows/weekly-model-update.yml) CI job automates weekly refreshes and validates JSON syntax.
- Manual edits to the JSON file are discouraged; always use the scraper to regenerate the catalog.

## Frequently Asked Questions

### How often is the hf_models.json catalog updated automatically?

The catalog updates automatically once per week via the GitHub Actions workflow defined in [`.github/workflows/weekly-model-update.yml`](https://github.com/AlexsJones/llmfit/blob/main/.github/workflows/weekly-model-update.yml). The workflow runs the scraper, validates the JSON structure with `python -m json.tool`, and commits any changes back to the repository.

### Can I add a custom model to the catalog without using the scraper?

No. The documentation in [`docs/how-it-works.md`](https://github.com/AlexsJones/llmfit/blob/main/docs/how-it-works.md) explicitly prohibits manual edits to [`llmfit-core/data/hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/hf_models.json). To add a custom model, you must modify the `TARGET_MODELS` constant or discovery logic in [`scripts/scrape_hf_models.py`](https://github.com/AlexsJones/llmfit/blob/main/scripts/scrape_hf_models.py) and run the scraper to regenerate the file, ensuring the change persists through subsequent automated updates.

### Why does the Rust binary need to be rebuilt after regenerating the JSON?

The Rust code uses `include_str!` to embed [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) at compile time (see [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs), lines 1301–1302). This macro reads the file during compilation and inlines its contents as a string constant. Therefore, changes to the JSON file only take effect after running `cargo build` to produce a new binary with the updated catalog embedded.

### What metadata does the scraper extract from Hugging Face?

The scraper extracts parameter counts, estimated RAM and VRAM requirements, architecture details, supported quantization formats, use-case classifications, model capabilities, supported languages, and licensing information. This comprehensive metadata enables `llmfit` to match users with appropriate models based on hardware constraints and task requirements.