How to Add a New Model to llmfit's Catalog: A Step-by-Step Guide
Add the Hugging Face model ID to the TARGET_MODELS list in scripts/scrape_hf_models.py, run the scraper to regenerate llmfit-core/data/hf_models.json, and rebuild the Rust project with cargo build.
The llmfit project (AlexsJones/llmfit) maintains a curated catalog of Large Language Models to help users determine hardware compatibility. Adding a new model to llmfit's catalog requires updating the scraper configuration and regenerating the embedded JSON database that powers the CLI and TUI interfaces.
Understand the Catalog Architecture
The model catalog relies on two critical components according to the source code documented in AGENTS.md. The scraper script (scripts/scrape_hf_models.py) fetches metadata from Hugging Face, while the embedded catalog (llmfit-core/data/hf_models.json) stores the compiled data that the Rust core reads at compile time.
Modify the Scraper Configuration
Add the Model ID to TARGET_MODELS
Edit scripts/scrape_hf_models.py and locate the TARGET_MODELS list. Add the Hugging Face repository identifier for your model:
TARGET_MODELS = [
"meta-llama/Meta-Llama-3-8B-Instruct",
"mistralai/Mistral-7B-Instruct-v0.2",
# Add your new model here
"your-org/your-model-name",
]
Handle Gated Models with FALLBACK (Optional)
If the model requires authentication on Hugging Face, add a fallback entry to the FALLBACK dictionary in the same file. This ensures the scraper generates a placeholder record even when API access is restricted:
FALLBACK = {
"meta-llama/Meta-Llama-3-70B-Instruct": {
"provider": "hf",
"parameter_count": 70_000_000_000,
"memory_minimum_ram_gb": 140,
"memory_minimum_vram_gb": 80,
"architecture": "transformers",
"quantization": "none"
}
}
Regenerate the Catalog
Run the Scraper Script
Execute the Python script to fetch model metadata and regenerate the JSON catalog:
python3 scripts/scrape_hf_models.py
The script pulls data including parameter counts, quantization formats, and memory requirements, then writes the updated catalog to llmfit-core/data/hf_models.json.
Verify the JSON Output
Confirm the new model appears correctly in the generated catalog using jq or a text editor:
jq '.[] | select(.name | contains("your-model-name"))' llmfit-core/data/hf_models.json
Verify the record includes required fields: name, provider, parameter_count, memory_minimum_ram_gb, and memory_minimum_vram_gb.
Rebuild the Rust Project
Compile the Rust project to embed the updated catalog into the binary:
cargo build
Alternatively, run cargo test to verify everything compiles and functions correctly. The Rust code embeds hf_models.json at compile time, so changes only take effect after rebuilding.
Summary
- Catalog source:
llmfit-core/data/hf_models.jsonis generated byscripts/scrape_hf_models.py - Registration: Add model IDs to the
TARGET_MODELSlist in the scraper script - Gated access: Use the
FALLBACKdictionary for models requiring Hugging Face authentication - Generation: Run
python3 scripts/scrape_hf_models.pyto update the JSON catalog - Deployment: Execute
cargo buildto embed the new catalog into the llmfit binary
Frequently Asked Questions
How do I add a model that requires Hugging Face authentication?
Add the model's metadata to the FALLBACK dictionary in scripts/scrape_hf_models.py. This bypasses the API requirement and creates a manual entry with predefined parameter counts and memory requirements.
Where is the model catalog stored in the repository?
The compiled catalog lives at llmfit-core/data/hf_models.json. This file is auto-generated by the scraper script and embedded into the Rust binary at compile time. Do not edit this file manually; instead, modify the scraper and regenerate it.
Why do I need to run cargo build after updating the JSON?
The llmfit Rust core embeds the catalog directly into the binary using compile-time macros. Changes to hf_models.json are only visible after recompiling the project, as the JSON becomes part of the static binary rather than being read at runtime.
Can I add multiple models at once?
Yes. Add multiple entries to the TARGET_MODELS list in scripts/scrape_hf_models.py, then run the scraper once. The script will fetch metadata for all listed models and update the catalog file in a single execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →