# What Is the Schema for Models in llmfit's JSON Catalogs?

> Understand the JSON schema for models in llmfit catalogs like hf_models.json and onnx_models.json. Ensure consistent metadata for Hugging Face models with llmfit.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: api-reference
- Published: 2026-08-22

---

**The [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) catalog adheres to a formal JSON Schema defined in [`schema.json`](https://github.com/AlexsJones/llmfit/blob/main/schema.json), while [`onnx_models.json`](https://github.com/AlexsJones/llmfit/blob/main/onnx_models.json) follows an implicit schema hard-coded in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs), both standardizing metadata for Hugging Face and ONNX model discovery.**

The `llmfit` project by AlexsJones maintains two canonical JSON catalogs—[`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) and [`onnx_models.json`](https://github.com/AlexsJones/llmfit/blob/main/onnx_models.json)—that enumerate supported large language models. Understanding the **schema for models in llmfit's JSON catalogs** is essential for contributors adding new models or developers integrating the library, as each file serves distinct provider ecosystems with specific validation requirements.

## Hugging Face Model Schema (schema.json)

The Hugging Face catalog is governed by [`llmfit-core/data/schema.json`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/data/schema.json), which declares `"$id": "https://llmfit.dev/schemas/hf_models.schema.json"`. This schema defines an array of model objects with strict typing for hardware requirements and model capabilities.

### Required Fields

Every entry in [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) must include:

- `name`: String matching pattern `^.+/.+$` (e.g., `facebook/opt-125m`)
- `provider`: String (typically "huggingface")
- `parameter_count`: String (e.g., "1.3B")
- `min_ram_gb` and `recommended_ram_gb`: Numbers ≥ 0
- `quantization`: String (e.g., `q4_k_m`)
- `context_length`: Integer ≥ 1
- `use_case`: String describing the primary application

### Optional Metadata Fields

The schema permits extensive optional fields for advanced filtering:

- `min_vram_gb`: Number for GPU requirements
- `format`: String (e.g., "gguf")
- `capabilities`: Array of enums like "chat", "code", "vision"
- `languages`: Array of ISO codes
- `architecture`: String (e.g., "llama")
- MoE-specific fields: `is_moe`, `num_experts`, `active_experts`, `moe_intermediate_size`
- `gguf_sources`: Array of download objects
- `attention_layout`: String (e.g., "mha", "gqa")

### Schema Validation

The repository uses this schema to validate [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) at build time, ensuring all entries conform to the expected structure before the Rust parser ingests them.

## ONNX Model Schema (onnx_models.json)

Unlike the HF catalog, [`onnx_models.json`](https://github.com/AlexsJones/llmfit/blob/main/onnx_models.json) lacks a separate JSON Schema file. Instead, [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) enforces an implicit schema during deserialization, expecting a flatter structure optimized for ONNX runtime requirements.

### Core Structure

Each ONNX model entry requires:

- `id`: String in format `onnx-community/<model-name>`
- `name`: Human-readable string
- `parameters`: String representation of model size (e.g., "3.8B")
- `license`: SPDX identifier string
- `format`: Must be the literal string `"onnx"` (validated by the loader)

### The onnx_files Map

The critical distinction is the `onnx_files` object, which maps precision keys to file sizes in bytes:

```json
"onnx_files": {
  "fp32": 15300000000,
  "fp16": 7660000000,
  "q8": 3820000000,
  "int8": 3820000000,
  "uint8": 3820000000,
  "q4": 2390000000,
  "q4f16": 2440000000,
  "bnb4": 2280000000
}

```

Valid precision keys include `fp32`, `fp16`, `q8`, `int8`, `uint8`, `q4`, `q4f16`, and `bnb4`. The Rust parser validates that at least one precision entry exists.

### Validation Rules

The loader in [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs) performs runtime checks:

- Asserts `format == "onnx"`
- Verifies `onnx_files` contains at least one key
- Confirms file sizes are positive integers

## Parsing Both Schemas in Rust

The [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) file handles deserialization for both catalogs. While the HF schema relies on Serde's derives matching [`schema.json`](https://github.com/AlexsJones/llmfit/blob/main/schema.json), the ONNX schema uses custom validation logic.

Example deserialization structures:

```rust
use serde::Deserialize;
use std::collections::HashMap;

#[derive(Debug, Deserialize)]
pub struct HfModel {
    pub name: String,
    pub provider: String,
    pub parameter_count: String,
    pub min_ram_gb: f64,
    pub recommended_ram_gb: f64,
    pub quantization: String,
    pub context_length: u32,
    pub use_case: String,
    pub min_vram_gb: Option<f64>,
    pub format: Option<String>,
    pub capabilities: Option<Vec<String>>,
    pub is_moe: Option<bool>,
    pub num_experts: Option<u32>,
    pub gguf_sources: Option<Vec<GgufSource>>,
}

#[derive(Debug, Deserialize)]
pub struct OnnxModel {
    pub id: String,
    pub name: String,
    pub parameters: String,
    pub license: String,
    pub format: String,
    pub onnx_files: HashMap<String, u64>,
}

impl OnnxModel {
    pub fn validate(&self) -> Result<(), String> {
        if self.format != "onnx" {
            return Err("ONNX models must have format='onnx'".into());
        }
        if self.onnx_files.is_empty() {
            return Err("At least one precision required in onnx_files".into());
        }
        Ok(())
    }
}

```

The library embeds both JSON files via `include_str!` and parses them at initialization, merging valid entries into the internal `LlmModel` representation used for hardware compatibility analysis.

## Summary

- **[`schema.json`](https://github.com/AlexsJones/llmfit/blob/main/schema.json)** provides formal validation for [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) with 20+ fields covering hardware specs, MoE configuration, and GGUF sources.
- **[`onnx_models.json`](https://github.com/AlexsJones/llmfit/blob/main/onnx_models.json)** uses an implicit schema requiring `format: "onnx"` and an `onnx_files` map with precision-specific byte sizes.
- **[`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs)** implements the Rust parser that deserializes both formats and enforces ONNX-specific validation rules.
- Both catalogs are embedded at compile time and merged into a unified internal model representation for the llmfit analysis pipeline.

## Frequently Asked Questions

### What is the difference between the HF and ONNX schemas in llmfit?

The HF schema is formally defined in [`schema.json`](https://github.com/AlexsJones/llmfit/blob/main/schema.json) and includes extensive metadata for quantization, context length, and MoE architecture. The ONNX schema is simpler and enforced by code, focusing on file sizes per precision format rather than architectural details.

### Where is the JSON Schema file for ONNX models?

There is no separate JSON Schema file for ONNX models. The structure is implicitly defined in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs), which expects the `format` field to equal `"onnx"` and requires an `onnx_files` object containing at least one precision key.

### How does llmfit validate the quantization field?

For Hugging Face models, the `quantization` field is validated against the [`schema.json`](https://github.com/AlexsJones/llmfit/blob/main/schema.json) definition as a required string. For ONNX models, quantization is represented as keys within the `onnx_files` map (e.g., `q4`, `q8`), validated by checking for the presence of these keys at runtime.

### Can I add custom fields to the model catalogs?

Adding custom fields to [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) requires updating [`schema.json`](https://github.com/AlexsJones/llmfit/blob/main/schema.json) to maintain validation compliance. For [`onnx_models.json`](https://github.com/AlexsJones/llmfit/blob/main/onnx_models.json), you must modify the `OnnxModel` struct in [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs) and update the validation logic to handle additional fields without breaking existing parsers.