# How MTPLX Manages Model Compatibility Using Its Tiered RAM-Based System

> Discover how MTPLX ensures model compatibility with its innovative RAM-tiered system, prioritizing smaller, unquantized models via deterministic sorting.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: architecture
- Published: 2026-09-05

---

**MTPLX manages model compatibility through a RAM-tiered ranking system that prioritizes smaller, unquantized models over larger quantized variants using deterministic sorting logic in** [`mtplx/model_catalog.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_catalog.py).

The MTPLX project implements a sophisticated model catalog system to handle diverse AI model specifications through structured metadata. By utilizing a tiered approach based on memory requirements and quantization status, MTPLX ensures optimal model selection for varying hardware constraints. This compatibility system centers on the `ModelInfo` dataclass and intelligent ranking algorithms that process JSON catalog files.

## Understanding the Model Catalog Architecture

### The ModelInfo Dataclass Structure

At the core of MTPLX's compatibility system lies the `ModelInfo` dataclass defined in [`mtplx/model_catalog.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_catalog.py). This immutable container stores essential metadata including `model_id`, `family`, `version`, `parameters` (in billions), `quantized` status, and `tags`. The frozen dataclass design ensures that model metadata remains constant throughout the application lifecycle, preventing accidental mutations that could break compatibility checks.

### Configurable Catalog Discovery

The system locates model definitions through JSON files stored in a configurable directory path. The `_MODEL_CATALOG_PATH` variable reads from the `MTPLX_MODEL_CATALOG_PATH` environment variable, defaulting to a `model_catalogs` directory relative to the module location. The `_discover_catalog_files()` function recursively scans this path for `.json` files, enabling decentralized model definitions across multiple catalog files.

## The RAM-Tiered Compatibility Ranking System

### How the Tiered Ranking Algorithm Works

The `_rank_by_ram_tier()` function implements the core compatibility logic using a tuple-based sorting key: `(i.quantized, i.parameters, i.model_id)`. This creates a three-tier priority system where:

- **Unquantized models** sort before quantized models (False < True)
- **Smaller parameter counts** receive priority over larger models
- **Alphabetical model_id** breaks ties for deterministic ordering

### Memory Footprint Prioritization

The ranking explicitly favors models with lower memory footprints by sorting `parameters` in ascending order. A 7-billion parameter model ranks higher than a 70-billion parameter variant. Combined with the quantization tier (where unquantized models precede quantized ones), this creates a hardware-friendly default selection that minimizes RAM requirements for initial deployments.

## Loading and Validating Model Metadata

### JSON Catalog Processing

The `load_model_catalog()` function orchestrates the discovery and parsing pipeline. It iterates through discovered JSON files, validates entries using `_parse_model_entry()`, and returns a deterministically sorted list of `ModelInfo` objects. The parser enforces required fields (`model_id` and `family`) while coercing optional fields like `quantized` to boolean types and normalizing `tags` into sorted tuples.

### Version String Validation

MTPLX enforces strict version formatting through the `_VERSION_RE` regular expression pattern: `^(\d+(?:\.\d+)?)$`. This accepts simple version strings like "1.2" or "3" while rejecting malformed entries. The validation occurs during entry parsing, ensuring that only properly versioned models enter the compatibility tier system.

## Implementation in Practice

### Retrieving Recommended Models

The `recommended_models()` function provides the primary API for accessing the tiered compatibility system. Building upon `load_model_catalog()` and `_rank_by_ram_tier()`, it returns a list of `model_id` strings ordered by RAM efficiency. For backwards compatibility, the module exposes `recommended_models_ids` as an alias to this function.

```python
from mtplx.model_catalog import recommended_models

# Get models ordered by RAM tier (unquantized, small parameters first)

compatible_models = recommended_models()
print(compatible_models)  # ['llama-2-7b', 'llama-2-13b', 'llama-2-70b-q4', ...]

```

### Customizing the Catalog Path

Deployers can override the default catalog location by setting the environment variable before import:

```python
import os
os.environ["MTPLX_MODEL_CATALOG_PATH"] = "/path/to/custom/catalogs"

from mtplx.model_catalog import load_model_catalog
models = load_model_catalog()

```

## Summary

- **MTPLX uses a RAM-tiered ranking system** in [`mtplx/model_catalog.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_catalog.py) that prioritizes unquantized, low-parameter models for optimal hardware compatibility.
- **The `_rank_by_ram_tier()` function** applies a three-level sort: quantization status (False before True), parameter count (ascending), and model_id (alphabetical).
- **Immutable `ModelInfo` dataclasses** store canonical metadata including version strings validated against the `_VERSION_RE` pattern.
- **Configurable catalog discovery** via `MTPLX_MODEL_CATALOG_PATH` enables flexible deployment configurations across different environments.
- **Backwards-compatible APIs** like `recommended_models_ids` ensure existing integrations remain functional as the tiered system evolves.

## Frequently Asked Questions

### How does MTPLX prioritize models with the same parameter count?

When models share identical `parameters` values and quantization status, the tiered system falls back to alphabetical ordering by `model_id`. This deterministic tie-breaking ensures consistent results across repeated queries, as implemented in the `_rank_by_ram_tier()` function's tuple sorting key.

### What happens if a model catalog JSON file contains invalid entries?

The `_parse_model_entry()` function silently skips malformed entries by returning `None` when validation fails or required fields are missing. This fault-tolerant approach allows partial catalog loading where valid models load successfully even if some entries contain corrupted data or unsupported version strings.

### Can the tiered ranking system be customized to prioritize quantized models?

The current implementation in [`mtplx/model_catalog.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_catalog.py) hardcodes the preference for unquantized models (False sorts before True). To reverse this priority, you would need to modify the sorting key in `_rank_by_ram_tier()` to use `(not i.quantized, i.parameters, i.model_id)` or implement a custom sorting wrapper around `load_model_catalog()`.

### How does MTPLX handle model versioning?

The system validates version strings using the `_VERSION_RE` regular expression pattern `^(\d+(?:\.\d+)?)$`, accepting simple numeric formats like "2" or "3.1". Parsed versions store as strings in the `ModelInfo.version` field, though the current tiered ranking does not use version information for compatibility sorting—focusing instead on RAM requirements and quantization status.