# Which Models Are Supported by the MTPLX Backend Architecture? A Complete Catalog Guide

> Discover MTPLX backend architecture supported models including Qwen, Laguna, DeepSeek V4, and HY-V3. Explore our complete catalog guide for MTP heads and hardware filtering.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: catalog-guide
- Published: 2026-09-08

---

**The MTPLX backend architecture supports a curated catalog of MLX-compatible models including quantized Qwen variants (3.5B to 27B parameters), Laguna, DeepSeek V4, and HY-V3, all featuring multi-token prediction (MTP) heads and automatic memory-based hardware filtering.**

The MTPLX inference engine, hosted in the `youssofal/MTPLX` repository, implements a hardware-aware model catalog that determines which machine learning architectures can run efficiently on Apple Silicon devices. Understanding which models are supported by the MTPLX backend architecture requires examining the `OFFICIAL_CATALOG` tuple defined in [`mtplx/model_catalog.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_catalog.py), which enumerates every compatible model alongside its memory requirements, quantization profiles, and recommended hardware tiers.

## How the MTPLX Model Catalog Works

The backend initializes support by loading the `OFFICIAL_CATALOG` from [`mtplx/model_catalog.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_catalog.py), a Python tuple containing `CatalogModel` dataclass instances. Each entry specifies the Hugging Face model identifier, peak memory consumption in GiB, recommended hardware tiers, and quantization strategy.

At runtime, MTPLX filters this catalog using a `MEMORY_SAFETY_FACTOR` of 1.5 multiplied against each model's `peak_memory_gib` value. Only models whose adjusted memory footprint fits within the available device memory appear as selectable options. The `recommended_tiers` field further categorizes compatibility into **Modern** (Apple Silicon M3+), **Legacy** (M1/M2), **Intel**, or **Unknown** classifications.

## Complete Catalog of Supported Models

The MTPLX backend officially supports 18+ model configurations spanning the Qwen family and specialized architectures. Each model exposes an **MTP head** enabling multi-token prediction, though the system also supports pure autoregressive models via the `ar_only` flag or `--no-mtp` command-line option.

| Model ID | Display Name | Precision | Peak Memory | Tier |
|----------|--------------|-----------|-------------|------|
| `qwen35-4b-optimized-speed` | Qwen 3.5 4B Optimized Speed | 4-bit | 2.86 GiB | Modern |
| `qwen35-4b-optimized-quality` | Qwen 3.5 4B Optimized Quality | 8-bit | 4.75 GiB | Modern |
| `qwen35-9b-optimized-speed` | Qwen 3.5 9B Optimized Speed | 6-bit | 10.0 GiB | Modern |
| `qwen35-9b-optimized-speed-fp16` | Qwen 3.5 9B Optimized Speed FP16 | FP16 | 10.5 GiB | Legacy |
| `qwen36-35b-a3b-optimized-speed` | Qwen 3.6 35B-A3B Optimized Speed | 4-bit | 28.0 GiB | Modern |
| `qwen36-35b-a3b-optimized-speed-fp16` | Qwen 3.6 35B-A3B Optimized Speed FP16 | FP16 | 28.5 GiB | Legacy |
| `qwen36-35b-a3b-optimized-balance` | Qwen 3.6 35B-A3B Optimized Balance | 6-bit | 32.0 GiB | Modern |
| `qwen38-27b-bare-speed` | Qwen 3.8 27B Bare Speed | BF16 + MTP side-car | 20.0 GiB | Modern |
| `qwen38-27b-optimized-speed` | Qwen 3.8 27B Optimized Speed | 4-bit dynamic | 25.0 GiB | Modern |
| `qwen38-27b-optimized-quality` | Qwen 3.8 27B Optimized Quality | 8-bit dynamic | 33.0 GiB | Modern |
| `qwen38-27b-bare-speed-fp16` | Qwen 3.8 27B Bare Speed FP16 | FP16 | 20.0 GiB | Legacy |
| `qwen38-27b-optimized-speed-fp16` | Qwen 3.8 27B Optimized Speed FP16 | 4-bit + FP16 | 25.0 GiB | Legacy |
| `flash-next-bare-speed` | Qwen 3.8 Flash-Next Bare Speed | 4-bit flat | 78.0 GiB | Modern |
| `flash-next-optimized-speed` | Qwen 3.8 Flash-Next Optimized Speed | 4-bit + 8-bit attention | 87.0 GiB | Modern |
| `optimized-speed-v2` | Qwen 3.6 27B Optimized Speed V2 | Hybrid 4-bit | 21.5 GiB | Modern |
| `laguna` | Laguna (custom MLX) | Native MLX BF16 | ~12 GiB | Modern |
| `deepseek_v4` | DeepSeek V4 | BF16 + MTP draft lane | ~40 GiB | Modern |
| `hy_v3` | HY-V3 (vendored) | BF16 | ~14 GiB | Modern |

The `deepseek_v4` implementation specifically utilizes an MTP draft lane for speculative decoding, while the `laguna` model represents a custom native MLX architecture optimized for the MTPLX backend.

## Loading Models from the Catalog

Developers interact with supported models through the `load_model` function in [`mtplx/hf_loader.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/hf_loader.py), which patches standard MLX model classes with the MTP implementation from [`mtplx/mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_patch.py).

List compatible models based on device constraints:

```python
from mtplx.model_catalog import OFFICIAL_CATALOG

def list_supported_models():
    """Return models fitting current device memory."""
    safety_factor = 1.5
    device_memory = get_device_memory_gib()  # MTPLX runtime utility

    
    supported = [
        m for m in OFFICIAL_CATALOG 
        if m.peak_memory_gib * safety_factor < device_memory
    ]
    
    for model in supported:
        print(f"{model.display_name}: {model.hf_model_id}")

```

Initialize a specific model with MTP enabled:

```python
from mtplx.hf_loader import load_model
from mtplx.model_catalog import OFFICIAL_CATALOG

def load_optimized_qwen():
    """Load the Qwen 3.5 4B Optimized Speed variant."""
    model_id = "qwen35-4b-optimized-speed"
    entry = next(m for m in OFFICIAL_CATALOG if m.id == model_id)
    
    # load_model patches the class to expose MTP functionality

    model, config = load_model(
        entry.hf_model_id,
        get_model_classes=lambda cfg: (Model, ModelArgs)
    )
    return model

```

## Backend Architecture Components

The MTPLX backend architecture supporting these models consists of several interconnected modules:

- **[`mtplx/model_catalog.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_catalog.py)**: Defines `CatalogModel` dataclass and `OFFICIAL_CATALOG` tuple containing all supported model metadata.
- **[`mtplx/hf_loader.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/hf_loader.py)**: Handles Hugging Face model retrieval and class patching to inject MTP capabilities.
- **[`mtplx/mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_patch.py)**: Core logic that modifies standard MLX model forward passes to enable multi-token prediction.
- **[`mtplx/models/qwen4_exp.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/models/qwen4_exp.py)**: Implementation of Qwen 4-generation text and vision models.
- **[`mtplx/models/laguna.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/models/laguna.py)**: Custom Laguna model definition for native MLX execution.
- **[`mtplx/models/deepseek_v4.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/models/deepseek_v4.py)**: Speculative-draft DeepSeek V4 variant with MTP draft lanes.
- **[`mtplx/vendored_hy_v3.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vendored_hy_v3.py)**: Vendored HY-V3 model implementation for benchmarking purposes.

## Summary

- The **MTPLX backend architecture** supports 18+ quantized model variants stored in `OFFICIAL_CATALOG` inside [`mtplx/model_catalog.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_catalog.py).
- Supported families include **Qwen 3.5/3.6/3.8** (4B to 27B parameters), **Flash-Next**, **Laguna**, **DeepSeek V4**, and **HY-V3**.
- Models are filtered at runtime using a **1.5x memory safety factor** against `peak_memory_gib` values and categorized by **Modern** (M3+) or **Legacy** (M1/M2) hardware tiers.
- All catalog models feature **MLX-compatible MTP heads** for multi-token prediction, with fallback support for pure autoregressive mode via the `ar_only` flag or `--no-mtp` option.
- The loading pipeline in [`mtplx/hf_loader.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/hf_loader.py) automatically patches model classes to expose MTPLX-specific functionality.

## Frequently Asked Questions

### Which hardware tiers does MTPLX support?

MTPLX categorizes devices into **Modern** (Apple Silicon M3 and later), **Legacy** (M1/M2 chips requiring FP16 fallbacks), **Intel** processors, and **Unknown** classifications. The `recommended_tiers` field in each `CatalogModel` entry determines which hardware can safely run specific model variants without performance degradation or out-of-memory errors.

### Can I run models not listed in the OFFICIAL_CATALOG?

While the backend architecture is optimized for catalog models, MTPLX can load arbitrary MLX-compatible models using the `--no-mtp` flag. However, these non-catalog models lack the multi-token prediction head and must operate in pure autoregressive (AR) mode, bypassing the [`mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtp_patch.py) injection system entirely.

### How does the memory safety factor affect model availability?

The backend multiplies each model's `peak_memory_gib` value by a global `MEMORY_SAFETY_FACTOR` of 1.5 to calculate a safe memory estimate. If this adjusted value exceeds available device memory as reported by `get_device_memory_gib()`, the model is filtered from the user interface. This prevents out-of-memory crashes during inference when accounting for MTP overhead and activation caching.

### What is the difference between "Optimized Speed" and "Optimized Quality" variants?

**Optimized Speed** variants typically use 4-bit or 6-bit quantization to minimize latency and memory footprint (e.g., 2.86 GiB for Qwen 3.5 4B), while **Optimized Quality** variants employ 8-bit quantization or BF16 precision for higher fidelity at the cost of increased memory usage (up to 33 GiB for Qwen 3.8 27B). The catalog explicitly defines these trade-offs in the `quantization` and `peak_memory_gib` fields of each `CatalogModel` entry.