Which Models Are Supported by the MTPLX Backend Architecture? A Complete Catalog Guide

The MTPLX backend architecture supports a curated catalog of MLX-compatible models including quantized Qwen variants (3.5B to 27B parameters), Laguna, DeepSeek V4, and HY-V3, all featuring multi-token prediction (MTP) heads and automatic memory-based hardware filtering.

The MTPLX inference engine, hosted in the youssofal/MTPLX repository, implements a hardware-aware model catalog that determines which machine learning architectures can run efficiently on Apple Silicon devices. Understanding which models are supported by the MTPLX backend architecture requires examining the OFFICIAL_CATALOG tuple defined in mtplx/model_catalog.py, which enumerates every compatible model alongside its memory requirements, quantization profiles, and recommended hardware tiers.

How the MTPLX Model Catalog Works

The backend initializes support by loading the OFFICIAL_CATALOG from mtplx/model_catalog.py, a Python tuple containing CatalogModel dataclass instances. Each entry specifies the Hugging Face model identifier, peak memory consumption in GiB, recommended hardware tiers, and quantization strategy.

At runtime, MTPLX filters this catalog using a MEMORY_SAFETY_FACTOR of 1.5 multiplied against each model's peak_memory_gib value. Only models whose adjusted memory footprint fits within the available device memory appear as selectable options. The recommended_tiers field further categorizes compatibility into Modern (Apple Silicon M3+), Legacy (M1/M2), Intel, or Unknown classifications.

Complete Catalog of Supported Models

The MTPLX backend officially supports 18+ model configurations spanning the Qwen family and specialized architectures. Each model exposes an MTP head enabling multi-token prediction, though the system also supports pure autoregressive models via the ar_only flag or --no-mtp command-line option.

Model ID Display Name Precision Peak Memory Tier
qwen35-4b-optimized-speed Qwen 3.5 4B Optimized Speed 4-bit 2.86 GiB Modern
qwen35-4b-optimized-quality Qwen 3.5 4B Optimized Quality 8-bit 4.75 GiB Modern
qwen35-9b-optimized-speed Qwen 3.5 9B Optimized Speed 6-bit 10.0 GiB Modern
qwen35-9b-optimized-speed-fp16 Qwen 3.5 9B Optimized Speed FP16 FP16 10.5 GiB Legacy
qwen36-35b-a3b-optimized-speed Qwen 3.6 35B-A3B Optimized Speed 4-bit 28.0 GiB Modern
qwen36-35b-a3b-optimized-speed-fp16 Qwen 3.6 35B-A3B Optimized Speed FP16 FP16 28.5 GiB Legacy
qwen36-35b-a3b-optimized-balance Qwen 3.6 35B-A3B Optimized Balance 6-bit 32.0 GiB Modern
qwen38-27b-bare-speed Qwen 3.8 27B Bare Speed BF16 + MTP side-car 20.0 GiB Modern
qwen38-27b-optimized-speed Qwen 3.8 27B Optimized Speed 4-bit dynamic 25.0 GiB Modern
qwen38-27b-optimized-quality Qwen 3.8 27B Optimized Quality 8-bit dynamic 33.0 GiB Modern
qwen38-27b-bare-speed-fp16 Qwen 3.8 27B Bare Speed FP16 FP16 20.0 GiB Legacy
qwen38-27b-optimized-speed-fp16 Qwen 3.8 27B Optimized Speed FP16 4-bit + FP16 25.0 GiB Legacy
flash-next-bare-speed Qwen 3.8 Flash-Next Bare Speed 4-bit flat 78.0 GiB Modern
flash-next-optimized-speed Qwen 3.8 Flash-Next Optimized Speed 4-bit + 8-bit attention 87.0 GiB Modern
optimized-speed-v2 Qwen 3.6 27B Optimized Speed V2 Hybrid 4-bit 21.5 GiB Modern
laguna Laguna (custom MLX) Native MLX BF16 ~12 GiB Modern
deepseek_v4 DeepSeek V4 BF16 + MTP draft lane ~40 GiB Modern
hy_v3 HY-V3 (vendored) BF16 ~14 GiB Modern

The deepseek_v4 implementation specifically utilizes an MTP draft lane for speculative decoding, while the laguna model represents a custom native MLX architecture optimized for the MTPLX backend.

Loading Models from the Catalog

Developers interact with supported models through the load_model function in mtplx/hf_loader.py, which patches standard MLX model classes with the MTP implementation from mtplx/mtp_patch.py.

List compatible models based on device constraints:

from mtplx.model_catalog import OFFICIAL_CATALOG

def list_supported_models():
    """Return models fitting current device memory."""
    safety_factor = 1.5
    device_memory = get_device_memory_gib()  # MTPLX runtime utility

    
    supported = [
        m for m in OFFICIAL_CATALOG 
        if m.peak_memory_gib * safety_factor < device_memory
    ]
    
    for model in supported:
        print(f"{model.display_name}: {model.hf_model_id}")

Initialize a specific model with MTP enabled:

from mtplx.hf_loader import load_model
from mtplx.model_catalog import OFFICIAL_CATALOG

def load_optimized_qwen():
    """Load the Qwen 3.5 4B Optimized Speed variant."""
    model_id = "qwen35-4b-optimized-speed"
    entry = next(m for m in OFFICIAL_CATALOG if m.id == model_id)
    
    # load_model patches the class to expose MTP functionality

    model, config = load_model(
        entry.hf_model_id,
        get_model_classes=lambda cfg: (Model, ModelArgs)
    )
    return model

Backend Architecture Components

The MTPLX backend architecture supporting these models consists of several interconnected modules:

Summary

  • The MTPLX backend architecture supports 18+ quantized model variants stored in OFFICIAL_CATALOG inside mtplx/model_catalog.py.
  • Supported families include Qwen 3.5/3.6/3.8 (4B to 27B parameters), Flash-Next, Laguna, DeepSeek V4, and HY-V3.
  • Models are filtered at runtime using a 1.5x memory safety factor against peak_memory_gib values and categorized by Modern (M3+) or Legacy (M1/M2) hardware tiers.
  • All catalog models feature MLX-compatible MTP heads for multi-token prediction, with fallback support for pure autoregressive mode via the ar_only flag or --no-mtp option.
  • The loading pipeline in mtplx/hf_loader.py automatically patches model classes to expose MTPLX-specific functionality.

Frequently Asked Questions

Which hardware tiers does MTPLX support?

MTPLX categorizes devices into Modern (Apple Silicon M3 and later), Legacy (M1/M2 chips requiring FP16 fallbacks), Intel processors, and Unknown classifications. The recommended_tiers field in each CatalogModel entry determines which hardware can safely run specific model variants without performance degradation or out-of-memory errors.

Can I run models not listed in the OFFICIAL_CATALOG?

While the backend architecture is optimized for catalog models, MTPLX can load arbitrary MLX-compatible models using the --no-mtp flag. However, these non-catalog models lack the multi-token prediction head and must operate in pure autoregressive (AR) mode, bypassing the mtp_patch.py injection system entirely.

How does the memory safety factor affect model availability?

The backend multiplies each model's peak_memory_gib value by a global MEMORY_SAFETY_FACTOR of 1.5 to calculate a safe memory estimate. If this adjusted value exceeds available device memory as reported by get_device_memory_gib(), the model is filtered from the user interface. This prevents out-of-memory crashes during inference when accounting for MTP overhead and activation caching.

What is the difference between "Optimized Speed" and "Optimized Quality" variants?

Optimized Speed variants typically use 4-bit or 6-bit quantization to minimize latency and memory footprint (e.g., 2.86 GiB for Qwen 3.5 4B), while Optimized Quality variants employ 8-bit quantization or BF16 precision for higher fidelity at the cost of increased memory usage (up to 33 GiB for Qwen 3.8 27B). The catalog explicitly defines these trade-offs in the quantization and peak_memory_gib fields of each CatalogModel entry.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →