Which Models Are Supported by the MTPLX Backend Architecture? A Complete Catalog Guide
The MTPLX backend architecture supports a curated catalog of MLX-compatible models including quantized Qwen variants (3.5B to 27B parameters), Laguna, DeepSeek V4, and HY-V3, all featuring multi-token prediction (MTP) heads and automatic memory-based hardware filtering.
The MTPLX inference engine, hosted in the youssofal/MTPLX repository, implements a hardware-aware model catalog that determines which machine learning architectures can run efficiently on Apple Silicon devices. Understanding which models are supported by the MTPLX backend architecture requires examining the OFFICIAL_CATALOG tuple defined in mtplx/model_catalog.py, which enumerates every compatible model alongside its memory requirements, quantization profiles, and recommended hardware tiers.
How the MTPLX Model Catalog Works
The backend initializes support by loading the OFFICIAL_CATALOG from mtplx/model_catalog.py, a Python tuple containing CatalogModel dataclass instances. Each entry specifies the Hugging Face model identifier, peak memory consumption in GiB, recommended hardware tiers, and quantization strategy.
At runtime, MTPLX filters this catalog using a MEMORY_SAFETY_FACTOR of 1.5 multiplied against each model's peak_memory_gib value. Only models whose adjusted memory footprint fits within the available device memory appear as selectable options. The recommended_tiers field further categorizes compatibility into Modern (Apple Silicon M3+), Legacy (M1/M2), Intel, or Unknown classifications.
Complete Catalog of Supported Models
The MTPLX backend officially supports 18+ model configurations spanning the Qwen family and specialized architectures. Each model exposes an MTP head enabling multi-token prediction, though the system also supports pure autoregressive models via the ar_only flag or --no-mtp command-line option.
| Model ID | Display Name | Precision | Peak Memory | Tier |
|---|---|---|---|---|
qwen35-4b-optimized-speed |
Qwen 3.5 4B Optimized Speed | 4-bit | 2.86 GiB | Modern |
qwen35-4b-optimized-quality |
Qwen 3.5 4B Optimized Quality | 8-bit | 4.75 GiB | Modern |
qwen35-9b-optimized-speed |
Qwen 3.5 9B Optimized Speed | 6-bit | 10.0 GiB | Modern |
qwen35-9b-optimized-speed-fp16 |
Qwen 3.5 9B Optimized Speed FP16 | FP16 | 10.5 GiB | Legacy |
qwen36-35b-a3b-optimized-speed |
Qwen 3.6 35B-A3B Optimized Speed | 4-bit | 28.0 GiB | Modern |
qwen36-35b-a3b-optimized-speed-fp16 |
Qwen 3.6 35B-A3B Optimized Speed FP16 | FP16 | 28.5 GiB | Legacy |
qwen36-35b-a3b-optimized-balance |
Qwen 3.6 35B-A3B Optimized Balance | 6-bit | 32.0 GiB | Modern |
qwen38-27b-bare-speed |
Qwen 3.8 27B Bare Speed | BF16 + MTP side-car | 20.0 GiB | Modern |
qwen38-27b-optimized-speed |
Qwen 3.8 27B Optimized Speed | 4-bit dynamic | 25.0 GiB | Modern |
qwen38-27b-optimized-quality |
Qwen 3.8 27B Optimized Quality | 8-bit dynamic | 33.0 GiB | Modern |
qwen38-27b-bare-speed-fp16 |
Qwen 3.8 27B Bare Speed FP16 | FP16 | 20.0 GiB | Legacy |
qwen38-27b-optimized-speed-fp16 |
Qwen 3.8 27B Optimized Speed FP16 | 4-bit + FP16 | 25.0 GiB | Legacy |
flash-next-bare-speed |
Qwen 3.8 Flash-Next Bare Speed | 4-bit flat | 78.0 GiB | Modern |
flash-next-optimized-speed |
Qwen 3.8 Flash-Next Optimized Speed | 4-bit + 8-bit attention | 87.0 GiB | Modern |
optimized-speed-v2 |
Qwen 3.6 27B Optimized Speed V2 | Hybrid 4-bit | 21.5 GiB | Modern |
laguna |
Laguna (custom MLX) | Native MLX BF16 | ~12 GiB | Modern |
deepseek_v4 |
DeepSeek V4 | BF16 + MTP draft lane | ~40 GiB | Modern |
hy_v3 |
HY-V3 (vendored) | BF16 | ~14 GiB | Modern |
The deepseek_v4 implementation specifically utilizes an MTP draft lane for speculative decoding, while the laguna model represents a custom native MLX architecture optimized for the MTPLX backend.
Loading Models from the Catalog
Developers interact with supported models through the load_model function in mtplx/hf_loader.py, which patches standard MLX model classes with the MTP implementation from mtplx/mtp_patch.py.
List compatible models based on device constraints:
from mtplx.model_catalog import OFFICIAL_CATALOG
def list_supported_models():
"""Return models fitting current device memory."""
safety_factor = 1.5
device_memory = get_device_memory_gib() # MTPLX runtime utility
supported = [
m for m in OFFICIAL_CATALOG
if m.peak_memory_gib * safety_factor < device_memory
]
for model in supported:
print(f"{model.display_name}: {model.hf_model_id}")
Initialize a specific model with MTP enabled:
from mtplx.hf_loader import load_model
from mtplx.model_catalog import OFFICIAL_CATALOG
def load_optimized_qwen():
"""Load the Qwen 3.5 4B Optimized Speed variant."""
model_id = "qwen35-4b-optimized-speed"
entry = next(m for m in OFFICIAL_CATALOG if m.id == model_id)
# load_model patches the class to expose MTP functionality
model, config = load_model(
entry.hf_model_id,
get_model_classes=lambda cfg: (Model, ModelArgs)
)
return model
Backend Architecture Components
The MTPLX backend architecture supporting these models consists of several interconnected modules:
mtplx/model_catalog.py: DefinesCatalogModeldataclass andOFFICIAL_CATALOGtuple containing all supported model metadata.mtplx/hf_loader.py: Handles Hugging Face model retrieval and class patching to inject MTP capabilities.mtplx/mtp_patch.py: Core logic that modifies standard MLX model forward passes to enable multi-token prediction.mtplx/models/qwen4_exp.py: Implementation of Qwen 4-generation text and vision models.mtplx/models/laguna.py: Custom Laguna model definition for native MLX execution.mtplx/models/deepseek_v4.py: Speculative-draft DeepSeek V4 variant with MTP draft lanes.mtplx/vendored_hy_v3.py: Vendored HY-V3 model implementation for benchmarking purposes.
Summary
- The MTPLX backend architecture supports 18+ quantized model variants stored in
OFFICIAL_CATALOGinsidemtplx/model_catalog.py. - Supported families include Qwen 3.5/3.6/3.8 (4B to 27B parameters), Flash-Next, Laguna, DeepSeek V4, and HY-V3.
- Models are filtered at runtime using a 1.5x memory safety factor against
peak_memory_gibvalues and categorized by Modern (M3+) or Legacy (M1/M2) hardware tiers. - All catalog models feature MLX-compatible MTP heads for multi-token prediction, with fallback support for pure autoregressive mode via the
ar_onlyflag or--no-mtpoption. - The loading pipeline in
mtplx/hf_loader.pyautomatically patches model classes to expose MTPLX-specific functionality.
Frequently Asked Questions
Which hardware tiers does MTPLX support?
MTPLX categorizes devices into Modern (Apple Silicon M3 and later), Legacy (M1/M2 chips requiring FP16 fallbacks), Intel processors, and Unknown classifications. The recommended_tiers field in each CatalogModel entry determines which hardware can safely run specific model variants without performance degradation or out-of-memory errors.
Can I run models not listed in the OFFICIAL_CATALOG?
While the backend architecture is optimized for catalog models, MTPLX can load arbitrary MLX-compatible models using the --no-mtp flag. However, these non-catalog models lack the multi-token prediction head and must operate in pure autoregressive (AR) mode, bypassing the mtp_patch.py injection system entirely.
How does the memory safety factor affect model availability?
The backend multiplies each model's peak_memory_gib value by a global MEMORY_SAFETY_FACTOR of 1.5 to calculate a safe memory estimate. If this adjusted value exceeds available device memory as reported by get_device_memory_gib(), the model is filtered from the user interface. This prevents out-of-memory crashes during inference when accounting for MTP overhead and activation caching.
What is the difference between "Optimized Speed" and "Optimized Quality" variants?
Optimized Speed variants typically use 4-bit or 6-bit quantization to minimize latency and memory footprint (e.g., 2.86 GiB for Qwen 3.5 4B), while Optimized Quality variants employ 8-bit quantization or BF16 precision for higher fidelity at the cost of increased memory usage (up to 33 GiB for Qwen 3.8 27B). The catalog explicitly defines these trade-offs in the quantization and peak_memory_gib fields of each CatalogModel entry.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →