Which Model Architectures Are Supported by MTPLX? Complete 2024 Support Matrix
MTPLX supports eight distinct model families—Qwen 3.5, Qwen 3.8, Qwen 4 Exp, DeepSeek V4, Nemotron H, GLM4 MoE Lite, Qwen 3 Next, and LLaMA—identified via model_type strings in config.json and mapped through mtplx/default_models.py and mtplx/commands/public.py according to the youssofal/MTPLX source code.
MTPLX is an MLX-based inference engine that routes prompts to specific model implementations based on architecture identifiers. Understanding which model architectures are currently supported by MTPLX is essential for loading compatible checkpoints and optimizing inference performance on Apple Silicon hardware.
Supported Model Architecture Families
The runtime recognizes models by reading the model_type field from a checkpoint's config.json and dispatching to the appropriate handler. The following architectures are fully supported in the current codebase.
Qwen 3.5 Series
The Qwen 3.5 family represents the primary optimized path for MTPLX inference, supporting multiple variants:
- Base:
qwen3_5— Standard dense transformer - MoE:
qwen3_5_moe— Mixture-of-Experts variant - MTP:
qwen3_5_mtp— Multi-Token Prediction architecture - Speed-Optimized:
qwen3_5with 6-bit quantization flags - Speed-Optimized FP16:
qwen3_5variants with forced FP16 precision
Qwen 3.8 Variants
MTPLX supports the Qwen 3.8 experimental line with distinct optimization targets:
- Optimized-Speed:
qwen3_8_optimized_speed— Latency-optimized kernels - Bare-Speed:
qwen3_8_bare_speed— Minimal overhead mode - Optimized-Quality:
qwen3_8_optimized_quality— Balanced accuracy/precision - FP16 Siblings: Each above variant supports a
*_fp16orfp16suffixedmodel_typefor full-precision inference
Qwen 4 Exp (Experimental)
Early-access architectures for next-generation inference:
- Flash-Next Variants:
qwen4_exp,qwen4_exp_text - Handled by
_model_config_is_qwen4_exp()inmtplx/commands/public.pyfor specialized routing logic
DeepSeek V4 and V3
Legacy and current DeepSeek architectures:
- DeepSeek V4:
deepseek_v4— Current recommended identifier - DeepSeek V3:
deepseek_v3— Backward compatibility for older checkpoints - Supports low-bit quantization variants through the same type identifiers
Nemotron H
Single-head transformer architecture:
- Identifier:
nemotron_h - Optimized for specific enterprise inference patterns
GLM4 MoE Lite
Lightweight mixture-of-experts implementation:
- Identifier:
glm4_moe_lite - Memory-efficient routing for constrained environments
Qwen 3 Next
Early-access rollout architecture:
- Identifier:
qwen3_next - Pre-production testing channel for upcoming Qwen features
LLaMA (Legacy)
Compatibility stub primarily used in test suites:
- Identifier:
llama - Minimal implementation for validation and CI testing
How MTPLX Maps Architectures to Implementations
The architecture detection logic resides in two critical files. The select_default_model() function in mtplx/commands/public.py queries hardware capabilities and returns a DefaultModelSelection dataclass containing the appropriate model path, HuggingFace ID, and precision settings.
Architecture-specific validation occurs through helper functions like _model_config_is_qwen4_exp(), which inspect configuration dictionaries to determine if a checkpoint belongs to the experimental Qwen 4 family. The public_model_id_for_ref() function translates local checkpoint paths into OpenAI-compatible model identifiers based on these detected types.
Practical Code Examples for Supported Architectures
Selecting the Default Model for Current Hardware
Use select_default_model() to automatically choose the optimal supported architecture based on available memory and compute:
from mtplx.commands.public import select_default_model
# Returns a DefaultModelSelection dataclass
default = select_default_model()
print(f"Chosen architecture: {default.display_name}")
print(f"Selection reason: {default.reason}")
Resolving Model Types from Checkpoints
Extract the model_type from config.json and resolve the public identifier:
import json
from pathlib import Path
from mtplx.commands.public import public_model_id_for_ref
model_dir = Path("/path/to/qwen3_5_moe")
config = json.loads((model_dir / "config.json").read_text())
model_type = config.get("model_type") # e.g., "qwen3_5_moe"
public_id = public_model_id_for_ref(model_dir)
print(f"Type {model_type} maps to public ID: {public_id}")
Overriding Precision Variants
Force specific precision modes using environment variables before initialization:
import os
from mtplx.commands.public import select_default_model
# Force FP16 variant selection
os.environ["MTPLX_DEFAULT_MODEL_VARIANT"] = "fp16"
default_fp16 = select_default_model()
print(default_fp16.display_name) # Outputs: "Qwen3.5 9B Optimized Speed FP16"
Verifying Default Model Status
Check if a local checkpoint is a verified default model:
from mtplx.default_models import is_verified_default_model_ref
model_path = "/my/models/qwen3_8_optimized_speed"
if is_verified_default_model_ref(model_path):
print("Verified default model - optimized paths available")
else:
print("Custom model - using generic inference pipeline")
Key Implementation Files
The architecture support matrix is defined in the following source locations:
mtplx/default_models.py: Contains the central registry of supportedmodel_typevalues, hardware-aware routing logic, and default model selection algorithmsmtplx/commands/public.py: Implements architecture detection helpers including_model_config_is_qwen4_exp()and thepublic_model_id_for_ref()mapping functionmtplx/vision/qwen3_vl_tower.py: Vision-tower implementation specific to Qwen 3.5 VL multimodal architecturestests/test_runtime_model_alias.py: Validation suite ensuring correctmodel_typealiasing for all supported architectures
Summary
- MTPLX supports eight primary architectures: Qwen 3.5/3.8/4 Exp, DeepSeek V4/V3, Nemotron H, GLM4 MoE Lite, Qwen 3 Next, and LLaMA
- Architecture detection relies on
model_typestrings read fromconfig.jsonand processed throughmtplx/default_models.py - Hardware-aware selection occurs via
select_default_model()inmtplx/commands/public.py - Variant control (speed vs. quality vs. FP16) is managed through environment variables and specific
model_typesuffixes - Experimental architectures like Qwen 4 Exp use specialized validation functions before routing to implementation handlers
Frequently Asked Questions
How does MTPLX determine which architecture a checkpoint uses?
MTPLX reads the model_type field from the checkpoint's config.json file. This string is then compared against internal registries in mtplx/default_models.py and validated through helper functions in mtplx/commands/public.py to determine the appropriate model class and inference pipeline.
Can I force MTPLX to use a specific precision variant like FP16?
Yes. Set the environment variable MTPLX_DEFAULT_MODEL_VARIANT to "fp16" before calling select_default_model(). Alternatively, specify the exact model_type variant (e.g., qwen3_8_optimized_speed with FP16 flags) in your configuration to override the hardware-default selection.
Are custom or fine-tuned models based on supported architectures compatible?
Custom models are compatible if they retain the base architecture's model_type identifier and structure. Use is_verified_default_model_ref() from mtplx/default_models.py to check if your specific checkpoint receives optimized default-model routing, or falls back to the generic compatible inference path.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →