Which Model Architectures Are Supported by MTPLX? Complete 2024 Support Matrix

MTPLX supports eight distinct model families—Qwen 3.5, Qwen 3.8, Qwen 4 Exp, DeepSeek V4, Nemotron H, GLM4 MoE Lite, Qwen 3 Next, and LLaMA—identified via model_type strings in config.json and mapped through mtplx/default_models.py and mtplx/commands/public.py according to the youssofal/MTPLX source code.

MTPLX is an MLX-based inference engine that routes prompts to specific model implementations based on architecture identifiers. Understanding which model architectures are currently supported by MTPLX is essential for loading compatible checkpoints and optimizing inference performance on Apple Silicon hardware.

Supported Model Architecture Families

The runtime recognizes models by reading the model_type field from a checkpoint's config.json and dispatching to the appropriate handler. The following architectures are fully supported in the current codebase.

Qwen 3.5 Series

The Qwen 3.5 family represents the primary optimized path for MTPLX inference, supporting multiple variants:

  • Base: qwen3_5 — Standard dense transformer
  • MoE: qwen3_5_moe — Mixture-of-Experts variant
  • MTP: qwen3_5_mtp — Multi-Token Prediction architecture
  • Speed-Optimized: qwen3_5 with 6-bit quantization flags
  • Speed-Optimized FP16: qwen3_5 variants with forced FP16 precision

Qwen 3.8 Variants

MTPLX supports the Qwen 3.8 experimental line with distinct optimization targets:

  • Optimized-Speed: qwen3_8_optimized_speed — Latency-optimized kernels
  • Bare-Speed: qwen3_8_bare_speed — Minimal overhead mode
  • Optimized-Quality: qwen3_8_optimized_quality — Balanced accuracy/precision
  • FP16 Siblings: Each above variant supports a *_fp16 or fp16 suffixed model_type for full-precision inference

Qwen 4 Exp (Experimental)

Early-access architectures for next-generation inference:

  • Flash-Next Variants: qwen4_exp, qwen4_exp_text
  • Handled by _model_config_is_qwen4_exp() in mtplx/commands/public.py for specialized routing logic

DeepSeek V4 and V3

Legacy and current DeepSeek architectures:

  • DeepSeek V4: deepseek_v4 — Current recommended identifier
  • DeepSeek V3: deepseek_v3 — Backward compatibility for older checkpoints
  • Supports low-bit quantization variants through the same type identifiers

Nemotron H

Single-head transformer architecture:

  • Identifier: nemotron_h
  • Optimized for specific enterprise inference patterns

GLM4 MoE Lite

Lightweight mixture-of-experts implementation:

  • Identifier: glm4_moe_lite
  • Memory-efficient routing for constrained environments

Qwen 3 Next

Early-access rollout architecture:

  • Identifier: qwen3_next
  • Pre-production testing channel for upcoming Qwen features

LLaMA (Legacy)

Compatibility stub primarily used in test suites:

  • Identifier: llama
  • Minimal implementation for validation and CI testing

How MTPLX Maps Architectures to Implementations

The architecture detection logic resides in two critical files. The select_default_model() function in mtplx/commands/public.py queries hardware capabilities and returns a DefaultModelSelection dataclass containing the appropriate model path, HuggingFace ID, and precision settings.

Architecture-specific validation occurs through helper functions like _model_config_is_qwen4_exp(), which inspect configuration dictionaries to determine if a checkpoint belongs to the experimental Qwen 4 family. The public_model_id_for_ref() function translates local checkpoint paths into OpenAI-compatible model identifiers based on these detected types.

Practical Code Examples for Supported Architectures

Selecting the Default Model for Current Hardware

Use select_default_model() to automatically choose the optimal supported architecture based on available memory and compute:

from mtplx.commands.public import select_default_model

# Returns a DefaultModelSelection dataclass

default = select_default_model()
print(f"Chosen architecture: {default.display_name}")
print(f"Selection reason: {default.reason}")

Resolving Model Types from Checkpoints

Extract the model_type from config.json and resolve the public identifier:

import json
from pathlib import Path
from mtplx.commands.public import public_model_id_for_ref

model_dir = Path("/path/to/qwen3_5_moe")
config = json.loads((model_dir / "config.json").read_text())
model_type = config.get("model_type")  # e.g., "qwen3_5_moe"

public_id = public_model_id_for_ref(model_dir)
print(f"Type {model_type} maps to public ID: {public_id}")

Overriding Precision Variants

Force specific precision modes using environment variables before initialization:

import os
from mtplx.commands.public import select_default_model

# Force FP16 variant selection

os.environ["MTPLX_DEFAULT_MODEL_VARIANT"] = "fp16"

default_fp16 = select_default_model()
print(default_fp16.display_name)  # Outputs: "Qwen3.5 9B Optimized Speed FP16"

Verifying Default Model Status

Check if a local checkpoint is a verified default model:

from mtplx.default_models import is_verified_default_model_ref

model_path = "/my/models/qwen3_8_optimized_speed"
if is_verified_default_model_ref(model_path):
    print("Verified default model - optimized paths available")
else:
    print("Custom model - using generic inference pipeline")

Key Implementation Files

The architecture support matrix is defined in the following source locations:

Summary

  • MTPLX supports eight primary architectures: Qwen 3.5/3.8/4 Exp, DeepSeek V4/V3, Nemotron H, GLM4 MoE Lite, Qwen 3 Next, and LLaMA
  • Architecture detection relies on model_type strings read from config.json and processed through mtplx/default_models.py
  • Hardware-aware selection occurs via select_default_model() in mtplx/commands/public.py
  • Variant control (speed vs. quality vs. FP16) is managed through environment variables and specific model_type suffixes
  • Experimental architectures like Qwen 4 Exp use specialized validation functions before routing to implementation handlers

Frequently Asked Questions

How does MTPLX determine which architecture a checkpoint uses?

MTPLX reads the model_type field from the checkpoint's config.json file. This string is then compared against internal registries in mtplx/default_models.py and validated through helper functions in mtplx/commands/public.py to determine the appropriate model class and inference pipeline.

Can I force MTPLX to use a specific precision variant like FP16?

Yes. Set the environment variable MTPLX_DEFAULT_MODEL_VARIANT to "fp16" before calling select_default_model(). Alternatively, specify the exact model_type variant (e.g., qwen3_8_optimized_speed with FP16 flags) in your configuration to override the hardware-default selection.

Are custom or fine-tuned models based on supported architectures compatible?

Custom models are compatible if they retain the base architecture's model_type identifier and structure. Use is_verified_default_model_ref() from mtplx/default_models.py to check if your specific checkpoint receives optimized default-model routing, or falls back to the generic compatible inference path.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →