# Which Model Architectures Are Supported by MTPLX? Complete 2024 Support Matrix

> Explore the MTPLX model architectures supported in 2024 including Qwen, DeepSeek, Nemotron, GLM4, and LLaMA. Discover the complete support matrix for MTPLX.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: api-reference
- Published: 2026-09-06

---

**MTPLX supports eight distinct model families—Qwen 3.5, Qwen 3.8, Qwen 4 Exp, DeepSeek V4, Nemotron H, GLM4 MoE Lite, Qwen 3 Next, and LLaMA—identified via `model_type` strings in [`config.json`](https://github.com/youssofal/MTPLX/blob/main/config.json) and mapped through [`mtplx/default_models.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/default_models.py) and [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) according to the youssofal/MTPLX source code.**

MTPLX is an MLX-based inference engine that routes prompts to specific model implementations based on architecture identifiers. Understanding which model architectures are currently supported by MTPLX is essential for loading compatible checkpoints and optimizing inference performance on Apple Silicon hardware.

## Supported Model Architecture Families

The runtime recognizes models by reading the `model_type` field from a checkpoint's [`config.json`](https://github.com/youssofal/MTPLX/blob/main/config.json) and dispatching to the appropriate handler. The following architectures are fully supported in the current codebase.

### Qwen 3.5 Series

The Qwen 3.5 family represents the primary optimized path for MTPLX inference, supporting multiple variants:

- **Base**: `qwen3_5` — Standard dense transformer
- **MoE**: `qwen3_5_moe` — Mixture-of-Experts variant
- **MTP**: `qwen3_5_mtp` — Multi-Token Prediction architecture
- **Speed-Optimized**: `qwen3_5` with 6-bit quantization flags
- **Speed-Optimized FP16**: `qwen3_5` variants with forced FP16 precision

### Qwen 3.8 Variants

MTPLX supports the Qwen 3.8 experimental line with distinct optimization targets:

- **Optimized-Speed**: `qwen3_8_optimized_speed` — Latency-optimized kernels
- **Bare-Speed**: `qwen3_8_bare_speed` — Minimal overhead mode
- **Optimized-Quality**: `qwen3_8_optimized_quality` — Balanced accuracy/precision
- **FP16 Siblings**: Each above variant supports a `*_fp16` or `fp16` suffixed `model_type` for full-precision inference

### Qwen 4 Exp (Experimental)

Early-access architectures for next-generation inference:

- **Flash-Next Variants**: `qwen4_exp`, `qwen4_exp_text`
- Handled by `_model_config_is_qwen4_exp()` in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) for specialized routing logic

### DeepSeek V4 and V3

Legacy and current DeepSeek architectures:

- **DeepSeek V4**: `deepseek_v4` — Current recommended identifier
- **DeepSeek V3**: `deepseek_v3` — Backward compatibility for older checkpoints
- Supports low-bit quantization variants through the same type identifiers

### Nemotron H

Single-head transformer architecture:

- **Identifier**: `nemotron_h`
- Optimized for specific enterprise inference patterns

### GLM4 MoE Lite

Lightweight mixture-of-experts implementation:

- **Identifier**: `glm4_moe_lite`
- Memory-efficient routing for constrained environments

### Qwen 3 Next

Early-access rollout architecture:

- **Identifier**: `qwen3_next`
- Pre-production testing channel for upcoming Qwen features

### LLaMA (Legacy)

Compatibility stub primarily used in test suites:

- **Identifier**: `llama`
- Minimal implementation for validation and CI testing

## How MTPLX Maps Architectures to Implementations

The architecture detection logic resides in two critical files. The `select_default_model()` function in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) queries hardware capabilities and returns a `DefaultModelSelection` dataclass containing the appropriate model path, HuggingFace ID, and precision settings.

Architecture-specific validation occurs through helper functions like `_model_config_is_qwen4_exp()`, which inspect configuration dictionaries to determine if a checkpoint belongs to the experimental Qwen 4 family. The `public_model_id_for_ref()` function translates local checkpoint paths into OpenAI-compatible model identifiers based on these detected types.

## Practical Code Examples for Supported Architectures

### Selecting the Default Model for Current Hardware

Use `select_default_model()` to automatically choose the optimal supported architecture based on available memory and compute:

```python
from mtplx.commands.public import select_default_model

# Returns a DefaultModelSelection dataclass

default = select_default_model()
print(f"Chosen architecture: {default.display_name}")
print(f"Selection reason: {default.reason}")

```

### Resolving Model Types from Checkpoints

Extract the `model_type` from [`config.json`](https://github.com/youssofal/MTPLX/blob/main/config.json) and resolve the public identifier:

```python
import json
from pathlib import Path
from mtplx.commands.public import public_model_id_for_ref

model_dir = Path("/path/to/qwen3_5_moe")
config = json.loads((model_dir / "config.json").read_text())
model_type = config.get("model_type")  # e.g., "qwen3_5_moe"

public_id = public_model_id_for_ref(model_dir)
print(f"Type {model_type} maps to public ID: {public_id}")

```

### Overriding Precision Variants

Force specific precision modes using environment variables before initialization:

```python
import os
from mtplx.commands.public import select_default_model

# Force FP16 variant selection

os.environ["MTPLX_DEFAULT_MODEL_VARIANT"] = "fp16"

default_fp16 = select_default_model()
print(default_fp16.display_name)  # Outputs: "Qwen3.5 9B Optimized Speed FP16"

```

### Verifying Default Model Status

Check if a local checkpoint is a verified default model:

```python
from mtplx.default_models import is_verified_default_model_ref

model_path = "/my/models/qwen3_8_optimized_speed"
if is_verified_default_model_ref(model_path):
    print("Verified default model - optimized paths available")
else:
    print("Custom model - using generic inference pipeline")

```

## Key Implementation Files

The architecture support matrix is defined in the following source locations:

- **[`mtplx/default_models.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/default_models.py)**: Contains the central registry of supported `model_type` values, hardware-aware routing logic, and default model selection algorithms
- **[`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py)**: Implements architecture detection helpers including `_model_config_is_qwen4_exp()` and the `public_model_id_for_ref()` mapping function
- **[`mtplx/vision/qwen3_vl_tower.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision/qwen3_vl_tower.py)**: Vision-tower implementation specific to Qwen 3.5 VL multimodal architectures
- **[`tests/test_runtime_model_alias.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_runtime_model_alias.py)**: Validation suite ensuring correct `model_type` aliasing for all supported architectures

## Summary

- **MTPLX supports eight primary architectures**: Qwen 3.5/3.8/4 Exp, DeepSeek V4/V3, Nemotron H, GLM4 MoE Lite, Qwen 3 Next, and LLaMA
- **Architecture detection relies on `model_type` strings** read from [`config.json`](https://github.com/youssofal/MTPLX/blob/main/config.json) and processed through [`mtplx/default_models.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/default_models.py)
- **Hardware-aware selection** occurs via `select_default_model()` in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py)
- **Variant control** (speed vs. quality vs. FP16) is managed through environment variables and specific `model_type` suffixes
- **Experimental architectures** like Qwen 4 Exp use specialized validation functions before routing to implementation handlers

## Frequently Asked Questions

### How does MTPLX determine which architecture a checkpoint uses?

MTPLX reads the `model_type` field from the checkpoint's [`config.json`](https://github.com/youssofal/MTPLX/blob/main/config.json) file. This string is then compared against internal registries in [`mtplx/default_models.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/default_models.py) and validated through helper functions in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) to determine the appropriate model class and inference pipeline.

### Can I force MTPLX to use a specific precision variant like FP16?

Yes. Set the environment variable `MTPLX_DEFAULT_MODEL_VARIANT` to `"fp16"` before calling `select_default_model()`. Alternatively, specify the exact `model_type` variant (e.g., `qwen3_8_optimized_speed` with FP16 flags) in your configuration to override the hardware-default selection.

### Are custom or fine-tuned models based on supported architectures compatible?

Custom models are compatible if they retain the base architecture's `model_type` identifier and structure. Use `is_verified_default_model_ref()` from [`mtplx/default_models.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/default_models.py) to check if your specific checkpoint receives optimized default-model routing, or falls back to the generic compatible inference path.