# MTPLX Experimental Backends: Architecture Support and Enablement Guide

> Discover the eight experimental backends MTPLX supports, including DeepSeek V3-MTP and GLM-MTP. Explore architecture support and enablement with this comprehensive guide.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: architecture
- Published: 2026-09-05

---

**MTPLX supports eight experimental backends—including DeepSeek V3-MTP, DeepSeek V4 (AR), and GLM-MTP—that are gated behind the `--experimental-mtp-cohorts` flag and defined in [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py).**

The MTPLX repository extends high-performance inference to specialized transformer architectures through its modular backend system. These **experimental backends** provide native support for multi-token prediction (MTP) and auto-regressive (AR) model families that require explicit opt-in for testing and development.

## Overview of MTPLX Experimental Backends

MTPLX categorizes experimental backends by their `support_level` metadata in the architecture-compatibility registry. According to the source code in [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py), backends with `support_level` strings beginning with `experimental-` represent cutting-edge implementations that are not enabled by default.

The framework identifies three distinct experimental support tiers:

- **`experimental-native-contract-gated`**: Backends that implement contract-gated multi-token prediction
- **`experimental-native-ar-only`**: Backends restricted to auto-regressive generation without MTP heads  
- **`experimental-mlx-lm-ar-only`**: Backends for MLX-LM model families lacking native MTP support

## Complete List of MTPLX Experimental Backends

The following architectures are available when experimental mode is activated:

| Architecture | Backend Module | Support Level | Description |
|--------------|----------------|---------------|-------------|
| **DeepSeek V3-MTP** | `deepseek_v3_mtp` | `experimental-native-contract-gated` | Contract-gated MTP for DeepSeek-V3 families |
| **DeepSeek V4 (AR)** | `deepseek_v4` | `experimental-native-ar-only` | AR-only native backend for DeepSeek-V4-Flash models |
| **GLM-MTP** | `glm_mtp` | `experimental-native-contract-gated` | Contract-gated MTP for GLM-MoE families |
| **MiMo-MTP** | `mimo_mtp` | `experimental-native-contract-gated` | Contract-gated MTP for MiMo family |
| **Nemotron-H-MTP** | `nemotron_h_mtp` | `experimental-native-contract-gated` | Contract-gated MTP for Nemotron-H models |
| **Step3.5-MTP** | `step3p5_mtp` | `experimental-native-contract-gated` | Contract-gated MTP for Step-3.5 models |
| **Hy-V3-MTP** | `hy_v3_mtp` | `experimental-native-contract-gated` | Contract-gated MTP for Hy-V3 models |
| **MLX-LM-AR** | `mlx_lm_ar` | `experimental-mlx-lm-ar-only` | AR-only backend for Llama-AR, iQuest-AR, LFM2-MoE-AR, and other MLX-LM families |

These entries map architecture IDs to their corresponding backend modules in [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py) at lines 146, 229, and 300.

## How to Enable MTPLX Experimental Backends

Experimental backends are **opt-in only** and disabled by default. You must activate them via command-line flag or configuration file before the server will expose these architectures.

### Command-Line Activation

Pass the `--experimental-mtp-cohorts` flag when starting the MTPLX server:

```bash
mtplx start --experimental-mtp-cohorts

```

### Configuration File Method

Set the boolean flag permanently in your MTPLX configuration:

```bash
mtplx config set experimental_mtp_cohorts true

```

This updates the `experimental_mtp_cohorts` value in [`mtplx/config.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/config.py) and persists across restarts.

### Python Client Verification

Once enabled, the server exposes experimental backends through the standard API:

```python
import requests
import json

resp = requests.post(
    "http://127.0.0.1:8000/v1/chat/completions",
    json={
        "model": "mtplx", 
        "messages": [{"role": "user", "content": "Hello"}]
    },
    headers={"Content-Type": "application/json"}
)
print(json.dumps(resp.json(), indent=2))

```

## Technical Implementation Details

According to the MTPLX source code, experimental backend identification relies on string parsing in the compatibility registry.

In [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py), the system checks the `support_level` field for the prefix `experimental-`. For example, line 300 defines the DeepSeek V4 backend with `support_level="experimental-native-ar-only"`, while lines 229 and 146 handle the contract-gated and MLX-LM variants respectively.

The [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) file defines the `--experimental-mtp-cohorts` argument, which propagates to [`mtplx/config.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/config.py) where the `experimental_mtp_cohorts` boolean controls scheduler behavior. When enabled, the scheduler enters `mtp_cohort_experimental` mode as documented in [`docs/concurrency.md`](https://github.com/youssofal/MTPLX/blob/main/docs/concurrency.md), allowing the dispatcher to route requests to these specialized backends.

## Summary

- **Eight experimental backends** are defined in [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py), including DeepSeek V3-MTP, DeepSeek V4 (AR), GLM-MTP, and MLX-LM-AR variants.
- **Three support tiers** classify experimental capabilities: `experimental-native-contract-gated`, `experimental-native-ar-only`, and `experimental-mlx-lm-ar-only`.
- **Opt-in activation** requires setting `--experimental-mtp-cohorts` via CLI or `experimental_mtp_cohorts` in configuration.
- **Developer preview status** means these backends are intended for testing and early adoption, not production workloads.

## Frequently Asked Questions

### How do I enable experimental backends in MTPLX?

You must start the server with the `--experimental-mtp-cohorts` command-line flag or set `experimental_mtp_cohorts` to `true` in your MTPLX configuration file. This flag is defined in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) and stored in [`mtplx/config.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/config.py).

### What is the difference between contract-gated and AR-only experimental backends?

**Contract-gated backends** (`experimental-native-contract-gated`) implement multi-token prediction with contract verification for models like DeepSeek V3 and GLM-MTP. **AR-only backends** (`experimental-native-ar-only` or `experimental-mlx-lm-ar-only`) support auto-regressive generation only, lacking MTP heads, as seen in DeepSeek V4-Flash and MLX-LM model families.

### Which model families are supported by the MLX-LM experimental backend?

The `experimental-mlx-lm-ar-only` backend in `mlx_lm_ar` supports Llama-AR, iQuest-AR, LFM2-MoE-AR, and other MLX-LM families that do not have native MTP capabilities, according to the registry at line 146 of [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py).

### Are MTPLX experimental backends production-ready?

No. As indicated by their `experimental-` prefix in the support level registry, these backends are not enabled by default and their performance is not guaranteed. They are designed for developers and early adopters testing cutting-edge model parallelism features.