MTPLX Experimental Backends: Architecture Support and Enablement Guide

MTPLX supports eight experimental backends—including DeepSeek V3-MTP, DeepSeek V4 (AR), and GLM-MTP—that are gated behind the --experimental-mtp-cohorts flag and defined in mtplx/backends/registry.py.

The MTPLX repository extends high-performance inference to specialized transformer architectures through its modular backend system. These experimental backends provide native support for multi-token prediction (MTP) and auto-regressive (AR) model families that require explicit opt-in for testing and development.

Overview of MTPLX Experimental Backends

MTPLX categorizes experimental backends by their support_level metadata in the architecture-compatibility registry. According to the source code in mtplx/backends/registry.py, backends with support_level strings beginning with experimental- represent cutting-edge implementations that are not enabled by default.

The framework identifies three distinct experimental support tiers:

  • experimental-native-contract-gated: Backends that implement contract-gated multi-token prediction
  • experimental-native-ar-only: Backends restricted to auto-regressive generation without MTP heads
  • experimental-mlx-lm-ar-only: Backends for MLX-LM model families lacking native MTP support

Complete List of MTPLX Experimental Backends

The following architectures are available when experimental mode is activated:

Architecture Backend Module Support Level Description
DeepSeek V3-MTP deepseek_v3_mtp experimental-native-contract-gated Contract-gated MTP for DeepSeek-V3 families
DeepSeek V4 (AR) deepseek_v4 experimental-native-ar-only AR-only native backend for DeepSeek-V4-Flash models
GLM-MTP glm_mtp experimental-native-contract-gated Contract-gated MTP for GLM-MoE families
MiMo-MTP mimo_mtp experimental-native-contract-gated Contract-gated MTP for MiMo family
Nemotron-H-MTP nemotron_h_mtp experimental-native-contract-gated Contract-gated MTP for Nemotron-H models
Step3.5-MTP step3p5_mtp experimental-native-contract-gated Contract-gated MTP for Step-3.5 models
Hy-V3-MTP hy_v3_mtp experimental-native-contract-gated Contract-gated MTP for Hy-V3 models
MLX-LM-AR mlx_lm_ar experimental-mlx-lm-ar-only AR-only backend for Llama-AR, iQuest-AR, LFM2-MoE-AR, and other MLX-LM families

These entries map architecture IDs to their corresponding backend modules in mtplx/backends/registry.py at lines 146, 229, and 300.

How to Enable MTPLX Experimental Backends

Experimental backends are opt-in only and disabled by default. You must activate them via command-line flag or configuration file before the server will expose these architectures.

Command-Line Activation

Pass the --experimental-mtp-cohorts flag when starting the MTPLX server:

mtplx start --experimental-mtp-cohorts

Configuration File Method

Set the boolean flag permanently in your MTPLX configuration:

mtplx config set experimental_mtp_cohorts true

This updates the experimental_mtp_cohorts value in mtplx/config.py and persists across restarts.

Python Client Verification

Once enabled, the server exposes experimental backends through the standard API:

import requests
import json

resp = requests.post(
    "http://127.0.0.1:8000/v1/chat/completions",
    json={
        "model": "mtplx", 
        "messages": [{"role": "user", "content": "Hello"}]
    },
    headers={"Content-Type": "application/json"}
)
print(json.dumps(resp.json(), indent=2))

Technical Implementation Details

According to the MTPLX source code, experimental backend identification relies on string parsing in the compatibility registry.

In mtplx/backends/registry.py, the system checks the support_level field for the prefix experimental-. For example, line 300 defines the DeepSeek V4 backend with support_level="experimental-native-ar-only", while lines 229 and 146 handle the contract-gated and MLX-LM variants respectively.

The mtplx/cli.py file defines the --experimental-mtp-cohorts argument, which propagates to mtplx/config.py where the experimental_mtp_cohorts boolean controls scheduler behavior. When enabled, the scheduler enters mtp_cohort_experimental mode as documented in docs/concurrency.md, allowing the dispatcher to route requests to these specialized backends.

Summary

  • Eight experimental backends are defined in mtplx/backends/registry.py, including DeepSeek V3-MTP, DeepSeek V4 (AR), GLM-MTP, and MLX-LM-AR variants.
  • Three support tiers classify experimental capabilities: experimental-native-contract-gated, experimental-native-ar-only, and experimental-mlx-lm-ar-only.
  • Opt-in activation requires setting --experimental-mtp-cohorts via CLI or experimental_mtp_cohorts in configuration.
  • Developer preview status means these backends are intended for testing and early adoption, not production workloads.

Frequently Asked Questions

How do I enable experimental backends in MTPLX?

You must start the server with the --experimental-mtp-cohorts command-line flag or set experimental_mtp_cohorts to true in your MTPLX configuration file. This flag is defined in mtplx/cli.py and stored in mtplx/config.py.

What is the difference between contract-gated and AR-only experimental backends?

Contract-gated backends (experimental-native-contract-gated) implement multi-token prediction with contract verification for models like DeepSeek V3 and GLM-MTP. AR-only backends (experimental-native-ar-only or experimental-mlx-lm-ar-only) support auto-regressive generation only, lacking MTP heads, as seen in DeepSeek V4-Flash and MLX-LM model families.

Which model families are supported by the MLX-LM experimental backend?

The experimental-mlx-lm-ar-only backend in mlx_lm_ar supports Llama-AR, iQuest-AR, LFM2-MoE-AR, and other MLX-LM families that do not have native MTP capabilities, according to the registry at line 146 of mtplx/backends/registry.py.

Are MTPLX experimental backends production-ready?

No. As indicated by their experimental- prefix in the support level registry, these backends are not enabled by default and their performance is not guaranteed. They are designed for developers and early adopters testing cutting-edge model parallelism features.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →