# MTPLX Model Backend Support Tiers: A Complete Guide to Backend Descriptors

> Explore MTPLX model backend support tiers and their capabilities. Learn about MTP pre-fill depth, tool-call policies, and hardware optimizations in this comprehensive guide.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: api-reference
- Published: 2026-09-04

---

**MTPLX organizes model backends into distinct support tiers defined in [`mtplx/backends/descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py), where each tier maps a backend identifier to specific capabilities like MTP pre-fill depth, tool-call policies, and hardware optimizations.**

The youssofal/MTPLX repository implements a tiered backend architecture to route models to optimized inference implementations. Understanding these **model backend support tiers** is essential for configuring which capabilities—such as multi-token prediction (MTP) depth or tool-calling modes—are available to specific model families. Each tier is represented by a backend descriptor that ties a backend identifier to the runtime capabilities required for execution.

## How Backend Support Tiers Work in MTPLX

The support tier system centers on backend descriptors stored in [`mtplx/backends/descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py). These descriptors act as capability contracts: they store metadata including the required **tool-prompt mode**, default sampling codecs, and hardware-specific flags (e.g., GPU-only execution requirements).

When the engine initializes a request, it resolves the backend ID via `descriptor_for_backend_id`—exposed publicly through [`mtplx/backends/__init__.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/__init__.py)—and applies the corresponding tier’s defaults. This resolution process guarantees that requests route to backends capable of fulfilling specific capabilities, such as MTP-style pre-fill or specialized attention routing.

## The Nine Core Model Backend Support Tiers

MTPLX defines nine primary support tiers, each optimized for distinct model architectures and inference patterns:

- **`native_mtp`** — The reference native MTP implementation that works with any model following the standard MTPLX API. This tier serves as the baseline for safe, universal inference but does not include model-specific optimizations.

- **`step3p5_mtp`** — Implements depth-one MTP support for prefill-only operations. This tier targets models requiring shallow MTP without full-length support.

- **`nemotron_h_mtp`** — Provides Nemotron-H specific MTP support with depth-one capabilities and specialized attention routing optimizations tailored to the Nemotron-H family.

- **`mimo_mtp`** — Offers depth-one MTP with multi-input-multi-output (MIMO) handling, designed for models that produce multiple output streams per request.

- **`gemma4_assistant`** — Includes draft-position handling and tool-prompt policy enforcement. This tier supports Gemma-4 models deployed as assistants with tool-calling capabilities.

- **`qwen3_next`** — Delivers the newest Qwen-3 features, including advanced token routing mechanisms specific to the Qwen-3-Next series.

- **`deepseek_mtp`** — Supports DeepSeek-V4 specific optimizations, including attention-island kernels and adaptive-width implementations.

- **`glm_mtp`** — Specialized for GLM-style model architectures, providing optimized kernels for GLM family inference.

- **`hy_v3_mtp`** — Integrates Hy-V3 architecture custom kernels for models built on the Hy-V3 specification.

## Resolving and Configuring Backend Tiers

Developers interact with support tiers programmatically through the backend resolution API. The following example demonstrates retrieving a descriptor by its backend ID:

```python

# Resolve a backend descriptor by its ID

from mtplx.backends import descriptor_for_backend_id

# Choose the Gemma-4 assistant backend

gemma4_descr = descriptor_for_backend_id("gemma4_assistant")
print(gemma4_descr.backend_id)        # → "gemma4_assistant"

print(gemma4_descr.required_tool_prompt_mode)  # → "hybrid"

```

For command-line interfaces, MTPLX applies backend-specific defaults using the `_apply_backend_serve_defaults` function in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py). This function automatically selects appropriate tiers based on model inspection data:

```python

# Apply backend-specific defaults to an argparse.Namespace

from mtplx.commands.public import _apply_backend_serve_defaults
args = argparse.Namespace(backend_id=None)
_apply_backend_serve_defaults(args, inspection)   # inspection = model inspection data

print(args.backend_id)                 # backend is auto-selected based on model

```

To force a specific tier for a model, use the registry acquisition pattern defined in [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py):

```python

# Example: forcing a specific tier for a model

from mtplx.backends import registry, _BackendSpec

spec = _BackendSpec(model_type="mpt", backend_name="step3p5_mtp")
with registry._acquire(spec) as backend:
    # backend is an instance of the Step-3.5 MTP class

    backend.prefill(...)               # uses depth-one MTP logic

```

## Backend Tier Metadata and Capabilities

Each descriptor stores critical metadata that determines runtime behavior. According to the implementation in [`mtplx/backends/descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py), descriptors track:

- **Tool-prompt modes**: Specifies whether the backend requires "hybrid" or other prompt formatting for tool calls.
- **Sampling codecs**: Default encoding strategies for token generation.
- **Hardware flags**: Execution environment requirements, such as GPU-only constraints or specific kernel availability.

These metadata fields ensure that the engine routes requests to hardware-compatible backends with appropriate policy enforcement.

## Summary

- MTPLX defines **nine core support tiers** in [`mtplx/backends/descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py), ranging from `native_mtp` (baseline) to model-specific tiers like `deepseek_mtp` and `gemma4_assistant`.
- **Backend descriptors** store capability metadata including tool-prompt modes, sampling codecs, and hardware requirements.
- The resolution API `descriptor_for_backend_id` in [`mtplx/backends/__init__.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/__init__.py) connects model requests to appropriate tiers.
- **Registry acquisition** via [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py) allows explicit instantiation of specific backend classes for controlled inference pipelines.
- **Command-line defaults** are applied through `_apply_backend_serve_defaults` in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py), enabling automatic tier selection based on model inspection.

## Frequently Asked Questions

### What is the difference between native_mtp and step3p5_mtp support tiers?

The **`native_mtp`** tier provides a universal baseline implementation compatible with all MTPLX-compliant models but lacks specialized optimizations. The **`step3p5_mtp`** tier specifically implements depth-one multi-token prediction for prefill-only operations, offering targeted performance improvements for models that support shallow MTP contexts but not full-length prediction.

### How does MTPLX select the appropriate backend tier for a model?

MTPLX selects tiers through a resolution pipeline: the engine calls `descriptor_for_backend_id` with a backend identifier, or uses `_apply_backend_serve_defaults` (defined in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py)) to auto-select based on model inspection data. The registry in [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py) then instantiates the concrete backend class associated with the resolved descriptor.

### What hardware-specific flags are stored in backend descriptors?

Backend descriptors store flags indicating execution environment requirements, such as **GPU-only execution** constraints or availability of specific optimized kernels (e.g., attention-island kernels for DeepSeek models or adaptive-width kernels for Nemotron-H). These flags ensure the engine routes requests to compatible hardware configurations.

### Where are the concrete implementations of each support tier located?

Each support tier’s runtime behavior is implemented in dedicated backend modules within `mtplx/backends/`. For example, [`gemma4_assistant.py`](https://github.com/youssofal/MTPLX/blob/main/gemma4_assistant.py) implements the Gemma-4 assistant tier’s draft-position handling, while [`step3p5_mtp.py`](https://github.com/youssofal/MTPLX/blob/main/step3p5_mtp.py) contains the depth-one MTP logic. The central registry in [`mtplx/backends/descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py) maps identifiers to these concrete implementations.