# DeepSeek V3 MTP and GLM MoE DSA Support in MTPLX: Complete Implementation Guide

> Explore the complete implementation of DeepSeek V3 MTP and GLM MoE DSA support in MTPLX. Learn how MTPLX automatically injects MTP heads for enhanced performance.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: deep-dive
- Published: 2026-09-04

---

**MTPLX provides full Multi-Token Prediction (MTP) support for both the DeepSeek V3 family and the GLM MoE DSA family, automatically detecting compatible configurations and injecting MTP heads at runtime when `num_nextn_predict_layers` is greater than zero.**

The MTPLX library (youssofal/MTPLX) implements comprehensive **Multi-Token Prediction** capabilities for modern mixture-of-experts architectures. Both **DeepSeek V3 MTP** and **GLM MoE DSA** models are fully supported through specialized detection logic in [`mtplx/deepseek_mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/deepseek_mtp_patch.py) and [`mtplx/glm_mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/glm_mtp_patch.py), combined with runtime weight rewriting mechanisms that validate checkpoint integrity before injection.

## Configuration Detection for MTP Support

MTPLX employs distinct configuration predicates to identify whether a model checkpoint requires MTP processing. These functions inspect the `model_type` field and the `num_nextn_predict_layers` parameter to determine eligibility.

### DeepSeek V3 Detection (`is_deepseek_mtp_config`)

In **[`mtplx/deepseek_mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/deepseek_mtp_patch.py)**, the `is_deepseek_mtp_config` function (lines L14-L38) validates MTP eligibility by checking if the `model_type` is one of `{"deepseek_v3", "deepseek_v32", "glm_moe_dsa"}` and that `num_nextn_predict_layers` is greater than zero. This ensures that only models explicitly configured for next-token prediction receive the MTP treatment.

### GLM MoE DSA Detection (`is_glm_mtp_config`)

For the GLM family, **[`mtplx/glm_mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/glm_mtp_patch.py)** contains the `is_glm_mtp_config` function (lines L14-L36), which recognizes model types `{"glm4_moe", "glm4_moe_lite"}` when paired with a positive `num_nextn_predict_layers` value. This detection mechanism separates standard MoE configurations from those requiring DSA (Dynamic Speculative Allocation) MTP support.

## Runtime Injection Mechanisms

Once detected, MTPLX rewrites checkpoint weights and attaches **MTPHead** instances through specialized injector functions dispatched from the central runtime.

### DeepSeek V3 Weight Rewriting (`inject_deepseek_mtp_support`)

The `inject_deepseek_mtp_support` function, called from **[`mtplx/runtime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime.py)** (lines L42-L55), creates an `MTPHead` on the model only when the checkpoint contains the required `mtp.*` tensors. This function internally calls `_rewrite_deepseek_mtp_weights` to transform the native DeepSeek weight layout into the MTPLX-expected format before attachment.

### GLM MoE DSA Weight Rewriting (`inject_glm_mtp_support`)

For GLM models, **[`mtplx/glm_mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/glm_mtp_patch.py)** provides `inject_glm_mtp_support` (lines L307-L322), which first invokes `_rewrite_glm_mtp_weights` to convert checkpoint tensors to the internal MTP format, then attaches the `MTPHead`. This two-step process ensures compatibility between GLM's native MoE structure and MTPLX's speculative decoding pipeline.

### Central Dispatch Logic ([`mtplx/runtime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime.py))

The **[`mtplx/runtime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime.py)** file serves as the central dispatcher (lines L718-L747), routing to either `inject_deepseek_mtp_support` or `inject_glm_mtp_support` based on the configuration predicates. This unified entry point guarantees that models declaring non-zero `num_nextn_predict_layers` are automatically equipped with MTP capabilities without manual intervention.

## Implementing MTP in Your Applications

MTPLX exposes MTP functionality through the standard `load_model` interface. The library automatically handles weight rewriting and head injection based on the provided configuration.

### Loading DeepSeek V3 with MTP

```python
from mtplx import load_model, MTPContract

model_path = "/path/to/deepseek_v3"
config = {
    "model_type": "deepseek_v3",
    "num_nextn_predict_layers": 1,  # >0 enables MTP

}
contract = MTPContract()

model = load_model(model_path, config, contract=contract)

# The model now has a `model.mtp` attribute (MTPHead) ready for speculative decoding

```

### Loading GLM-4 MoE-Lite with MTP

```python
from mtplx import load_model, MTPContract

model_path = "/path/to/glm4_moe_lite"
config = {
    "model_type": "glm4_moe_lite",
    "num_nextn_predict_layers": 1,
}
contract = MTPContract()

model = load_model(model_path, config, contract=contract)

# `model.mtp` is injected and the weight file is rewritten on-the-fly

```

## Test Coverage and Payload Validation

MTPLX validates MTP implementations through extensive test suites that verify both successful injection and proper rejection of incomplete checkpoints.

### DeepSeek V3 Test Suite

The **[`tests/test_deepseek_v4_mtp.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_deepseek_v4_mtp.py)** file validates that payload guards correctly recognize complete DeepSeek MTP payloads and that injection succeeds only when required `mtp.*` tensors are present. Additional coverage in **[`tests/test_runtime_deepseek_v4_o_lora.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_runtime_deepseek_v4_o_lora.py)** and **[`tests/test_deepseek_v4_spec.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_deepseek_v4_spec.py)** confirms runtime compatibility and speculative decoding behavior.

### GLM MoE DSA Test Suite

For GLM models, **[`tests/test_public_cli.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_public_cli.py)**, **[`tests/test_forge_cli.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_forge_cli.py)**, and **[`tests/test_artifacts.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_artifacts.py)** contain assertions such as `assert "glm4-moe-lite-mtp" in payload["verified_runtime_arch_ids"]` (lines L5973-L5999 and L1156-L1174), confirming that the runtime contract correctly generates verified architecture IDs for MTP variants.

## Summary

- **MTPLX fully supports DeepSeek V3 MTP and GLM MoE DSA MTP** through automatic configuration detection and runtime injection.
- **Detection relies on `model_type` and `num_nextn_predict_layers`** in [`mtplx/deepseek_mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/deepseek_mtp_patch.py) and [`mtplx/glm_mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/glm_mtp_patch.py).
- **Weight rewriting** transforms native checkpoint formats to MTPLX-internal layouts before `MTPHead` attachment.
- **Central dispatch** in [`mtplx/runtime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime.py#L718-L747) routes to the appropriate injector based on model family.
- **Comprehensive test coverage** validates payload completeness, weight mapping correctness, and runtime contract generation.

## Frequently Asked Questions

### Which model types support MTP in MTPLX?

MTPLX supports `deepseek_v3`, `deepseek_v32`, and `glm_moe_dsa` through the DeepSeek patch, plus `glm4_moe` and `glm4_moe_lite` through the GLM patch. All require `num_nextn_predict_layers > 0` in their configuration.

### How does MTPLX handle weight format conversion for MTP models?

MTPLX uses `_rewrite_deepseek_mtp_weights` for DeepSeek V3 models and `_rewrite_glm_mtp_weights` for GLM MoE models. These functions transform native checkpoint tensors into the internal MTP format expected by the `MTPHead` before runtime injection.

### What configuration parameter enables MTP support?

The `num_nextn_predict_layers` parameter controls MTP activation. When this integer is greater than zero, the detection predicates `is_deepseek_mtp_config` or `is_glm_mtp_config` return `True`, triggering automatic MTP head injection during model loading.

### Where is the MTP head attached in the model architecture?

The MTP head is attached as a `model.mtp` attribute (an instance of `MTPHead`) after the injection functions `inject_deepseek_mtp_support` or `inject_glm_mtp_support` complete their execution, making it available for speculative decoding operations immediately after `load_model` returns.