DeepSeek V3 MTP and GLM MoE DSA Support in MTPLX: Complete Implementation Guide

MTPLX provides full Multi-Token Prediction (MTP) support for both the DeepSeek V3 family and the GLM MoE DSA family, automatically detecting compatible configurations and injecting MTP heads at runtime when num_nextn_predict_layers is greater than zero.

The MTPLX library (youssofal/MTPLX) implements comprehensive Multi-Token Prediction capabilities for modern mixture-of-experts architectures. Both DeepSeek V3 MTP and GLM MoE DSA models are fully supported through specialized detection logic in mtplx/deepseek_mtp_patch.py and mtplx/glm_mtp_patch.py, combined with runtime weight rewriting mechanisms that validate checkpoint integrity before injection.

Configuration Detection for MTP Support

MTPLX employs distinct configuration predicates to identify whether a model checkpoint requires MTP processing. These functions inspect the model_type field and the num_nextn_predict_layers parameter to determine eligibility.

DeepSeek V3 Detection (is_deepseek_mtp_config)

In mtplx/deepseek_mtp_patch.py, the is_deepseek_mtp_config function (lines L14-L38) validates MTP eligibility by checking if the model_type is one of {"deepseek_v3", "deepseek_v32", "glm_moe_dsa"} and that num_nextn_predict_layers is greater than zero. This ensures that only models explicitly configured for next-token prediction receive the MTP treatment.

GLM MoE DSA Detection (is_glm_mtp_config)

For the GLM family, mtplx/glm_mtp_patch.py contains the is_glm_mtp_config function (lines L14-L36), which recognizes model types {"glm4_moe", "glm4_moe_lite"} when paired with a positive num_nextn_predict_layers value. This detection mechanism separates standard MoE configurations from those requiring DSA (Dynamic Speculative Allocation) MTP support.

Runtime Injection Mechanisms

Once detected, MTPLX rewrites checkpoint weights and attaches MTPHead instances through specialized injector functions dispatched from the central runtime.

DeepSeek V3 Weight Rewriting (inject_deepseek_mtp_support)

The inject_deepseek_mtp_support function, called from mtplx/runtime.py (lines L42-L55), creates an MTPHead on the model only when the checkpoint contains the required mtp.* tensors. This function internally calls _rewrite_deepseek_mtp_weights to transform the native DeepSeek weight layout into the MTPLX-expected format before attachment.

GLM MoE DSA Weight Rewriting (inject_glm_mtp_support)

For GLM models, mtplx/glm_mtp_patch.py provides inject_glm_mtp_support (lines L307-L322), which first invokes _rewrite_glm_mtp_weights to convert checkpoint tensors to the internal MTP format, then attaches the MTPHead. This two-step process ensures compatibility between GLM's native MoE structure and MTPLX's speculative decoding pipeline.

Central Dispatch Logic (mtplx/runtime.py)

The mtplx/runtime.py file serves as the central dispatcher (lines L718-L747), routing to either inject_deepseek_mtp_support or inject_glm_mtp_support based on the configuration predicates. This unified entry point guarantees that models declaring non-zero num_nextn_predict_layers are automatically equipped with MTP capabilities without manual intervention.

Implementing MTP in Your Applications

MTPLX exposes MTP functionality through the standard load_model interface. The library automatically handles weight rewriting and head injection based on the provided configuration.

Loading DeepSeek V3 with MTP

from mtplx import load_model, MTPContract

model_path = "/path/to/deepseek_v3"
config = {
    "model_type": "deepseek_v3",
    "num_nextn_predict_layers": 1,  # >0 enables MTP

}
contract = MTPContract()

model = load_model(model_path, config, contract=contract)

# The model now has a `model.mtp` attribute (MTPHead) ready for speculative decoding

Loading GLM-4 MoE-Lite with MTP

from mtplx import load_model, MTPContract

model_path = "/path/to/glm4_moe_lite"
config = {
    "model_type": "glm4_moe_lite",
    "num_nextn_predict_layers": 1,
}
contract = MTPContract()

model = load_model(model_path, config, contract=contract)

# `model.mtp` is injected and the weight file is rewritten on-the-fly

Test Coverage and Payload Validation

MTPLX validates MTP implementations through extensive test suites that verify both successful injection and proper rejection of incomplete checkpoints.

DeepSeek V3 Test Suite

The tests/test_deepseek_v4_mtp.py file validates that payload guards correctly recognize complete DeepSeek MTP payloads and that injection succeeds only when required mtp.* tensors are present. Additional coverage in tests/test_runtime_deepseek_v4_o_lora.py and tests/test_deepseek_v4_spec.py confirms runtime compatibility and speculative decoding behavior.

GLM MoE DSA Test Suite

For GLM models, tests/test_public_cli.py, tests/test_forge_cli.py, and tests/test_artifacts.py contain assertions such as assert "glm4-moe-lite-mtp" in payload["verified_runtime_arch_ids"] (lines L5973-L5999 and L1156-L1174), confirming that the runtime contract correctly generates verified architecture IDs for MTP variants.

Summary

  • MTPLX fully supports DeepSeek V3 MTP and GLM MoE DSA MTP through automatic configuration detection and runtime injection.
  • Detection relies on model_type and num_nextn_predict_layers in mtplx/deepseek_mtp_patch.py and mtplx/glm_mtp_patch.py.
  • Weight rewriting transforms native checkpoint formats to MTPLX-internal layouts before MTPHead attachment.
  • Central dispatch in mtplx/runtime.py routes to the appropriate injector based on model family.
  • Comprehensive test coverage validates payload completeness, weight mapping correctness, and runtime contract generation.

Frequently Asked Questions

Which model types support MTP in MTPLX?

MTPLX supports deepseek_v3, deepseek_v32, and glm_moe_dsa through the DeepSeek patch, plus glm4_moe and glm4_moe_lite through the GLM patch. All require num_nextn_predict_layers > 0 in their configuration.

How does MTPLX handle weight format conversion for MTP models?

MTPLX uses _rewrite_deepseek_mtp_weights for DeepSeek V3 models and _rewrite_glm_mtp_weights for GLM MoE models. These functions transform native checkpoint tensors into the internal MTP format expected by the MTPHead before runtime injection.

What configuration parameter enables MTP support?

The num_nextn_predict_layers parameter controls MTP activation. When this integer is greater than zero, the detection predicates is_deepseek_mtp_config or is_glm_mtp_config return True, triggering automatic MTP head injection during model loading.

Where is the MTP head attached in the model architecture?

The MTP head is attached as a model.mtp attribute (an instance of MTPHead) after the injection functions inject_deepseek_mtp_support or inject_glm_mtp_support complete their execution, making it available for speculative decoding operations immediately after load_model returns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →