Supported MTP Paths for Nemotron-H in MTPLX: Configuration Constraints and Valid Patterns

MTPLX enforces strict constraints on Nemotron-H Multi-Token Prediction, allowing only single-layer configurations with concat_order="embedding_hidden", mtp_depth=1, and patterns containing exclusively * (attention) and E (MoE) characters.

The MTPLX library provides a native backend specifically designed for Nemotron-H model architectures. Understanding the supported MTP paths for Nemotron-H in MTPLX is critical for configuring inference pipelines correctly, as the implementation deviates from standard MTP configurability by imposing hard limits on concatenation order, prediction depth, and layer composition.

MTP Concatenation Order Requirements

In mtplx/backends/nemotron_h_mtp.py, the mtp_forward method validates the concat_order parameter against a restricted whitelist. The implementation accepts only "embedding_hidden" or None, where None defaults internally to "embedding_hidden".

Any alternative ordering—such as "hidden_embedding"—immediately fails validation. This constraint is enforced before tensor operations begin, ensuring the MTP block output concatenation follows the exact architecture-specific layout required by Nemotron-H models.

MTP Depth and Layer Limitations

The Nemotron-H backend hardcodes a maximum prediction depth of exactly one step. Within mtp_forward, the code explicitly checks int(mtp_depth) > 1 and raises a ValueError when larger values are supplied.

Additionally, the configuration must specify precisely one MTP hidden layer. The validation function is_nemotron_h_mtp_config verifies that _num_mtp_layers(config) == 1, rejecting any attempt to initialize multi-layer MTP stacks regardless of the pattern specified.

Valid MTP Patterns for Nemotron-H

The MTP pattern defines the internal block structure using a domain-specific string notation parsed from the model configuration. According to the validation logic in is_nemotron_h_mtp_config, patterns must satisfy the constraint set(pattern).issubset({"*", "E"}).

  • * represents standard multi-head attention blocks
  • E represents Mixture-of-Experts (MoE) layers

Valid configurations include patterns like "*E*", "E", or "**". Any string containing characters outside this binary set invalidates the configuration and prevents model initialization.

Implementation Source Files

Three critical files in the youssofal/MTPLX repository define these constraints:

Practical Configuration Examples

The following demonstrates valid usage according to the MTPLX source constraints:


# Valid Nemotron-H MTP invocation

output = model.mtp_forward(
    hidden_states=backbone_output,
    next_token_ids=candidate_tokens,
    concat_order="embedding_hidden",  # Only accepted value

    mtp_depth=1,                      # Must be exactly 1

    return_hidden=False,
)

Attempting to use unsupported parameters triggers immediate runtime exceptions:


# Invalid: Raises ValueError for depth > 1 and unsupported concat order

model.mtp_forward(
    hidden_states,
    next_token_ids,
    concat_order="hidden_embedding",   # ❌ Not in {None, "embedding_hidden"}

    mtp_depth=2,                      # ❌ ValueError: depth must be 1

)

Summary

  • Concatenation order must be "embedding_hidden"; no other tensor arrangements are supported.
  • MTP depth is strictly limited to 1; any higher value raises ValueError in mtp_forward.
  • Pattern characters are restricted to * (attention) and E (MoE) exclusively.
  • Layer count must be exactly one; multi-layer configurations fail validation in is_nemotron_h_mtp_config.
  • Validation occurs in mtplx/backends/nemotron_h_mtp.py and mtplx/nemotron_h_mtp_patch.py.

Frequently Asked Questions

What error does MTPLX raise when mtp_depth exceeds 1 for Nemotron-H?

The mtp_forward function in mtplx/backends/nemotron_h_mtp.py raises a ValueError immediately upon detecting int(mtp_depth) > 1. The Nemotron-H backend is architected exclusively for single-step multi-token prediction, and no multi-depth execution path exists in the current implementation.

Can I customize the concatenation order in Nemotron-H MTP configurations?

No. The validation logic explicitly checks against the set {None, "embedding_hidden"}. Supplying alternatives like "hidden_embedding" causes the forward pass to fail before computation begins, as the underlying CUDA kernels and tensor layouts expect the specific embedding_hidden ordering.

What characters are permitted in the MTP pattern string?

Only two symbols are valid: * representing attention blocks and E representing MoE layers. The configuration validator in is_nemotron_h_mtp_config uses set(pattern).issubset({"*", "E"}) to verify compliance, rejecting any pattern containing digits, lowercase letters, or other symbols.

Where is the single-layer requirement enforced in the codebase?

The check occurs in two locations: is_nemotron_h_mtp_config within mtplx/backends/nemotron_h_mtp.py verifies _num_mtp_layers(config) == 1, while mtplx/nemotron_h_mtp_patch.py handles the runtime injection logic that physically constrains the stack to a single layer during model initialization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →