HY V3 MTP Features in MTPLX: Architecture and Implementation Guide
HY V3 MTP is an experimental-native-contract-gated backend that adds exactly one appended MTP (NextN) layer with a 192-expert MoE to HY V3 models, enabling verified execution through exact rejection sampling.
The MTPLX framework by youssofal provides specialized backends for hybrid model architectures. HY V3 MTP extends the base HY V3 model with a Medusa-style token prediction layer, implementing specific constraints for runtime compatibility and verification. This backend is registered as arch_id="hy-v3-mtp" and integrates deeply with MTPLX's contract-gated runtime system.
Architecture Registration and Support Level
In mtplx/backends/registry.py at line 580, HY V3 MTP is cataloged with the architecture identifier hy-v3-mtp and the user-friendly name "HY V3 MTP". The registration specifies a support level of experimental-native-contract-gated, indicating the backend operates under a verified runtime contract while remaining in experimental status.
The registry entry defines the following critical capabilities:
- Runtime compatibility:
native-contract-gated— requires the native contract-gated runtime - Verified execution:
can_run_verified=True— supports execution under MTPLX's verification framework - MTP depth limit:
mtp_depth_max = 1— supports exactly one MTP layer, no more
Core Backend Implementation
The HyV3MTPBackend class in mtplx/backends/hy_v3_mtp.py (line 25) serves as the primary façade for this architecture. This implementation handles model loading, health reporting, and integration with the generation pipeline.
The backend wires draft generation and verification logic to the MTPLX runtime, specifically managing how the MTP layer consumes intermediate states from the base HY V3 trunk model.
MTP Architecture Specifications
HY V3 MTP implements a single appended NextN layer with highly specific architectural constraints:
- MoE Configuration: 192-expert Mixture-of-Experts MLP
- Routing mechanism: Sigmoid top-8 routing with expert bias
- Projection layer:
eh_projprojects the concatenation of normalized embeddings and hidden states - Weight sharing: Shares the trunk's embeddings and LM head with the base HY V3 model
- State consumption: The draft layer specifically consumes the trunk's post-final-norm hidden state, achieving approximately 0.773 agreement versus 0.387 with pre-norm states
The verification mechanism employs exact rejection sampling at temperature 0, ensuring deterministic output validation.
Runtime Compatibility and Verification Support
HY V3 MTP requires the native-contract-gated runtime for execution. This runtime requirement ensures that the model adheres to strict contracts regarding memory access and computational boundaries.
The backend supports verified execution where can_run_verified=True enables the model to run under MTPLX's verification framework. This capability is essential for production deployments requiring correctness guarantees, utilizing exact rejection sampling for draft token validation.
Health Reporting and Diagnostic Features
The health() method at line 48 in mtplx/backends/hy_v3_mtp.py exposes comprehensive diagnostic information through a dictionary containing:
arch_id: The architecture identifier (hy-v3-mtp)runtime_path: Execution path includingmtplx.runtime + mtplx.hy_v3_mtp_patch + mtplx.generationsupport_level: Current experimental statuscontract_required: Boolean indicating runtime contract requirementssupported_model_types: List containing["hy_v3"]mtp_depth_max: Maximum supported MTP layers (1)notes: Concise description of the MTP layer implementationreferences: Pointers to reference implementations invllm/model_executor/models/hy_v3_mtp.pyandmlx_lm/models/hy_v3.py
Loading and Using HY V3 MTP Models
To instantiate a HY V3 MTP model within MTPLX, use the standard loader interface which automatically detects the architecture:
from mtplx import MTPLX
mtplx = MTPLX()
model = mtplx.load("path/to/hy_v3_mtp_checkpoint")
To inspect backend capabilities and verify feature support programmatically:
health = mtplx.backends["hy_v3_mtp"].health()
print(health)
This returns a structured dictionary including architecture metadata, runtime requirements, and implementation references as defined in the backend's health method.
Summary
- HY V3 MTP adds exactly one appended MTP (NextN) layer to base HY V3 models in MTPLX
- Implementation resides in
HyV3MTPBackendwithinmtplx/backends/hy_v3_mtp.py - Runtime requirement is
native-contract-gatedwith verified execution support enabled - Architecture specifics include 192-expert MoE, sigmoid top-8 routing, and post-final-norm state consumption
- Registration occurs in
mtplx/backends/registry.pywitharch_id="hy-v3-mtp"and experimental support level - Diagnostics are available through the
health()method, exposing contract requirements and reference implementations
Frequently Asked Questions
What is the maximum MTP depth supported by HY V3 MTP?
HY V3 MTP supports exactly one MTP layer, enforced by mtp_depth_max = 1 in the registry configuration. The backend implements a single appended NextN layer and does not support speculative decoding with multiple draft layers.
What runtime compatibility does HY V3 MTP require?
The backend requires the native-contract-gated runtime, as specified by runtime_compatibility in the registry entry. This runtime ensures the model executes within verified memory and computational boundaries, enabling the can_run_verified=True capability for exact rejection sampling verification.
How does HY V3 MTP handle verification?
HY V3 MTP implements exact rejection sampling at temperature 0 for verification. The backend consumes the trunk's post-final-norm hidden state (achieving ~0.773 agreement) to generate draft tokens, which are then validated against the base model's outputs through deterministic rejection sampling.
Where are the reference implementations for HY V3 MTP?
The health() method in mtplx/backends/hy_v3_mtp.py references two canonical implementations: the vLLM implementation at vllm/model_executor/models/hy_v3_mtp.py and the mlx-lm implementation at mlx_lm/models/hy_v3.py. These serve as the authoritative references for the MTP layer behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →