# HY V3 MTP Features in MTPLX: Architecture and Implementation Guide

> Discover HY V3 MTP features in MTPLX. Learn how this native-contract-gated backend with a 192-expert MoE enables verified execution through exact rejection sampling. Read the implementation guide.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: architecture
- Published: 2026-09-04

---

**HY V3 MTP is an experimental-native-contract-gated backend that adds exactly one appended MTP (NextN) layer with a 192-expert MoE to HY V3 models, enabling verified execution through exact rejection sampling.**

The MTPLX framework by youssofal provides specialized backends for hybrid model architectures. HY V3 MTP extends the base HY V3 model with a Medusa-style token prediction layer, implementing specific constraints for runtime compatibility and verification. This backend is registered as `arch_id="hy-v3-mtp"` and integrates deeply with MTPLX's contract-gated runtime system.

## Architecture Registration and Support Level

In [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py) at line 580, HY V3 MTP is cataloged with the **architecture identifier** `hy-v3-mtp` and the user-friendly name "HY V3 MTP". The registration specifies a **support level** of `experimental-native-contract-gated`, indicating the backend operates under a verified runtime contract while remaining in experimental status.

The registry entry defines the following critical capabilities:

- **Runtime compatibility**: `native-contract-gated` — requires the native contract-gated runtime
- **Verified execution**: `can_run_verified=True` — supports execution under MTPLX's verification framework
- **MTP depth limit**: `mtp_depth_max = 1` — supports exactly one MTP layer, no more

## Core Backend Implementation

The `HyV3MTPBackend` class in [`mtplx/backends/hy_v3_mtp.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/hy_v3_mtp.py) (line 25) serves as the primary façade for this architecture. This implementation handles model loading, health reporting, and integration with the generation pipeline.

The backend wires draft generation and verification logic to the MTPLX runtime, specifically managing how the MTP layer consumes intermediate states from the base HY V3 trunk model.

## MTP Architecture Specifications

HY V3 MTP implements a **single appended NextN layer** with highly specific architectural constraints:

- **MoE Configuration**: 192-expert Mixture-of-Experts MLP
- **Routing mechanism**: Sigmoid top-8 routing with expert bias
- **Projection layer**: `eh_proj` projects the concatenation of normalized embeddings and hidden states
- **Weight sharing**: Shares the trunk's embeddings and LM head with the base HY V3 model
- **State consumption**: The draft layer specifically consumes the trunk's post-final-norm hidden state, achieving approximately 0.773 agreement versus 0.387 with pre-norm states

The **verification mechanism** employs exact rejection sampling at temperature 0, ensuring deterministic output validation.

## Runtime Compatibility and Verification Support

HY V3 MTP requires the **native-contract-gated runtime** for execution. This runtime requirement ensures that the model adheres to strict contracts regarding memory access and computational boundaries.

The backend supports **verified execution** where `can_run_verified=True` enables the model to run under MTPLX's verification framework. This capability is essential for production deployments requiring correctness guarantees, utilizing exact rejection sampling for draft token validation.

## Health Reporting and Diagnostic Features

The `health()` method at line 48 in [`mtplx/backends/hy_v3_mtp.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/hy_v3_mtp.py) exposes comprehensive diagnostic information through a dictionary containing:

- `arch_id`: The architecture identifier (`hy-v3-mtp`)
- `runtime_path`: Execution path including `mtplx.runtime + mtplx.hy_v3_mtp_patch + mtplx.generation`
- `support_level`: Current experimental status
- `contract_required`: Boolean indicating runtime contract requirements
- `supported_model_types`: List containing `["hy_v3"]`
- `mtp_depth_max`: Maximum supported MTP layers (1)
- `notes`: Concise description of the MTP layer implementation
- `references`: Pointers to reference implementations in [`vllm/model_executor/models/hy_v3_mtp.py`](https://github.com/youssofal/MTPLX/blob/main/vllm/model_executor/models/hy_v3_mtp.py) and [`mlx_lm/models/hy_v3.py`](https://github.com/youssofal/MTPLX/blob/main/mlx_lm/models/hy_v3.py)

## Loading and Using HY V3 MTP Models

To instantiate a HY V3 MTP model within MTPLX, use the standard loader interface which automatically detects the architecture:

```python
from mtplx import MTPLX

mtplx = MTPLX()
model = mtplx.load("path/to/hy_v3_mtp_checkpoint")

```

To inspect backend capabilities and verify feature support programmatically:

```python
health = mtplx.backends["hy_v3_mtp"].health()
print(health)

```

This returns a structured dictionary including architecture metadata, runtime requirements, and implementation references as defined in the backend's health method.

## Summary

- **HY V3 MTP** adds exactly one appended MTP (NextN) layer to base HY V3 models in MTPLX
- **Implementation** resides in `HyV3MTPBackend` within [`mtplx/backends/hy_v3_mtp.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/hy_v3_mtp.py)
- **Runtime requirement** is `native-contract-gated` with verified execution support enabled
- **Architecture specifics** include 192-expert MoE, sigmoid top-8 routing, and post-final-norm state consumption
- **Registration** occurs in [`mtplx/backends/registry.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/registry.py) with `arch_id="hy-v3-mtp"` and experimental support level
- **Diagnostics** are available through the `health()` method, exposing contract requirements and reference implementations

## Frequently Asked Questions

### What is the maximum MTP depth supported by HY V3 MTP?

HY V3 MTP supports **exactly one** MTP layer, enforced by `mtp_depth_max = 1` in the registry configuration. The backend implements a single appended NextN layer and does not support speculative decoding with multiple draft layers.

### What runtime compatibility does HY V3 MTP require?

The backend requires the **native-contract-gated runtime**, as specified by `runtime_compatibility` in the registry entry. This runtime ensures the model executes within verified memory and computational boundaries, enabling the `can_run_verified=True` capability for exact rejection sampling verification.

### How does HY V3 MTP handle verification?

HY V3 MTP implements **exact rejection sampling** at temperature 0 for verification. The backend consumes the trunk's post-final-norm hidden state (achieving ~0.773 agreement) to generate draft tokens, which are then validated against the base model's outputs through deterministic rejection sampling.

### Where are the reference implementations for HY V3 MTP?

The `health()` method in [`mtplx/backends/hy_v3_mtp.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/hy_v3_mtp.py) references two canonical implementations: the vLLM implementation at [`vllm/model_executor/models/hy_v3_mtp.py`](https://github.com/youssofal/MTPLX/blob/main/vllm/model_executor/models/hy_v3_mtp.py) and the mlx-lm implementation at [`mlx_lm/models/hy_v3.py`](https://github.com/youssofal/MTPLX/blob/main/mlx_lm/models/hy_v3.py). These serve as the authoritative references for the MTP layer behavior.