# What Is the Role of MTP Draft Heads in MTPLX? A Deep Dive into the Draft-Then-Verify Pipeline

> Discover the role of MTP draft heads in MTPLX. Learn how these lightweight heads boost throughput by proposing token candidates for a faster draft-then-verify pipeline.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: deep-dive
- Published: 2026-09-08

---

**MTP draft heads in MTPLX are lightweight language model heads that propose token candidates in a fast "draft" pass, enabling higher throughput by letting an expensive "verify" head only evaluate pre-selected tokens rather than computing over the full vocabulary.**

The MTPLX repository implements a **Multi-Token-Prediction (MTP)** system that separates token generation into two stages: a fast draft stage that "guesses" likely next tokens, and a slower verification stage that confirms correctness. This article explains how MTP draft heads function within this architecture, based on the source code in `youssofal/MTPLX`.

## How MTP Draft Heads Fit into the Generation Pipeline

MTPLX uses a **draft-then-verify** pattern to accelerate autoregressive generation. The draft head operates as a lightweight proxy for the full model, producing token logits that are sampled to create candidate sequences. The primary verify head—typically a larger, more accurate model—then checks these candidates and accepts or rejects them.

This design trades **slight increases in memory and coordination overhead** for **significant reductions in per-token compute cost**, since the verify head runs far less frequently than it would in standard generation.

### The Core Files

Three files define the draft head behavior:

| File | Purpose |
|------|---------|
| [`mtplx/frspec_draft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/frspec_draft.py) | Implements draft head installation, pruning, and full-vocabulary wrapping |
| [`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py) | Tracks `draft_head_identity` for session-scoped resource management |
| [`mtplx/server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server_openai.py) | Exposes runtime controls for draft sampler parameters |

## Installing and Configuring Draft Heads

Draft heads are attached to models after loading via **`install_frspec_draft_head`** in [`mtplx/frspec_draft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/frspec_draft.py). This function wraps a potentially pruned head with a full-vocabulary interface so downstream components need not handle shape mismatches.

### Activation via Environment Variables

The draft head is controlled through three environment variables:

- **`MTPLX_FRSPEC_DRAFT`** — enables the pruned draft head when set to `"1"`
- **`MTPLX_FRSPEC_VOCAB`** — specifies which vocabulary subset to use (e.g., `"builtin:qwen38-code-64k"`)
- **`MTPLX_FRSPEC_N`** — sets the number of vocabulary rows to retain (default typically 65,536)

```python
import os
from mtplx.frspec_draft import install_frspec_draft_head

# Enable frequency-ranked pruned draft head

os.environ["MTPLX_FRSPEC_DRAFT"] = "1"
os.environ["MTPLX_FRSPEC_VOCAB"] = "builtin:qwen38-code-64k"

# After loading your model (variable `text`), install the draft head

report = install_frspec_draft_head(text)
print(report)

# Output: {'installed': True, 'n': 65536, 'vocab_rows': 65536, 'bytes_ratio': 0.25, ...}

```

The `report` dictionary returned by `install_frspec_draft_head` includes metadata about the pruning ratio and memory savings, allowing callers to verify the configuration took effect.

## FR-Spec: Pruning Draft Heads for Speed

MTPLX implements **Frequency-Ranked Speculative decoding (FR-Spec)** to reduce draft head compute. Rather than computing logits over the full vocabulary, the draft head operates on a pruned subset containing the most frequent tokens.

### Probability-Ratio Correction

Pruning to a vocabulary subset would normally bias sampling toward high-frequency tokens. MTPLX preserves correctness through **probability-ratio correction**—adjusting the draft logits so that the acceptance probability in the verify stage accounts for the missing mass.

The `_full_vocab_head` wrapper (lines 42–74 in [`mtplx/frspec_draft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/frspec_draft.py)) handles the interface boundary: it expands pruned outputs back to full vocabulary size by inserting `-inf` logits for excluded tokens. This lets the verify stage treat draft outputs identically to standard LM head outputs.

### Legacy Mode for Compatibility

When **`MTPLX_FRSPEC_LEGACY`** is enabled, the draft head is swapped globally so older code paths expecting direct `draft_lm_head` access continue functioning. This is implemented in the same installation function and minimizes breaking changes across MTPLX versions.

## Session-State Tracking and Resource Management

Draft heads are not stateless utilities—they carry identity and configuration that must persist across a generation session. The `SessionBank` class in [`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py) tracks this via **`draft_head_identity`**, a field passed when creating new sessions.

```python
from mtplx.session_bank import SessionBank

bank = SessionBank(...)
session = bank.new_session(
    draft_head_identity="my-draft-head",
    # ... other session parameters

)

```

This identity ensures that:
- Draft-specific caches are correctly scoped
- Multiple concurrent sessions using different draft configurations do not interfere
- Resource cleanup targets the correct head when sessions terminate

## Runtime Sampler Controls

Draft heads respect the same sampling parameters as primary generation, exposed through the MTPLX server API. In [`mtplx/server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server_openai.py), the `/mtplx/settings` endpoint accepts draft-specific overrides:

```python
import requests

response = requests.post(
    "http://localhost:8000/mtplx/settings",
    json={
        "draft_temperature": 0.62,
        "draft_top_p": 0.95,
        "draft_top_k": 12,
    }
)
print(response.json())  # Confirms applied draft sampler configuration

```

These controls affect only the draft sampling stage—the verify head operates under its own independent configuration, allowing fine-grained trade-offs between draft diversity and verification load.

## Performance Characteristics of MTP Draft Heads

| Aspect | Behavior | Source Location |
|--------|----------|-----------------|
| Memory footprint | Reduced by `bytes_ratio` factor via pruning | [`mtplx/frspec_draft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/frspec_draft.py) |
| Compute per draft step | O(n) where n = `MTPLX_FRSPEC_N` (≤ 65K) vs. full vocab | Environment configuration |
| Accept rate | Depends on draft/verify model agreement; corrected by probability ratio | [`mtplx/frspec_draft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/frspec_draft.py) |
| Throughput gain | Higher when verify head is large and draft head is aggressively pruned | Architectural design |

The FR-Spec approach specifically targets **memory-bandwidth-bound** inference scenarios, where reducing the output dimension directly translates to faster token generation.

## Summary

- **MTP draft heads generate candidate tokens quickly** using lightweight, optionally pruned language model heads, forming the "draft" stage of MTPLX's draft-then-verify pipeline.

- **Installation happens via `install_frspec_draft_head`** in [`mtplx/frspec_draft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/frspec_draft.py), which wraps pruned heads with a full-vocabulary interface for compatibility.

- **FR-Spec pruning** reduces compute by restricting the draft head to frequent tokens, with correctness preserved through probability-ratio correction and the `_full_vocab_head` wrapper.

- **Session tracking** via `draft_head_identity` in [`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py) ensures proper resource scoping across concurrent generation sessions.

- **Runtime controls** for temperature, top-p, and top-k sampling are exposed through the server's settings endpoint in [`mtplx/server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server_openai.py).

## Frequently Asked Questions

### What is the difference between a draft head and a verify head in MTPLX?

The draft head is a fast, lightweight model component that proposes token candidates using reduced compute—often via FR-Spec pruning to a subset of the vocabulary. The verify head is the primary, typically larger language model that evaluates draft candidates for correctness. Only accepted tokens are emitted, ensuring output quality matches the verify head's capability while reducing the number of expensive forward passes.

### How does FR-Spec pruning affect output quality?

FR-Spec pruning preserves output quality through **probability-ratio correction**. When the draft head operates on a pruned vocabulary, the missing probability mass is accounted for during verification, ensuring the combined system remains mathematically equivalent to sampling from the full vocabulary. The trade-off is in acceptance rate—more aggressive pruning may reduce how often the verify head accepts draft tokens, increasing verification overhead.

### Can I use multiple draft heads with different configurations in the same MTPLX deployment?

Yes, through the **session bank mechanism**. Each session carries a `draft_head_identity` that associates it with specific draft resources. The `SessionBank` class manages these identities, allowing concurrent sessions to use different draft heads, pruning levels, or sampler configurations without interference.

### What happens if I enable `MTPLX_FRSPEC_LEGACY`?

Enabling `MTPLX_FRSPEC_LEGACY` causes `install_frspec_draft_head` to perform a global swap of the draft LM head reference, ensuring backward compatibility with code paths that expect direct `draft_lm_head` access rather than the newer wrapped interface. This is implemented in lines 42–60 of [`mtplx/frspec_draft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/frspec_draft.py) and primarily affects older integrations or custom extensions.