# Understanding the Draft Verify Accept Reject Process in MTPLX's MTP Pipeline

> Uncover MTPLX's MTP pipeline: learn how draft, verify, accept, and reject processes efficiently generate candidate tokens and optimize text generation with fallback sampling.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: internals
- Published: 2026-09-04

---

**MTPLX's MTP pipeline accelerates generation by drafting candidate tokens with a lightweight head, verifying them against the full model's probability distribution, and committing only accepted tokens while trimming rejects and falling back to primary sampling.**

The **MTPLX** repository implements a *Multitoken Predict* (MTP) engine that decouples fast speculative drafting from rigorous verification. This workflow lives in the `mtplx/` package and relies on a tight coordination between the draft language model head, the verification kernels, and the telemetry layer defined in the test suite.

## Draft Generation with DraftLMHead

The process begins when the `DraftLMHead` class produces a block of speculative tokens ahead of the main generation cursor.

### The DraftLMHead Module

Located in [`mtplx/draft_lm_head.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/draft_lm_head.py), this module instantiates a reduced‑capacity or quantized language model head optimized for speed. It emits a configurable number of candidate tokens—controlled by the `draft_block_size` parameter—without updating the primary KV‑cache. These candidates form the *draft block* that enters the verification pipeline.

```python

# mtplx/draft_lm_head.py conceptual usage

draft_head = mtplx.draft_lm_head.DraftLMHead(state)
draft_tokens = draft_head.sample(block_size=state.args.draft_block_size)

```

## The Verification Step

Once the draft block is materialized, the system evaluates each token individually using the full‑scale verification model.

### Cross‑Row Verification Logic

The verification kernels reside in [`mtplx/verify_crossrow.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/verify_crossrow.py) (and the quantized variant [`mtplx/verify_qmv.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/verify_qmv.py)). These modules expose a stateless function—typically `verify_token`—that computes the probability of a draft token under the main model’s distribution. As noted in the source, this step is deliberately *stateless per token* to allow parallel evaluation of the entire draft block.

```python

# mtplx/verify_crossrow.py workflow

for i, token in enumerate(draft_tokens):
    prob = mtplx.verify_crossrow.verify_token(state, token)
    # Acceptance logic follows...

```

### Acceptance Criteria

A draft token is accepted when its verification probability exceeds a uniform random draw. This stochastic rule guarantees that the final sequence remains statistically identical to samples drawn directly from the full model, preserving unbiasedness while re‑using the cheap draft head for throughput.

## Accept vs. Reject Handling

After verification, the engine forks into two distinct paths: committing accepted tokens or repairing rejected ones.

### Committing Accepted Drafts

When `prob > random.random()`, the token is **committed** to the primary KV‑cache, the draft head’s internal pointer advances, and the runtime increments the `accepted_draft_tokens` telemetry counter. This counter is visible in the OpenAI‑compatible server tests; specifically, [`tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py) captures these metrics around lines 570–580 for observability.

```python

# Acceptance path

if prob > random.random():
    state.commit_token(token)
    state.telemetry["accepted_draft_tokens"] += 1

```

### Rejection and Repair Mechanics

If the token fails verification, the engine triggers a **rejection** sequence. The uncommitted tail of the draft cache is trimmed—referenced in the comment at line 549 of [`mtplx/draft_sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/draft_sampling.py) regarding "doomed‑draft verify work"—and the `rejected_draft_tokens` metric is incremented. The system then falls back to the primary sampler for that position. Complex rejection scenarios invoke the *rejection‑repair* logic tested in [`tests/test_snapshot_free_rejection_repair.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_snapshot_free_rejection_repair.py), which prunes dangling speculative state before resuming generation.

```python

# Rejection path

else:
    state.trim_draft_tail(i)  # Drop from index i onward

    state.telemetry["rejected_draft_tokens"] += 1
    # Fallback to primary generation

    token = state.primary_sampler.sample_one()
    state.commit_token(token)

```

## Configuration and Telemetry

Operators can tune the draft-verify trade-off via runtime settings and observe health through structured telemetry.

### Draft Control Settings

The `/settings` endpoint exposes `draft_block_size` and sampling parameters such as `draft_top_p`. Adjusting these values changes how aggressively the system speculates versus verifies, directly impacting the accept/reject ratio.

```python
import requests
resp = requests.post(
    "http://localhost:8000/settings",
    json={"draft_block_size": 3, "draft_top_p": 0.9},
)
print(resp.json()["draft_control"])

```

### Observability via Test Telemetry

The test suite in [`tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py) validates the accept/reject bookkeeping, ensuring that `accepted_draft_tokens` and `rejected_draft_tokens` aggregate correctly across generation cycles. Unit tests in [`tests/test_verify_crossrow.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_verify_crossrow.py) verify the statistical correctness of the acceptance kernel itself.

## Summary

- **Draft Phase:** [`mtplx/draft_lm_head.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/draft_lm_head.py) generates speculative blocks using a fast, lightweight model head.
- **Verification Phase:** [`mtplx/verify_crossrow.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/verify_crossrow.py) evaluates each token’s probability under the full model; acceptance requires `prob > random()`.
- **Accept Path:** Tokens are committed to the KV‑cache and recorded in `accepted_draft_tokens` telemetry.
- **Reject Path:** Failed tokens trigger cache trimming (see [`mtplx/draft_sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/draft_sampling.py) line 549) and fallback sampling, logged under `rejected_draft_tokens`.
- **Repair Logic:** [`tests/test_snapshot_free_rejection_repair.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_snapshot_free_rejection_repair.py) ensures clean state recovery after batch rejections.
- **Control:** Runtime parameters adjustable via the settings endpoint govern block size and sampling thresholds.

## Frequently Asked Questions

### What triggers the rejection repair mechanism in MTPLX?

Rejection repair activates when a draft token fails verification and the system must prune the uncommitted speculative tail. The repair logic, validated in [`tests/test_snapshot_free_rejection_repair.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_snapshot_free_rejection_repair.py), trims dangling cache entries to prevent partial sequences from corrupting the primary generation state.

### How does the verification step ensure unbiased sampling?

The verification kernel in [`mtplx/verify_crossrow.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/verify_crossrow.py) computes the exact probability of each draft token under the full model. Acceptance occurs only when this probability exceeds a uniform random threshold, mathematically guaranteeing that the marginal distribution of accepted tokens matches the target model’s distribution.

### Where is the draft token telemetry exposed in MTPLX?

Telemetry counters for `accepted_draft_tokens` and `rejected_draft_tokens` are captured in the server test suite, specifically within [`tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py) around lines 570–580, and can be surfaced through the OpenAI‑compatible API’s logging hooks or the `/settings` telemetry response.

### Can the draft block size be adjusted dynamically?

Yes. The `draft_block_size` parameter is mutable via the `/settings` endpoint; changes take effect on the next generation cycle, allowing clients to trade off between latency (larger blocks) and acceptance rate (smaller blocks) without restarting the inference engine.