Understanding the Draft Verify Accept Reject Process in MTPLX's MTP Pipeline

MTPLX's MTP pipeline accelerates generation by drafting candidate tokens with a lightweight head, verifying them against the full model's probability distribution, and committing only accepted tokens while trimming rejects and falling back to primary sampling.

The MTPLX repository implements a Multitoken Predict (MTP) engine that decouples fast speculative drafting from rigorous verification. This workflow lives in the mtplx/ package and relies on a tight coordination between the draft language model head, the verification kernels, and the telemetry layer defined in the test suite.

Draft Generation with DraftLMHead

The process begins when the DraftLMHead class produces a block of speculative tokens ahead of the main generation cursor.

The DraftLMHead Module

Located in mtplx/draft_lm_head.py, this module instantiates a reduced‑capacity or quantized language model head optimized for speed. It emits a configurable number of candidate tokens—controlled by the draft_block_size parameter—without updating the primary KV‑cache. These candidates form the draft block that enters the verification pipeline.


# mtplx/draft_lm_head.py conceptual usage

draft_head = mtplx.draft_lm_head.DraftLMHead(state)
draft_tokens = draft_head.sample(block_size=state.args.draft_block_size)

The Verification Step

Once the draft block is materialized, the system evaluates each token individually using the full‑scale verification model.

Cross‑Row Verification Logic

The verification kernels reside in mtplx/verify_crossrow.py (and the quantized variant mtplx/verify_qmv.py). These modules expose a stateless function—typically verify_token—that computes the probability of a draft token under the main model’s distribution. As noted in the source, this step is deliberately stateless per token to allow parallel evaluation of the entire draft block.


# mtplx/verify_crossrow.py workflow

for i, token in enumerate(draft_tokens):
    prob = mtplx.verify_crossrow.verify_token(state, token)
    # Acceptance logic follows...

Acceptance Criteria

A draft token is accepted when its verification probability exceeds a uniform random draw. This stochastic rule guarantees that the final sequence remains statistically identical to samples drawn directly from the full model, preserving unbiasedness while re‑using the cheap draft head for throughput.

Accept vs. Reject Handling

After verification, the engine forks into two distinct paths: committing accepted tokens or repairing rejected ones.

Committing Accepted Drafts

When prob > random.random(), the token is committed to the primary KV‑cache, the draft head’s internal pointer advances, and the runtime increments the accepted_draft_tokens telemetry counter. This counter is visible in the OpenAI‑compatible server tests; specifically, tests/test_server_openai.py captures these metrics around lines 570–580 for observability.


# Acceptance path

if prob > random.random():
    state.commit_token(token)
    state.telemetry["accepted_draft_tokens"] += 1

Rejection and Repair Mechanics

If the token fails verification, the engine triggers a rejection sequence. The uncommitted tail of the draft cache is trimmed—referenced in the comment at line 549 of mtplx/draft_sampling.py regarding "doomed‑draft verify work"—and the rejected_draft_tokens metric is incremented. The system then falls back to the primary sampler for that position. Complex rejection scenarios invoke the rejection‑repair logic tested in tests/test_snapshot_free_rejection_repair.py, which prunes dangling speculative state before resuming generation.


# Rejection path

else:
    state.trim_draft_tail(i)  # Drop from index i onward

    state.telemetry["rejected_draft_tokens"] += 1
    # Fallback to primary generation

    token = state.primary_sampler.sample_one()
    state.commit_token(token)

Configuration and Telemetry

Operators can tune the draft-verify trade-off via runtime settings and observe health through structured telemetry.

Draft Control Settings

The /settings endpoint exposes draft_block_size and sampling parameters such as draft_top_p. Adjusting these values changes how aggressively the system speculates versus verifies, directly impacting the accept/reject ratio.

import requests
resp = requests.post(
    "http://localhost:8000/settings",
    json={"draft_block_size": 3, "draft_top_p": 0.9},
)
print(resp.json()["draft_control"])

Observability via Test Telemetry

The test suite in tests/test_server_openai.py validates the accept/reject bookkeeping, ensuring that accepted_draft_tokens and rejected_draft_tokens aggregate correctly across generation cycles. Unit tests in tests/test_verify_crossrow.py verify the statistical correctness of the acceptance kernel itself.

Summary

  • Draft Phase: mtplx/draft_lm_head.py generates speculative blocks using a fast, lightweight model head.
  • Verification Phase: mtplx/verify_crossrow.py evaluates each token’s probability under the full model; acceptance requires prob > random().
  • Accept Path: Tokens are committed to the KV‑cache and recorded in accepted_draft_tokens telemetry.
  • Reject Path: Failed tokens trigger cache trimming (see mtplx/draft_sampling.py line 549) and fallback sampling, logged under rejected_draft_tokens.
  • Repair Logic: tests/test_snapshot_free_rejection_repair.py ensures clean state recovery after batch rejections.
  • Control: Runtime parameters adjustable via the settings endpoint govern block size and sampling thresholds.

Frequently Asked Questions

What triggers the rejection repair mechanism in MTPLX?

Rejection repair activates when a draft token fails verification and the system must prune the uncommitted speculative tail. The repair logic, validated in tests/test_snapshot_free_rejection_repair.py, trims dangling cache entries to prevent partial sequences from corrupting the primary generation state.

How does the verification step ensure unbiased sampling?

The verification kernel in mtplx/verify_crossrow.py computes the exact probability of each draft token under the full model. Acceptance occurs only when this probability exceeds a uniform random threshold, mathematically guaranteeing that the marginal distribution of accepted tokens matches the target model’s distribution.

Where is the draft token telemetry exposed in MTPLX?

Telemetry counters for accepted_draft_tokens and rejected_draft_tokens are captured in the server test suite, specifically within tests/test_server_openai.py around lines 570–580, and can be surfaced through the OpenAI‑compatible API’s logging hooks or the /settings telemetry response.

Can the draft block size be adjusted dynamically?

Yes. The draft_block_size parameter is mutable via the /settings endpoint; changes take effect on the next generation cycle, allowing clients to trade off between latency (larger blocks) and acceptance rate (smaller blocks) without restarting the inference engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →