Does MTPLX Use Greedy-Argmax or a Second Draft Model for Speculative Decoding?

MTPLX implements speculative decoding by running a second draft model to generate candidate tokens that are verified by the target model, rather than using a greedy-argmax approach.

Speculative decoding accelerates large language model inference by predicting multiple tokens in parallel. In the youssofal/MTPLX repository, this optimization relies on a distinct draft model architecture configured through the SamplerConfig class, utilizing an exact-speculative loop for verification rather than simple greedy selection.

Draft Model Architecture in MTPLX

Exact-Speculative Loop Design

The codebase implements an exact-speculative loop where a secondary draft model generates candidate sequences before the primary target model verifies them. This design appears in tests/test_tail_gemma4_stream_holdback.py, which references an exact-speculative assistant pattern that validates draft outputs against the target distribution. This approach differs from greedy-argmax methods that select only the highest probability token at each step without speculative lookahead.

Speculative Depth and Draft Configuration

MTPLX controls speculative generation through the speculative_depth parameter observed in tests/test_deepseek_v4_spec.py. This parameter determines how many tokens the draft model produces before the target model performs verification, allowing the system to balance parallelism against draft model accuracy. The configuration explicitly separates the draft and target responsibilities using distinct core identifiers.

Implementation in mtplx/sampling.py

SamplerConfig Parameters

The SamplerConfig class defines the draft and target model relationship through two key fields:

  • draft_core: Identifies the secondary model responsible for speculative token generation
  • verify_core: Specifies the target model that validates draft outputs and executes rollbacks on mismatch

The speculative_output_marginal Function

The core verification logic resides in speculative_output_marginal, which accepts target probabilities (target_p), draft probabilities (draft_q), and the configuration object. This function compares draft-generated candidates against the target distribution, accepting tokens sequentially until detecting a mismatch.

from mtplx.sampling import SamplerConfig, speculative_output_marginal

config = SamplerConfig(
    draft_core="stock",          # draft model identifier

    speculative_depth=2,         # number of draft tokens to generate

    verify_core="target",        # target model for verification

)

# Run decoding – draft model produces candidates, target model verifies

output = speculative_output_marginal(target_p, draft_q, config)

Key Source Files

  • tests/test_tail_gemma4_stream_holdback.py: Demonstrates the exact-speculative assistant pattern using a dedicated draft model stage before verification.
  • tests/test_deepseek_v4_spec.py: Validates the speculative_depth parameter controlling the volume of draft tokens generated before target model verification.
  • mtplx/sampling.py: Contains the speculative_output_marginal implementation and the SamplerConfig class definition that orchestrates the draft-target interaction.

Summary

  • MTPLX uses a second draft model configured via draft_core rather than greedy-argmax for speculative decoding.
  • The speculative_depth parameter in SamplerConfig controls how many tokens are generated before target model verification.
  • The speculative_output_marginal function in mtplx/sampling.py executes the verification logic and handles rollback to the last accepted token.
  • Test suites in test_tail_gemma4_stream_holdback.py and test_deepseek_v4_spec.py confirm the exact-speculative loop implementation with distinct draft and target stages.

Frequently Asked Questions

Does MTPLX use greedy-argmax for speculative decoding?

No. According to the youssofal/MTPLX source code, the framework employs a second draft model specified by the draft_core parameter to generate candidate tokens. The target model verifies these candidates against its own distribution, which differs fundamentally from greedy-argmax selection that would only consider the highest probability single token at each step.

What does the speculative_depth parameter control?

The speculative_depth parameter, evidenced in tests/test_deepseek_v4_spec.py, determines the number of tokens the draft model generates before the target model executes verification. This allows the system to balance between parallelization (higher depth) and the draft model's distribution alignment with the target.

How does speculative_output_marginal handle token verification?

This function, implemented in mtplx/sampling.py, accepts the target model probabilities (target_p), draft model probabilities (draft_q), and a SamplerConfig instance. It validates draft tokens sequentially against the target distribution, accepting matches until encountering a discrepancy, then rolls back to the last verified position.

Where is the draft model specified in MTPLX?

The draft model is configured through the draft_core field in the SamplerConfig class, as demonstrated in the sampling module and validated by the test suite. This separation between draft_core (generator) and verify_core (validator) enables the dual-model speculative architecture.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →