Does MTPLX Use Greedy-Argmax or a Second Draft Model for Speculative Decoding?
MTPLX implements speculative decoding by running a second draft model to generate candidate tokens that are verified by the target model, rather than using a greedy-argmax approach.
Speculative decoding accelerates large language model inference by predicting multiple tokens in parallel. In the youssofal/MTPLX repository, this optimization relies on a distinct draft model architecture configured through the SamplerConfig class, utilizing an exact-speculative loop for verification rather than simple greedy selection.
Draft Model Architecture in MTPLX
Exact-Speculative Loop Design
The codebase implements an exact-speculative loop where a secondary draft model generates candidate sequences before the primary target model verifies them. This design appears in tests/test_tail_gemma4_stream_holdback.py, which references an exact-speculative assistant pattern that validates draft outputs against the target distribution. This approach differs from greedy-argmax methods that select only the highest probability token at each step without speculative lookahead.
Speculative Depth and Draft Configuration
MTPLX controls speculative generation through the speculative_depth parameter observed in tests/test_deepseek_v4_spec.py. This parameter determines how many tokens the draft model produces before the target model performs verification, allowing the system to balance parallelism against draft model accuracy. The configuration explicitly separates the draft and target responsibilities using distinct core identifiers.
Implementation in mtplx/sampling.py
SamplerConfig Parameters
The SamplerConfig class defines the draft and target model relationship through two key fields:
draft_core: Identifies the secondary model responsible for speculative token generationverify_core: Specifies the target model that validates draft outputs and executes rollbacks on mismatch
The speculative_output_marginal Function
The core verification logic resides in speculative_output_marginal, which accepts target probabilities (target_p), draft probabilities (draft_q), and the configuration object. This function compares draft-generated candidates against the target distribution, accepting tokens sequentially until detecting a mismatch.
from mtplx.sampling import SamplerConfig, speculative_output_marginal
config = SamplerConfig(
draft_core="stock", # draft model identifier
speculative_depth=2, # number of draft tokens to generate
verify_core="target", # target model for verification
)
# Run decoding – draft model produces candidates, target model verifies
output = speculative_output_marginal(target_p, draft_q, config)
Key Source Files
tests/test_tail_gemma4_stream_holdback.py: Demonstrates the exact-speculative assistant pattern using a dedicated draft model stage before verification.tests/test_deepseek_v4_spec.py: Validates thespeculative_depthparameter controlling the volume of draft tokens generated before target model verification.mtplx/sampling.py: Contains thespeculative_output_marginalimplementation and theSamplerConfigclass definition that orchestrates the draft-target interaction.
Summary
- MTPLX uses a second draft model configured via
draft_corerather than greedy-argmax for speculative decoding. - The
speculative_depthparameter inSamplerConfigcontrols how many tokens are generated before target model verification. - The
speculative_output_marginalfunction inmtplx/sampling.pyexecutes the verification logic and handles rollback to the last accepted token. - Test suites in
test_tail_gemma4_stream_holdback.pyandtest_deepseek_v4_spec.pyconfirm the exact-speculative loop implementation with distinct draft and target stages.
Frequently Asked Questions
Does MTPLX use greedy-argmax for speculative decoding?
No. According to the youssofal/MTPLX source code, the framework employs a second draft model specified by the draft_core parameter to generate candidate tokens. The target model verifies these candidates against its own distribution, which differs fundamentally from greedy-argmax selection that would only consider the highest probability single token at each step.
What does the speculative_depth parameter control?
The speculative_depth parameter, evidenced in tests/test_deepseek_v4_spec.py, determines the number of tokens the draft model generates before the target model executes verification. This allows the system to balance between parallelization (higher depth) and the draft model's distribution alignment with the target.
How does speculative_output_marginal handle token verification?
This function, implemented in mtplx/sampling.py, accepts the target model probabilities (target_p), draft model probabilities (draft_q), and a SamplerConfig instance. It validates draft tokens sequentially against the target distribution, accepting matches until encountering a discrepancy, then rolls back to the last verified position.
Where is the draft model specified in MTPLX?
The draft model is configured through the draft_core field in the SamplerConfig class, as demonstrated in the sampling module and validated by the test suite. This separation between draft_core (generator) and verify_core (validator) enables the dual-model speculative architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →