# Does MTPLX Use Greedy-Argmax or a Second Draft Model for Speculative Decoding?

> Discover if MTPLX uses greedy-argmax or a second draft model for speculative decoding. Learn how MTPLX employs a second draft model for efficient token generation and verification.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: deep-dive
- Published: 2026-09-04

---

**MTPLX implements speculative decoding by running a second draft model to generate candidate tokens that are verified by the target model, rather than using a greedy-argmax approach.**

Speculative decoding accelerates large language model inference by predicting multiple tokens in parallel. In the `youssofal/MTPLX` repository, this optimization relies on a distinct draft model architecture configured through the `SamplerConfig` class, utilizing an exact-speculative loop for verification rather than simple greedy selection.

## Draft Model Architecture in MTPLX

### Exact-Speculative Loop Design

The codebase implements an exact-speculative loop where a secondary draft model generates candidate sequences before the primary target model verifies them. This design appears in [`tests/test_tail_gemma4_stream_holdback.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_tail_gemma4_stream_holdback.py), which references an exact-speculative assistant pattern that validates draft outputs against the target distribution. This approach differs from greedy-argmax methods that select only the highest probability token at each step without speculative lookahead.

### Speculative Depth and Draft Configuration

MTPLX controls speculative generation through the `speculative_depth` parameter observed in [`tests/test_deepseek_v4_spec.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_deepseek_v4_spec.py). This parameter determines how many tokens the draft model produces before the target model performs verification, allowing the system to balance parallelism against draft model accuracy. The configuration explicitly separates the draft and target responsibilities using distinct core identifiers.

## Implementation in mtplx/sampling.py

### SamplerConfig Parameters

The `SamplerConfig` class defines the draft and target model relationship through two key fields:
- `draft_core`: Identifies the secondary model responsible for speculative token generation
- `verify_core`: Specifies the target model that validates draft outputs and executes rollbacks on mismatch

### The speculative_output_marginal Function

The core verification logic resides in `speculative_output_marginal`, which accepts target probabilities (`target_p`), draft probabilities (`draft_q`), and the configuration object. This function compares draft-generated candidates against the target distribution, accepting tokens sequentially until detecting a mismatch.

```python
from mtplx.sampling import SamplerConfig, speculative_output_marginal

config = SamplerConfig(
    draft_core="stock",          # draft model identifier

    speculative_depth=2,         # number of draft tokens to generate

    verify_core="target",        # target model for verification

)

# Run decoding – draft model produces candidates, target model verifies

output = speculative_output_marginal(target_p, draft_q, config)

```

## Key Source Files

- [`tests/test_tail_gemma4_stream_holdback.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_tail_gemma4_stream_holdback.py): Demonstrates the exact-speculative assistant pattern using a dedicated draft model stage before verification.
- [`tests/test_deepseek_v4_spec.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_deepseek_v4_spec.py): Validates the `speculative_depth` parameter controlling the volume of draft tokens generated before target model verification.
- [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py): Contains the `speculative_output_marginal` implementation and the `SamplerConfig` class definition that orchestrates the draft-target interaction.

## Summary

- MTPLX uses a **second draft model** configured via `draft_core` rather than greedy-argmax for speculative decoding.
- The `speculative_depth` parameter in `SamplerConfig` controls how many tokens are generated before target model verification.
- The `speculative_output_marginal` function in [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py) executes the verification logic and handles rollback to the last accepted token.
- Test suites in [`test_tail_gemma4_stream_holdback.py`](https://github.com/youssofal/MTPLX/blob/main/test_tail_gemma4_stream_holdback.py) and [`test_deepseek_v4_spec.py`](https://github.com/youssofal/MTPLX/blob/main/test_deepseek_v4_spec.py) confirm the exact-speculative loop implementation with distinct draft and target stages.

## Frequently Asked Questions

### Does MTPLX use greedy-argmax for speculative decoding?

No. According to the `youssofal/MTPLX` source code, the framework employs a second draft model specified by the `draft_core` parameter to generate candidate tokens. The target model verifies these candidates against its own distribution, which differs fundamentally from greedy-argmax selection that would only consider the highest probability single token at each step.

### What does the speculative_depth parameter control?

The `speculative_depth` parameter, evidenced in [`tests/test_deepseek_v4_spec.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_deepseek_v4_spec.py), determines the number of tokens the draft model generates before the target model executes verification. This allows the system to balance between parallelization (higher depth) and the draft model's distribution alignment with the target.

### How does speculative_output_marginal handle token verification?

This function, implemented in [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py), accepts the target model probabilities (`target_p`), draft model probabilities (`draft_q`), and a `SamplerConfig` instance. It validates draft tokens sequentially against the target distribution, accepting matches until encountering a discrepancy, then rolls back to the last verified position.

### Where is the draft model specified in MTPLX?

The draft model is configured through the `draft_core` field in the `SamplerConfig` class, as demonstrated in the sampling module and validated by the test suite. This separation between `draft_core` (generator) and `verify_core` (validator) enables the dual-model speculative architecture.