# How the MTPLX Auto-Tune System Measures and Selects Optimal MTP Depth

> Discover how the MTPLX auto-tune system benchmarks MTP depths against an AR baseline, measuring throughput and quality to select optimal settings for peak performance.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: internals
- Published: 2026-09-13

---

**The MTPLX auto-tune system benchmarks each candidate native-MTP depth against an autoregressive (AR) baseline, measuring throughput and draft-acceptance quality to select the deepest depth that outperforms AR without degrading output quality.**

MTPLX (youssofal/MTPLX) implements Multi-Token Prediction (MTP) to accelerate large language model inference, but optimal performance depends on selecting the right **native-MTP depth** for each specific model and hardware configuration. The `mtplx tune` command eliminates manual guesswork by automatically discovering the fastest valid depth through a systematic three-stage measurement pipeline.

## The Three-Stage Auto-Tuning Pipeline

The auto-tune system orchestrates candidate preparation, controlled benchmarking, and quality-aware selection through tightly coupled stages implemented in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py).

### Stage 1: Candidate Preparation with `_parse_tune_candidate_values`

The pipeline begins by resolving which depths to test. The `_parse_tune_candidate_values` function reads the `--depths` CLI argument (defaulting to `1,2,3`) and constructs a candidate list that always includes the autoregressive "AR" baseline as the control.

The system extracts the specific control field that governs MTP depth from the model’s **tune support payload** (`_tune_control_field`), ensuring the benchmark manipulates the correct internal parameter for each candidate run.

### Stage 2: Benchmark Execution via `_run_tune_candidates`

For each candidate depth, `_run_tune_candidates` executes a lightweight probe using [`mtplx/scripts/mtp_cost_probe.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/scripts/mtp_cost_probe.py). Before timing begins, the system establishes thermal stability by invoking **MaxSession** from [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py) to pin cooling fans to max-fan mode, preventing thermal throttling from skewing latency measurements.

During each probe run, the system collects three critical telemetry signals via `forkev_telemetry`:

- **Throughput**: Raw tokens-per-second generation speed
- **Draft-acceptance rate**: Percentage of MTP draft tokens retained in the final output
- **Quality flag**: Binary indicator marking whether any draft row failed quality validation

Any depth exhibiting quality failures during the probe is immediately flagged for exclusion regardless of speed.

### Stage 3: Decision Logic in `_select_best_mtp_depth`

After all candidates complete, `_select_best_mtp_depth` executes a deterministic selection algorithm:

1. **Quality filtering**: Removes any depth with failed quality flags
2. **Speed comparison**: Identifies which remaining depths exceed the AR baseline throughput
3. **Depth maximization**: Applies the "deeper-depth-within-noise-band" rule, selecting the deepest depth that matches or beats the fastest speed within measurement tolerance
4. **Fallback handling**: If no depth passes quality filters or beats AR speed, the system falls back to AR mode and emits: `No quality-passed MTP depth beat AR; falling back to AR`

The chosen depth is persisted to the model’s profile in [`mtplx/profiles.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py) and optionally cached in a tune-state file for subsequent inference sessions.

## Running the Auto-Tune Command

You can initiate the measurement pipeline using the `mtplx tune` CLI:

```bash

# Run with default depths (1,2,3) against a specific model

mtplx tune --model Qwen3.8-27B-MTPLX-Optimized-Speed

# Restrict the search space to specific depths

mtplx tune --depths 2,4,6 --model my-model

# Preview the candidate selection without executing benchmarks

mtplx tune --dry-run --depths 1,2,3 --model my-model

```

Typical successful output concludes with:

```

[tune] best MTP depth: 3 (speed 23.5 t/s, draft-accept 92 %)

```

When the system selects a depth, future inference automatically uses this value:

```python
from mtplx.generation import generate

# Uses the depth discovered by `mtplx tune` without manual configuration

generate(prompt="Explain quantum tunnelling.", model="my-model")

```

## How the Selection Algorithm Handles Edge Cases

The decision logic in [`public.py`](https://github.com/youssofal/MTPLX/blob/main/public.py) implements specific safeguards to ensure robust performance across hardware variations.

**Thermal Stability Guarantees**
Because clock speeds fluctuate with temperature, the MaxSession context manager in [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py) forces maximum fan speed during benchmarking. This eliminates thermal throttling as a confounding variable when comparing shallow versus deep MTP depths that may have different power consumption profiles.

**Quality-First Filtering**
The system prioritizes output integrity over raw speed. If depth 4 achieves 30% higher throughput than AR but exhibits any quality-failed draft rows, it is disqualified before speed comparisons occur. This prevents the auto-tune system from selecting a depth that corrupts model outputs.

**Noise-Band Tie Breaking**
When multiple depths yield statistically similar throughputs (within the configured noise band), the algorithm prefers the **deepest** valid depth. This heuristic maximizes speculative execution potential while maintaining the verified performance floor.

## Summary

- The MTPLX auto-tune system is invoked via `mtplx tune` and executes a three-stage pipeline defined in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py).
- Candidate depths default to `1,2,3` but are configurable via `--depths`, always including an AR baseline for comparison.
- The benchmark uses [`mtp_cost_probe.py`](https://github.com/youssofal/MTPLX/blob/main/mtp_cost_probe.py) and `forkev_telemetry` to measure throughput and draft-acceptance rates while [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py) ensures thermal stability.
- Selection occurs in `_select_best_mtp_depth`, which filters for quality, requires beating AR speed, and applies a deepest-depth tie-breaker.
- Results persist in [`mtplx/profiles.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py), making the optimal depth available to the generation engine without manual configuration.

## Frequently Asked Questions

### What telemetry does MTPLX collect during auto-tuning?

The system collects **throughput** (tokens per second), **draft-acceptance rate** (percentage of MTP tokens kept), and a **quality flag** indicating whether any draft rows failed validation. These metrics are gathered via `forkev_telemetry` inside the [`mtp_cost_probe.py`](https://github.com/youssofal/MTPLX/blob/main/mtp_cost_probe.py) script.

### How does MTPLX prevent thermal throttling from affecting benchmark results?

Before executing timing measurements, the `_run_tune_candidates` function initializes a **MaxSession** from [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py) that pins system fans to maximum speed. This stabilizes CPU/GPU temperatures so that performance differences reflect actual MTP depth efficiency rather than thermal variability.

### What happens if no MTP depth beats the autoregressive baseline?

If all candidate depths either fail quality checks or fail to exceed AR throughput, the system falls back to AR mode, prints `No quality-passed MTP depth beat AR`, and stores the AR configuration in the model profile. Inference will proceed without MTP speculation for that model.

### Where does MTPLX store the selected optimal depth?

The chosen depth is written to the model’s profile managed by [`mtplx/profiles.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py) and optionally cached in a tune-state file. The `generate` function in [`mtplx/generation.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py) automatically reads this profile to configure the native-MTP depth for subsequent inference calls.