How the MTPLX Auto-Tune System Measures and Selects Optimal MTP Depth
The MTPLX auto-tune system benchmarks each candidate native-MTP depth against an autoregressive (AR) baseline, measuring throughput and draft-acceptance quality to select the deepest depth that outperforms AR without degrading output quality.
MTPLX (youssofal/MTPLX) implements Multi-Token Prediction (MTP) to accelerate large language model inference, but optimal performance depends on selecting the right native-MTP depth for each specific model and hardware configuration. The mtplx tune command eliminates manual guesswork by automatically discovering the fastest valid depth through a systematic three-stage measurement pipeline.
The Three-Stage Auto-Tuning Pipeline
The auto-tune system orchestrates candidate preparation, controlled benchmarking, and quality-aware selection through tightly coupled stages implemented in mtplx/commands/public.py.
Stage 1: Candidate Preparation with _parse_tune_candidate_values
The pipeline begins by resolving which depths to test. The _parse_tune_candidate_values function reads the --depths CLI argument (defaulting to 1,2,3) and constructs a candidate list that always includes the autoregressive "AR" baseline as the control.
The system extracts the specific control field that governs MTP depth from the model’s tune support payload (_tune_control_field), ensuring the benchmark manipulates the correct internal parameter for each candidate run.
Stage 2: Benchmark Execution via _run_tune_candidates
For each candidate depth, _run_tune_candidates executes a lightweight probe using mtplx/scripts/mtp_cost_probe.py. Before timing begins, the system establishes thermal stability by invoking MaxSession from mtplx/thermal.py to pin cooling fans to max-fan mode, preventing thermal throttling from skewing latency measurements.
During each probe run, the system collects three critical telemetry signals via forkev_telemetry:
- Throughput: Raw tokens-per-second generation speed
- Draft-acceptance rate: Percentage of MTP draft tokens retained in the final output
- Quality flag: Binary indicator marking whether any draft row failed quality validation
Any depth exhibiting quality failures during the probe is immediately flagged for exclusion regardless of speed.
Stage 3: Decision Logic in _select_best_mtp_depth
After all candidates complete, _select_best_mtp_depth executes a deterministic selection algorithm:
- Quality filtering: Removes any depth with failed quality flags
- Speed comparison: Identifies which remaining depths exceed the AR baseline throughput
- Depth maximization: Applies the "deeper-depth-within-noise-band" rule, selecting the deepest depth that matches or beats the fastest speed within measurement tolerance
- Fallback handling: If no depth passes quality filters or beats AR speed, the system falls back to AR mode and emits:
No quality-passed MTP depth beat AR; falling back to AR
The chosen depth is persisted to the model’s profile in mtplx/profiles.py and optionally cached in a tune-state file for subsequent inference sessions.
Running the Auto-Tune Command
You can initiate the measurement pipeline using the mtplx tune CLI:
# Run with default depths (1,2,3) against a specific model
mtplx tune --model Qwen3.8-27B-MTPLX-Optimized-Speed
# Restrict the search space to specific depths
mtplx tune --depths 2,4,6 --model my-model
# Preview the candidate selection without executing benchmarks
mtplx tune --dry-run --depths 1,2,3 --model my-model
Typical successful output concludes with:
[tune] best MTP depth: 3 (speed 23.5 t/s, draft-accept 92 %)
When the system selects a depth, future inference automatically uses this value:
from mtplx.generation import generate
# Uses the depth discovered by `mtplx tune` without manual configuration
generate(prompt="Explain quantum tunnelling.", model="my-model")
How the Selection Algorithm Handles Edge Cases
The decision logic in public.py implements specific safeguards to ensure robust performance across hardware variations.
Thermal Stability Guarantees
Because clock speeds fluctuate with temperature, the MaxSession context manager in mtplx/thermal.py forces maximum fan speed during benchmarking. This eliminates thermal throttling as a confounding variable when comparing shallow versus deep MTP depths that may have different power consumption profiles.
Quality-First Filtering The system prioritizes output integrity over raw speed. If depth 4 achieves 30% higher throughput than AR but exhibits any quality-failed draft rows, it is disqualified before speed comparisons occur. This prevents the auto-tune system from selecting a depth that corrupts model outputs.
Noise-Band Tie Breaking When multiple depths yield statistically similar throughputs (within the configured noise band), the algorithm prefers the deepest valid depth. This heuristic maximizes speculative execution potential while maintaining the verified performance floor.
Summary
- The MTPLX auto-tune system is invoked via
mtplx tuneand executes a three-stage pipeline defined inmtplx/commands/public.py. - Candidate depths default to
1,2,3but are configurable via--depths, always including an AR baseline for comparison. - The benchmark uses
mtp_cost_probe.pyandforkev_telemetryto measure throughput and draft-acceptance rates whilemtplx/thermal.pyensures thermal stability. - Selection occurs in
_select_best_mtp_depth, which filters for quality, requires beating AR speed, and applies a deepest-depth tie-breaker. - Results persist in
mtplx/profiles.py, making the optimal depth available to the generation engine without manual configuration.
Frequently Asked Questions
What telemetry does MTPLX collect during auto-tuning?
The system collects throughput (tokens per second), draft-acceptance rate (percentage of MTP tokens kept), and a quality flag indicating whether any draft rows failed validation. These metrics are gathered via forkev_telemetry inside the mtp_cost_probe.py script.
How does MTPLX prevent thermal throttling from affecting benchmark results?
Before executing timing measurements, the _run_tune_candidates function initializes a MaxSession from mtplx/thermal.py that pins system fans to maximum speed. This stabilizes CPU/GPU temperatures so that performance differences reflect actual MTP depth efficiency rather than thermal variability.
What happens if no MTP depth beats the autoregressive baseline?
If all candidate depths either fail quality checks or fail to exceed AR throughput, the system falls back to AR mode, prints No quality-passed MTP depth beat AR, and stores the AR configuration in the model profile. Inference will proceed without MTP speculation for that model.
Where does MTPLX store the selected optimal depth?
The chosen depth is written to the model’s profile managed by mtplx/profiles.py and optionally cached in a tune-state file. The generate function in mtplx/generation.py automatically reads this profile to configure the native-MTP depth for subsequent inference calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →