Expected Speedup from Tuning Draft Depth with MTPLX: Performance Guide
Tuning draft depth with MTPLX delivers approximately 15–20% faster decoding over baseline autoregressive generation, with hardware-specific optimizations reaching up to 2.24× speedup on Apple Silicon Macs.
MTPLX is a speculative decoding framework that accelerates large language model inference by leveraging the model's own Multi-Token Prediction (MTP) head for draft generation. By automatically calibrating the number of speculative steps (draft depth) to your specific hardware and prompt characteristics, MTPLX eliminates manual parameter tuning while maximizing throughput. This guide breaks down the exact performance gains you can expect and how the runtime selects the optimal configuration.
Understanding Draft Depth in MTPLX
What is Draft Depth?
In speculative decoding, draft depth refers to the number of tokens the model attempts to generate speculatively before running a verification pass. A higher depth increases the potential batch size for parallel verification but also raises computational overhead if the draft tokens are rejected.
The MTP Head Architecture
MTPLX uses a lightweight Multi-Token Prediction head implemented as an optimized kernel. According to the source code in mtplx/cli.py, this kernel is particularly efficient on Apple Silicon, where the Metal GPU implementation allows the draft head to run with minimal latency overhead compared to standard autoregressive forward passes.
Expected Speedup Benchmarks
The MTPLX runtime measures sustained throughput across different depths to determine the optimal configuration. Based on release benchmarks documented in docs/releases/v2.9.0.md, docs/releases/v2.7.0.md, and docs/releases/v2.9.2.md, the expected speedups are:
| Configuration | Hardware | Speedup vs. Autoregressive |
|---|---|---|
| Default depth (3) | 16 GB M4 Mac mini | ~1.6× faster |
| Default depth (3) | M5 Max | Up to 2.24× faster (~124% gain) |
| Greedy draft (depth = 1) | M5 Max (short prompts ≤8k) | +2.5% to +9.8% decode speed |
| Depth-tuned per-machine | Various | +19% over static depth-3 baseline |
| Optimal depth-3 drafting | Various | 15–20% faster than MTPLX v2.8 |
These measurements are derived from the benchmarks/repro_a3b_depth_default.py harness, which reports tokens-per-second and acceptance rates for each candidate depth.
Why Speedup Varies by Configuration
Hardware Optimization
The MTP draft head kernels are heavily optimized for Metal GPU execution. As noted in docs/releases/v2.9.0.md, Apple Silicon devices show higher relative gains because the speculative kernels efficiently utilize the unified memory architecture, minimizing the data transfer overhead that typically bottlenecks speculative decoding on discrete GPUs.
Prompt Length Dynamics
For short prompts (≤8,000 tokens), greedy drafting with depth = 1 provides measurable gains of 2.5% to 9.8%. However, the runtime disables greedy mode for longer contexts via the MTPLX_GREEDY_TRIO_MAX_CONTEXT fence, as the verification overhead outweighs the drafting benefits beyond this threshold.
Model-Specific Depth Caps
Most public model packs distributed with MTPLX enforce a maximum draft depth of 3 (mtp_depth_max = 3). Experiments with depths 4–8 were conducted but showed reduced throughput due to increased verification rejection rates, leading the project to ship depth-3 as the default optimum.
The Tuning Process
When you run mtplx tune, the system warms each candidate depth by loading model weights and compiling kernels before measuring sustained throughput. This ensures the selected depth reflects real-world performance rather than cold-start latency, accounting for the approximately 19% improvement over untuned configurations.
How to Tune Draft Depth in MTPLX
The CLI provides specific commands to detect and apply the optimal depth for your machine:
Run the automatic tuner to detect the best depth for your hardware:
mtplx tune
The output displays the optimal depth (typically "Depth = 3 (optimal)") and caches this configuration for subsequent runs.
Serve the model using the tuned depth automatically:
mtplx serve --generation-mode mtp
The runtime reads the cached depth from the configuration file; no manual --depth flag is required.
Force a specific depth for experimental comparison:
mtplx serve --generation-mode mtp --depth 2
Run a depth-sweep benchmark to see per-depth performance:
mtplx bench --harness depth-sweep --model qwen-3.8b
This executes the harness defined in benchmarks/repro_a3b_depth_default.py, printing a table of tok/s and acceptance rates for depths 1–3.
Key Source Files and Implementation Details
Understanding the following files clarifies how MTPLX determines and applies draft depth:
docs/releases/v2.9.0.md: Documents the depth tuner warming process and the validation that depth-3 configurations outperform deeper speculative chains.docs/releases/v2.7.0.md: Contains the +19% performance fix for tuned depth selection.docs/releases/v2.9.2.md: Details the M5 Max greedy drafting benchmarks showing 2.5–9.8% gains.mtplx/cli.py: Implements themtplx tunecommand and manages the runtime configuration cache.benchmarks/repro_a3b_depth_default.py: Houses the depth-sweep harness used for empirical throughput measurement.
Summary
- Tuning draft depth with MTPLX provides a 15–20% speedup over standard autoregressive decoding, with best-case scenarios reaching 2.0–2.24× on optimized hardware.
- The default depth of 3 represents the shipped optimum for most model packs, balancing draft acceptance rates against verification overhead.
- The
mtplx tunecommand detects hardware-specific optima by warming kernels and measuring sustained throughput, improving performance by approximately 19% compared to static defaults. - Greedy drafting (depth = 1) benefits short prompts (≤8k tokens) but is automatically gated by
MTPLX_GREEDY_TRIO_MAX_CONTEXTfor longer contexts. - Depths beyond 3 are experimentally supported but generally reduce throughput due to verification rejection penalties.
Frequently Asked Questions
What is the maximum draft depth supported by MTPLX?
Most production model packs cap the draft depth at 3 via the mtp_depth_max = 3 parameter. While the codebase experimentally supports depths up to 8, these configurations consistently showed lower throughput than depth-3 in benchmarks, leading the project to enforce depth-3 as the practical maximum for shipped packs.
How does mtplx tune improve performance over default settings?
The tuner improves performance by approximately 19% over static depth-3 baselines. As implemented in mtplx/cli.py, it warms up each candidate depth (loading weights and compiling kernels) before measurement, ensuring the selection reflects the fastest sustained throughput on your specific hardware rather than transient cold-start behavior.
When should I use greedy drafting (depth = 1) instead of depth tuning?
Greedy drafting provides a 2.5% to 9.8% speed increase specifically for short prompts under 8,000 tokens on high-end Apple Silicon like the M5 Max. The runtime automatically disables this mode for longer contexts via the MTPLX_GREEDY_TRIO_MAX_CONTEXT environment fence, as the verification overhead outweighs the benefits for extended sequences.
Can I override the tuned depth for specific workloads?
Yes. While mtplx serve --generation-mode mtp automatically uses the cached tuner result, you can override it with the --depth flag (e.g., --depth 2). This is useful for A/B testing or when working with prompts of unusual length that might benefit from non-standard depths not selected by the automatic tuner.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →