How to Auto-Tune Draft Depth in MTPLX for Maximum Inference Speed
MTPLX automatically determines the optimal draft (MTP) depth by executing a performance sweep on the real model at each supported depth, measuring tokens-per-second while pinning fans to eliminate thermal throttling contamination.
Auto-tuning draft depth in MTPLX ensures your Apple Silicon Mac runs multi-token prediction at peak efficiency. The open-source project youssofal/MTPLX implements a sophisticated benchmarking routine that empirically measures each depth level rather than relying on heuristics. This process guarantees that the selected depth actually improves inference speed on your specific hardware configuration.
The Five-Step Auto-Tuning Process
The auto-tuning workflow follows a rigorous measurement protocol to isolate the fastest configuration.
1. Triggering the Benchmark
The process initiates either during first-run onboarding or manually via the CLI. When triggered, MTPLX launches the selected model once per draft depth while keeping the fans pinned to ensure thermal throttling does not contaminate the timing data. As documented in the repository’s README, "MTPLX runs the real model on your machine at each depth, with fans pinned for clean timing"【/cache/repos/github.com/youssofal/MTPLX/main/README.md†L55-L61】.
2. Establishing the Baseline
Before testing draft depths, MTPLX measures the autoregressive (non-draft) decoding speed to establish a reference baseline. This value serves as the performance threshold that any draft depth must exceed to be considered viable.
3. Depth-Wise Performance Sweep
For each supported depth d = 1, 2, …, the model generates a fixed number of tokens (typically a few seconds of text). MTPLX records the tokens-per-second (tok/s) achieved at each specific draft depth, creating a performance profile unique to your Mac’s thermal and computational characteristics.
4. Selection Logic
The system selects the depth that yields the highest tok/s and improves upon the baseline measurement. According to the README, "If an MTP depth beats it, that depth is saved. If nothing beats the baseline, nothing is saved and the app says so"【/cache/repos/github.com/youssofal/MTPLX/main/README.md†L55-L62】. This prevents the system from adopting a configuration that would slow down inference.
5. Persisting the Configuration
Once identified, the optimal depth is written to the user configuration file at ~/.mtplx/config.toml under the draft_block_size parameter. This value persists across subsequent runs unless manually overridden or re-tuned.
Core Implementation Files
The auto-tuning logic is centralized in specific modules that handle orchestration, benchmarking, and persistence.
mtplx/thermal.py contains the primary implementation through the run_autotune function. This function orchestrates the sweep, collects timing data for each depth, and returns the optimal block size.
mtplx/__main__.py registers the CLI interface using the Typer framework. It maps the tune sub-command to the run_autotune routine, parsing arguments such as --model and --retune before invocation.
mtplx/config.py manages the persistence layer, handling both loading existing settings and saving the newly determined draft_block_size to disk.
Running Auto-Tuning via CLI and Python
You can trigger the auto-tuning process directly from the terminal or integrate it programmatically into larger workflows.
Execute the benchmark from the command line to automatically find and save the best depth:
# Automatic tuning for a model (or local path) – runs the sweep and saves the best depth
mtplx tune --model Qwen3-8B-Optimized-Speed --retune
For programmatic control, import the run_autotune function directly from the thermal module:
# Programmatic use of the autotuning routine (e.g., from a script)
from mtplx.thermal import run_autotune
from mtplx.config import load_config, save_config
cfg = load_config()
best_depth = run_autotune(model=cfg.model, retune=True)
cfg.draft_block_size = best_depth
save_config(cfg)
print(f"Auto‑tuned draft depth: {best_depth}")
Summary
- MTPLX determines optimal draft depth through empirical measurement rather than estimation, running the real model at each supported depth level.
- The
run_autotunefunction inmtplx/thermal.pygoverns the process, whilemtplx/__main__.pyexposes it via thetuneCLI command. - Benchmarking employs pinned fans to eliminate thermal throttling variables, ensuring accurate tokens-per-second measurements.
- Only depths that beat the autoregressive baseline are accepted; if none perform better, the configuration remains unchanged.
- Results persist in
~/.mtplx/config.tomlasdraft_block_size, applying automatically to future inference sessions.
Frequently Asked Questions
How does MTPLX prevent thermal throttling from affecting benchmark results?
During the auto-tuning sweep, MTPLX pins the Mac’s fans to a fixed speed to prevent thermal throttling from skewing the timing measurements. This ensures that performance differences between depths reflect actual computational efficiency rather than thermal variations, providing clean data for the selection algorithm in mtplx/thermal.py.
What happens if no draft depth improves on the baseline speed?
If no tested depth exceeds the autoregressive baseline tokens-per-second measurement, MTPLX preserves the existing configuration and notifies the user that no superior setting was found. This safeguard prevents the system from adopting a draft depth that would degrade inference performance compared to standard autoregressive decoding.
Where does MTPLX store the auto-tuned draft depth setting?
The chosen depth is persisted to the user configuration file located at ~/.mtplx/config.toml under the draft_block_size key. The save_config function in mtplx/config.py handles this write operation, ensuring the setting persists across application restarts.
Can I integrate auto-tuning into a custom Python script?
Yes, you can import run_autotune directly from mtplx.thermal and invoke it programmatically with your model configuration. This allows you to embed auto-tuning workflows into deployment scripts, automated testing pipelines, or custom applications that manage multiple model configurations dynamically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →