How to Auto-Tune Draft Depth in MTPLX for Maximum Inference Speed

MTPLX automatically determines the optimal draft (MTP) depth by executing a performance sweep on the real model at each supported depth, measuring tokens-per-second while pinning fans to eliminate thermal throttling contamination.

Auto-tuning draft depth in MTPLX ensures your Apple Silicon Mac runs multi-token prediction at peak efficiency. The open-source project youssofal/MTPLX implements a sophisticated benchmarking routine that empirically measures each depth level rather than relying on heuristics. This process guarantees that the selected depth actually improves inference speed on your specific hardware configuration.

The Five-Step Auto-Tuning Process

The auto-tuning workflow follows a rigorous measurement protocol to isolate the fastest configuration.

1. Triggering the Benchmark

The process initiates either during first-run onboarding or manually via the CLI. When triggered, MTPLX launches the selected model once per draft depth while keeping the fans pinned to ensure thermal throttling does not contaminate the timing data. As documented in the repository’s README, "MTPLX runs the real model on your machine at each depth, with fans pinned for clean timing"【/cache/repos/github.com/youssofal/MTPLX/main/README.md†L55-L61】.

2. Establishing the Baseline

Before testing draft depths, MTPLX measures the autoregressive (non-draft) decoding speed to establish a reference baseline. This value serves as the performance threshold that any draft depth must exceed to be considered viable.

3. Depth-Wise Performance Sweep

For each supported depth d = 1, 2, …, the model generates a fixed number of tokens (typically a few seconds of text). MTPLX records the tokens-per-second (tok/s) achieved at each specific draft depth, creating a performance profile unique to your Mac’s thermal and computational characteristics.

4. Selection Logic

The system selects the depth that yields the highest tok/s and improves upon the baseline measurement. According to the README, "If an MTP depth beats it, that depth is saved. If nothing beats the baseline, nothing is saved and the app says so"【/cache/repos/github.com/youssofal/MTPLX/main/README.md†L55-L62】. This prevents the system from adopting a configuration that would slow down inference.

5. Persisting the Configuration

Once identified, the optimal depth is written to the user configuration file at ~/.mtplx/config.toml under the draft_block_size parameter. This value persists across subsequent runs unless manually overridden or re-tuned.

Core Implementation Files

The auto-tuning logic is centralized in specific modules that handle orchestration, benchmarking, and persistence.

mtplx/thermal.py contains the primary implementation through the run_autotune function. This function orchestrates the sweep, collects timing data for each depth, and returns the optimal block size.

mtplx/__main__.py registers the CLI interface using the Typer framework. It maps the tune sub-command to the run_autotune routine, parsing arguments such as --model and --retune before invocation.

mtplx/config.py manages the persistence layer, handling both loading existing settings and saving the newly determined draft_block_size to disk.

Running Auto-Tuning via CLI and Python

You can trigger the auto-tuning process directly from the terminal or integrate it programmatically into larger workflows.

Execute the benchmark from the command line to automatically find and save the best depth:


# Automatic tuning for a model (or local path) – runs the sweep and saves the best depth

mtplx tune --model Qwen3-8B-Optimized-Speed --retune

For programmatic control, import the run_autotune function directly from the thermal module:


# Programmatic use of the autotuning routine (e.g., from a script)

from mtplx.thermal import run_autotune
from mtplx.config import load_config, save_config

cfg = load_config()
best_depth = run_autotune(model=cfg.model, retune=True)
cfg.draft_block_size = best_depth
save_config(cfg)
print(f"Auto‑tuned draft depth: {best_depth}")

Summary

  • MTPLX determines optimal draft depth through empirical measurement rather than estimation, running the real model at each supported depth level.
  • The run_autotune function in mtplx/thermal.py governs the process, while mtplx/__main__.py exposes it via the tune CLI command.
  • Benchmarking employs pinned fans to eliminate thermal throttling variables, ensuring accurate tokens-per-second measurements.
  • Only depths that beat the autoregressive baseline are accepted; if none perform better, the configuration remains unchanged.
  • Results persist in ~/.mtplx/config.toml as draft_block_size, applying automatically to future inference sessions.

Frequently Asked Questions

How does MTPLX prevent thermal throttling from affecting benchmark results?

During the auto-tuning sweep, MTPLX pins the Mac’s fans to a fixed speed to prevent thermal throttling from skewing the timing measurements. This ensures that performance differences between depths reflect actual computational efficiency rather than thermal variations, providing clean data for the selection algorithm in mtplx/thermal.py.

What happens if no draft depth improves on the baseline speed?

If no tested depth exceeds the autoregressive baseline tokens-per-second measurement, MTPLX preserves the existing configuration and notifies the user that no superior setting was found. This safeguard prevents the system from adopting a draft depth that would degrade inference performance compared to standard autoregressive decoding.

Where does MTPLX store the auto-tuned draft depth setting?

The chosen depth is persisted to the user configuration file located at ~/.mtplx/config.toml under the draft_block_size key. The save_config function in mtplx/config.py handles this write operation, ensuring the setting persists across application restarts.

Can I integrate auto-tuning into a custom Python script?

Yes, you can import run_autotune directly from mtplx.thermal and invoke it programmatically with your model configuration. This allows you to embed auto-tuning workflows into deployment scripts, automated testing pipelines, or custom applications that manage multiple model configurations dynamically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →