# How to Auto-Tune Draft Depth in MTPLX for Maximum Inference Speed

> Learn how to auto-tune draft depth in MTPLX. Optimize your inference speed by automatically finding the best MTP depth for maximum tokens per second. Get faster AI models today.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-08

---

**MTPLX automatically determines the optimal draft (MTP) depth by executing a performance sweep on the real model at each supported depth, measuring tokens-per-second while pinning fans to eliminate thermal throttling contamination.**

Auto-tuning draft depth in MTPLX ensures your Apple Silicon Mac runs multi-token prediction at peak efficiency. The open-source project **youssofal/MTPLX** implements a sophisticated benchmarking routine that empirically measures each depth level rather than relying on heuristics. This process guarantees that the selected depth actually improves inference speed on your specific hardware configuration.

## The Five-Step Auto-Tuning Process

The auto-tuning workflow follows a rigorous measurement protocol to isolate the fastest configuration.

### 1. Triggering the Benchmark

The process initiates either during first-run onboarding or manually via the CLI. When triggered, MTPLX launches the selected model once per draft depth while keeping the **fans pinned** to ensure thermal throttling does not contaminate the timing data. As documented in the repository’s README, "MTPLX runs the real model on your machine at each depth, with fans pinned for clean timing"【/cache/repos/github.com/youssofal/MTPLX/main/README.md†L55-L61】.

### 2. Establishing the Baseline

Before testing draft depths, MTPLX measures the **autoregressive (non-draft) decoding speed** to establish a reference baseline. This value serves as the performance threshold that any draft depth must exceed to be considered viable.

### 3. Depth-Wise Performance Sweep

For each supported depth `d = 1, 2, …`, the model generates a fixed number of tokens (typically a few seconds of text). MTPLX records the **tokens-per-second (tok/s)** achieved at each specific draft depth, creating a performance profile unique to your Mac’s thermal and computational characteristics.

### 4. Selection Logic

The system selects the depth that yields the highest tok/s **and** improves upon the baseline measurement. According to the README, "If an MTP depth beats it, that depth is saved. If nothing beats the baseline, nothing is saved and the app says so"【/cache/repos/github.com/youssofal/MTPLX/main/README.md†L55-L62】. This prevents the system from adopting a configuration that would slow down inference.

### 5. Persisting the Configuration

Once identified, the optimal depth is written to the user configuration file at `~/.mtplx/config.toml` under the `draft_block_size` parameter. This value persists across subsequent runs unless manually overridden or re-tuned.

## Core Implementation Files

The auto-tuning logic is centralized in specific modules that handle orchestration, benchmarking, and persistence.

**[`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py)** contains the primary implementation through the `run_autotune` function. This function orchestrates the sweep, collects timing data for each depth, and returns the optimal block size.

**[`mtplx/__main__.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/__main__.py)** registers the CLI interface using the Typer framework. It maps the `tune` sub-command to the `run_autotune` routine, parsing arguments such as `--model` and `--retune` before invocation.

**[`mtplx/config.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/config.py)** manages the persistence layer, handling both loading existing settings and saving the newly determined `draft_block_size` to disk.

## Running Auto-Tuning via CLI and Python

You can trigger the auto-tuning process directly from the terminal or integrate it programmatically into larger workflows.

Execute the benchmark from the command line to automatically find and save the best depth:

```bash

# Automatic tuning for a model (or local path) – runs the sweep and saves the best depth

mtplx tune --model Qwen3-8B-Optimized-Speed --retune

```

For programmatic control, import the `run_autotune` function directly from the thermal module:

```python

# Programmatic use of the autotuning routine (e.g., from a script)

from mtplx.thermal import run_autotune
from mtplx.config import load_config, save_config

cfg = load_config()
best_depth = run_autotune(model=cfg.model, retune=True)
cfg.draft_block_size = best_depth
save_config(cfg)
print(f"Auto‑tuned draft depth: {best_depth}")

```

## Summary

- **MTPLX** determines optimal draft depth through empirical measurement rather than estimation, running the real model at each supported depth level.
- The **`run_autotune`** function in [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py) governs the process, while [`mtplx/__main__.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/__main__.py) exposes it via the `tune` CLI command.
- Benchmarking employs **pinned fans** to eliminate thermal throttling variables, ensuring accurate tokens-per-second measurements.
- Only depths that **beat the autoregressive baseline** are accepted; if none perform better, the configuration remains unchanged.
- Results persist in **`~/.mtplx/config.toml`** as `draft_block_size`, applying automatically to future inference sessions.

## Frequently Asked Questions

### How does MTPLX prevent thermal throttling from affecting benchmark results?

During the auto-tuning sweep, MTPLX pins the Mac’s fans to a fixed speed to prevent thermal throttling from skewing the timing measurements. This ensures that performance differences between depths reflect actual computational efficiency rather than thermal variations, providing clean data for the selection algorithm in [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py).

### What happens if no draft depth improves on the baseline speed?

If no tested depth exceeds the autoregressive baseline tokens-per-second measurement, MTPLX preserves the existing configuration and notifies the user that no superior setting was found. This safeguard prevents the system from adopting a draft depth that would degrade inference performance compared to standard autoregressive decoding.

### Where does MTPLX store the auto-tuned draft depth setting?

The chosen depth is persisted to the user configuration file located at `~/.mtplx/config.toml` under the `draft_block_size` key. The `save_config` function in [`mtplx/config.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/config.py) handles this write operation, ensuring the setting persists across application restarts.

### Can I integrate auto-tuning into a custom Python script?

Yes, you can import `run_autotune` directly from `mtplx.thermal` and invoke it programmatically with your model configuration. This allows you to embed auto-tuning workflows into deployment scripts, automated testing pipelines, or custom applications that manage multiple model configurations dynamically.