# How to Auto-Tune Draft Depth in MTPLX for Specific Hardware: A Complete Optimization Guide

> Optimize MTPLX for your hardware. Learn how to auto-tune draft depth by detecting GPU capabilities and injecting runtime configurations for peak performance. Get the complete guide.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: optimization-guide
- Published: 2026-09-02

---

**MTPLX automatically determines the optimal draft depth for your GPU by detecting hardware capabilities, consulting benchmark tables, and injecting the derived value into the runtime configuration when `MTPLX_LONG_CONTEXT_MTP_DEPTH_POLICY` is set to `auto` and no explicit block size is provided.**

MTPLX is an open-source inference engine that accelerates large language model generation through speculative decoding. Learning how to auto-tune draft depth in MTPLX ensures you maximize token throughput without manual configuration across different hardware from consumer GPUs to data-center accelerators.

## How Auto-Tuning Works in MTPLX

The **draft depth** (also called *draft block size*) controls how many speculative tokens MTPLX generates per inference block. When left to automatic selection, MTPLX executes a three-stage hardware-aware initialization sequence at server startup.

### Hardware Capability Detection

First, MTPLX inspects GPU memory and model-specific context limits. In [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (lines 33930-33935), the backend reads the device-specific capability (`gemma_cap`) for Gemma-4 models and clamps the candidate draft depth to ensure it does not exceed hardware limits. This prevents out-of-memory errors when deploying on constrained devices.

### Benchmark-Driven Defaults

If the user has not specified a depth, MTPLX loads pre-computed performance benchmarks. The file [`mtplx/benchmarks/mtp_depth_sweep.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/mtp_depth_sweep.py) contains mappings between hardware classes and optimal block sizes. According to the source code in [`mtplx/backends/gemma4_assistant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/gemma4_assistant.py) (lines 484-497), the system falls back to `DEFAULT_DRAFT_BLOCK_SIZE` only when no benchmark entry matches the detected hardware.

### Runtime Configuration Injection

Finally, the derived value is written to `runtime.config.draft_block_size`. The UI metadata is simultaneously updated in [`mtplx/backends/descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py) (lines 986-1000), where the `draft_control` field's `minimum`, `maximum`, `default`, and `value_labels` properties are adjusted to reflect the auto-tuned range. This ensures client interfaces display valid options for the specific hardware.

## Enabling and Configuring Auto-Tune Draft Depth

The auto-tuning behavior is controlled through environment variables and CLI arguments defined in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) (line 1318).

### Activate via Environment Variable

Set `MTPLX_LONG_CONTEXT_MTP_DEPTH_POLICY` to `auto` before starting the server:

```bash
export MTPLX_LONG_CONTEXT_MTP_DEPTH_POLICY=auto
mtplx server start

```

The server logs the selected block size during initialization (for example, `draft_block_size=4` for an Nvidia A100 40GB).

### Override via CLI

To disable auto-tuning and force a specific depth, pass the `--draft-block-size` argument:

```bash
mtplx server start --draft-block_size 2

```

This overrides the automatic selection and uses the specified value regardless of hardware capabilities or benchmark tables.

### Access Through REST API and Python SDK

Once running, the auto-tuned configuration is exposed through the `draft_control` endpoint. Query the current settings to verify the hardware-derived limits:

```python
import requests

resp = requests.get("http://localhost:8000/v1/model-controls")
print(resp.json()["draft_control"]["value_labels"])

# Output: ["D1", "D2", "D3", "D4"] indicating the auto-chosen maximum

```

Using the Python SDK provides structured access to the same metadata:

```python
from mtplx import MTPLXClient

client = MTPLXClient()
info = client.model_controls()
print(info.draft_control)

# Output: DraftControl(maximum=4, default=1, ...)

```

## Verifying Auto-Tuned Depth in Inference

To confirm the active draft depth during actual inference, inspect the usage statistics returned by the API:

```python
response = client.chat_completion(
    model="gemma4",
    messages=[{"role": "user", "content": "Hello"}],
    draft_block_size=None  # Null value triggers auto-tune

)
print(response["usage"]["drafted_by_depth"])

# Output: [8, 8, 8] showing each block used the auto-tuned depth

```

When `draft_block_size` is explicitly set to `None` or omitted, MTPLX applies the hardware-specific optimum calculated during startup.

## Key Implementation Files

Understanding the source structure helps with debugging and custom deployments:

- **[`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py)** (lines 33930-33935): Clamps candidate depths to device-specific caps (`gemma_cap`) and injects values into runtime arguments.
- **[`mtplx/backends/gemma4_assistant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/gemma4_assistant.py)** (lines 484-497): Defines `DEFAULT_DRAFT_BLOCK_SIZE` and validates hardware-derived values against supported ranges.
- **[`mtplx/runtime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime.py)** (lines 601-610): Applies the benchmark-derived best block size to the active runtime configuration.
- **[`mtplx/backends/descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py)** (lines 986-1000): Constructs the `draft_control` UI metadata that reflects auto-tuned ranges in client interfaces.
- **[`mtplx/benchmarks/mtp_depth_sweep.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/mtp_depth_sweep.py)**: Stores the benchmark table mapping hardware classes to optimal draft depths.
- **[`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py)** (line 1318): Exposes the `--draft-block-size` CLI argument and related configuration options.

## Summary

- **Auto-tuning triggers** when `MTPLX_LONG_CONTEXT_MTP_DEPTH_POLICY=auto` is set and no `--draft-block-size` argument is passed.
- **Three-stage process**: Hardware capability detection ([`openai.py`](https://github.com/youssofal/MTPLX/blob/main/openai.py)), benchmark lookup ([`gemma4_assistant.py`](https://github.com/youssofal/MTPLX/blob/main/gemma4_assistant.py)), and runtime injection ([`descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/descriptors.py), [`runtime.py`](https://github.com/youssofal/MTPLX/blob/main/runtime.py)).
- **Override capability**: Explicit CLI arguments or API parameters disable auto-tuning and use the specified value instead.
- **Verification**: Check server logs at startup or inspect `usage.drafted_by_depth` in API responses to confirm the active depth.

## Frequently Asked Questions

### What hardware does MTPLX draft depth auto-tuning support?

MTPLX auto-detection primarily targets CUDA-capable GPUs with varying memory tiers. The system reads device-specific caps such as `gemma_cap` in the Gemma-4 backend to determine safe maximums, and consults benchmark tables in [`mtp_depth_sweep.py`](https://github.com/youssofal/MTPLX/blob/main/mtp_depth_sweep.py) that cover consumer cards through data-center accelerators like the A100.

### Can I use auto-tuning for some requests but override for others?

Once the server starts with auto-tuning enabled, the derived `draft_block_size` becomes the runtime default. However, individual API requests can override this by passing the `draft_block_size` parameter in the request body. Setting this to `null` or omitting it uses the auto-tuned value, while providing an integer forces that specific depth for that inference call.

### Where does MTPLX store the benchmark values for draft depth?

The benchmark mappings reside in [`mtplx/benchmarks/mtp_depth_sweep.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/mtp_depth_sweep.py). This file contains pre-computed sweeps that map hardware classes to optimal block sizes. If your specific GPU is not present in the benchmark table, MTPLX falls back to the `DEFAULT_DRAFT_BLOCK_SIZE` defined in [`mtplx/backends/gemma4_assistant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/gemma4_assistant.py).

### How do I debug which draft depth was auto-selected?

Check the server startup logs for entries like `draft_block_size=4`. Programmatically, query the `/v1/model-controls` endpoint and inspect the `draft_control` field, or check the `drafted_by_depth` array in completion response usage statistics to see the actual depth applied to each generation block.