# How to Tune MTPLX for Optimal Draft Depth on Your Machine

> Optimize MTPLX draft depth using --draft-block-size for static tuning or enable adaptive optimization with --adaptive-policy for automatic adjustment. Improve your machine's performance now.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-08

---

**Set your draft depth using the `--draft-block-size` CLI flag for static tuning or enable adaptive optimization with `--adaptive-policy expected_value` to let the engine automatically adjust speculation based on real-time acceptance rates.**

MTPLX is an open-source speculative decoding engine that accelerates LLM inference by generating draft tokens ahead of the target model. The **draft depth**—controlled via block size parameters—determines how many tokens are speculated per verification cycle, directly impacting throughput and latency on your specific hardware.

## Understanding Draft Depth Semantics

Draft depth represents the number of tokens the draft model generates before the target model verifies them. In [`mtplx/backends/descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py), the `draft_semantics_for_model` function (around line 568) defines model-specific constraints including default, minimum, and maximum depths. For example, some model descriptors cap the maximum depth at 3 to prevent excessive speculation waste, as implemented in the descriptor logic [source](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py#L568).

The server clamps incoming depth requests to these bounds in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (around line 35857), ensuring you cannot exceed hardware-safe limits [source](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py#L35857).

## Setting Static Draft Depth

### Via CLI Flags

The most direct method to tune MTPLX is passing the appropriate flag when starting the server. In [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) (around line 1347), the CLI maps `--draft-block-size` (for Gemma-4 models) and `--depth` (for other architectures) to the backend request fields [source](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py#L1347).

```bash

# For Gemma-4 based assistants

mtplx serve --model gemma4-assistant --draft-block-size 4

# For Qwen and other models

mtplx serve --model qwen3.5b --depth 3

```

### Via Environment Variables

For containerized deployments or shell scripts, override CLI settings using `MTPLX_DRAFT_BLOCK_SIZE`. This variable is read during runtime initialization in [`mtplx/runtime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime.py) (around line 573), taking precedence over configuration files [source](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime.py#L573).

```bash
export MTPLX_DRAFT_BLOCK_SIZE=4
export MTPLX_ADAPTIVE_VERIFY_COST_FEEDBACK=0  # Disable adaptive feedback for fixed depth

mtplx serve --model gemma4-assistant

```

## Enabling Adaptive Draft Depth (Recommended)

Rather than fixing depth manually, enable the adaptive policy to let MTPLX optimize dynamically. The `expected_value` policy observes **break-even acceptance** rates and verification costs, adjusting depth per-cycle to maximize throughput.

```bash
mtplx serve --model qwen3.5b \
  --adaptive-policy expected_value \
  --adaptive-ev-base-depth 2 \
  --adaptive-ev-warmup-full-depth-cycles 5 \
  --adaptive-ev-exploration-interval 17

```

This policy calculates the economic viability of speculation using the cost model defined in the test suite at [`tests/test_trace_diagnostics.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_trace_diagnostics.py) (around line 45), ensuring depth selection balances compute waste against verification overhead [source](https://github.com/youssofal/MTPLX/blob/main/tests/test_trace_diagnostics.py#L45).

According to the v2.11.1 release notes (line 84), this approach yields optimal results for Qwen-3.8, automatically settling on depth 3 after warmup cycles [source](https://github.com/youssofal/MTPLX/blob/main/docs/releases/v2.11.1.md#L84).

## Runtime Tuning via API

You can adjust draft depth without restarting the server by patching the settings endpoint:

```bash
curl -X PATCH http://localhost:8031/v1/mtplx/settings \
     -H "Content-Type: application/json" \
     -d '{"draft_block_size": 2}'

```

The server confirms the applied value in the response, as verified in [`tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py) (around line 2104) [source](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py#L2104).

## Benchmarking to Find Your Optimal Depth

To empirically determine the best depth for your GPU:

1. Run the built-in depth sweep benchmark:

```bash
mtplx bench depth-grid --model qwen3.5b

```

2. Identify the **inflection point** where throughput gains drop below 5% per additional draft token.

3. Use that depth as your `--draft-block-size` value or let the adaptive policy converge to it automatically.

The benchmark runner evaluates acceptance rates and verification latency to recommend hardware-specific settings.

## Summary

- **Static tuning** uses `--draft-block-size` or `MTPLX_DRAFT_BLOCK_SIZE` to fix depth at startup, suitable for consistent workloads.
- **Adaptive tuning** via `--adaptive-policy expected_value` dynamically optimizes depth based on real-time acceptance, ideal for variable prompt distributions.
- **Model constraints** defined in [`descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/descriptors.py) automatically clamp depth to safe maximums.
- **Runtime adjustment** is possible through the `/v1/mtplx/settings` REST API without service restart.
- **Benchmarking** with `mtplx bench depth-grid` reveals the throughput-optimal depth for your specific GPU memory bandwidth and compute.

## Frequently Asked Questions

### What is the default draft depth in MTPLX?

The default is pulled from the model descriptor's `draft_semantics.default` field in [`mtplx/backends/descriptors.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py). For most models, this defaults to 2 or 3 tokens depending on the architecture's speculation efficiency. You can override it via CLI, environment variables, or the settings API.

### How do I know if my draft depth is too high?

If the **acceptance rate** drops below the break-even threshold (typically 40-50% for most models), excessive draft tokens are being rejected, wasting compute. Check the live dashboard introduced in v2.11.1 (line 84) to view acceptance curves by depth; flat or declining throughput at higher depths indicates you've exceeded the optimal range [source](https://github.com/youssofal/MTPLX/blob/main/docs/releases/v2.11.1.md#L84).

### Can I change draft depth without restarting the server?

Yes. Send a PATCH request to `/v1/mtplx/settings` with the `draft_block_size` field. The change takes effect immediately for subsequent inference requests, as confirmed by the server response echoing the applied value [source](https://github.com/youssofal/MTPLX/blob/main/tests/test_server_openai.py#L2104).

### What is the maximum draft depth supported?

Maximum depth is model-specific and enforced by the descriptor's `draft_semantics.maximum` value. For example, Qwen-3.8 caps at 3 tokens (line 734 in descriptors.py). Attempting to set a higher value results in automatic clamping to this maximum at the server level [source](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/descriptors.py#L734).