# How to Retune the Draft Depth for a Model in MTPLX: 3 Configuration Methods

> Learn three ways to retune draft depth in MTPLX: CLI flag, environment variable, or runtime HTTP POST. Optimize your model performance with these configuration methods.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-02

---

**You can retune the draft depth in MTPLX using the `--draft_block_size` CLI flag, the `MTPLX_DRAFT_BLOCK_SIZE` environment variable, or by posting to the `/_mtplx/settings` HTTP endpoint at runtime.**

MTPLX accelerates inference through speculative decoding, generating "draft" tokens before finalizing outputs. The number of tokens generated per speculative step—referred to as **draft depth** or **draft block size**—is a critical tuning parameter that balances latency against computational overhead. This guide explains how to adjust this setting in the `youssofal/MTPLX` repository using three distinct configuration mechanisms.

## Understanding Draft Depth in MTPLX

During speculative decoding, MTPLX generates multiple candidate tokens in a single forward pass. The `draft_block_size` parameter controls how many tokens the model drafts before verification. Internally, this value populates `runtime.config.draft_block_size` and interfaces with `backend.draft_semantics` to manage the speculative generation loop, as defined in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py).

## Method 1: Configure via CLI Flag

The most direct approach is using the `--draft_block_size` argument (short form `-d`) when launching the server. According to the argument parser in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (line 34019), this flag directly updates `args.draft_block_size`, which then propagates to the runtime configuration.

```bash

# Start server with draft depth of 4 tokens per block

mtplx serve --model gemma4 --draft_block_size 4

# Or using the short flag

mtplx serve --model gemma4 -d 4

```

## Method 2: Set via Environment Variable

For Docker containers or automated deployments, set the `MTPLX_DRAFT_BLOCK_SIZE` environment variable. As implemented in [`mtplx/app_settings.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/app_settings.py) (line 45), this variable is read during application startup and overrides the default when no CLI flag is supplied.

```bash
export MTPLX_DRAFT_BLOCK_SIZE=3
mtplx serve --model gemma4

```

## Method 3: Adjust at Runtime via HTTP Endpoint

You can modify the draft depth without restarting the server by posting to the internal settings endpoint. The request handler in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (line 1994) processes these updates by modifying `state.args.draft_block_size` and synchronizing the active `runtime.config.draft_block_size` immediately.

```python
import requests

response = requests.post(
    "http://localhost:8080/_mtplx/settings",
    json={"draft_block_size": 5}
)
response.raise_for_status()
print(f"Draft depth updated to {response.json()['draft_block_size']}")

```

## Model-Specific Constraints and Validation

Certain model architectures enforce minimum draft depths. For Gemma-4 variants, [`mtplx/backends/gemma4_assistant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/gemma4_assistant.py) (line 512) validates the configuration and rejects values below the threshold:

```python
if self.draft_block_size < 2:
    raise ValueError("Gemma 4 assistant draft_block_size must be >= 2")

```

Attempting to set a non-compliant value raises a `ValueError` during either initialization or a runtime update.

## Verifying the Active Draft Depth

To confirm your configuration is active, inspect the runtime telemetry or status endpoints. The test suite in [`mtplx/tests/test_server_openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/tests/test_server_openai.py) (around line 1178) demonstrates that the server reports draft activity in telemetry envelopes under the key `drafted_by_depth`.

You can also query the current configuration directly:

```python
import requests

status = requests.get("http://localhost:8080/_mtplx/status").json()
current_depth = status["runtime"]["config"]["draft_block_size"]
print(f"Current draft depth: {current_depth}")

```

## Summary

- **Retune the draft depth** using CLI flags (`--draft_block_size` or `-d`), environment variables (`MTPLX_DRAFT_BLOCK_SIZE`), or the HTTP settings endpoint (`/_mtplx/settings`).
- The parameter maps to `runtime.config.draft_block_size` and controls speculative decoding batch size.
- **Gemma-4 models require a minimum draft depth of 2**, enforced in [`mtplx/backends/gemma4_assistant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/gemma4_assistant.py).
- Verify active settings through the `/_mtplx/status` endpoint or telemetry fields like `drafted_by_depth`.

## Frequently Asked Questions

### What is the default draft depth in MTPLX?

The default value is model-dependent and typically defined in [`mtplx/app_settings.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/app_settings.py) or by the specific backend implementation. You can determine the current default by querying the `/_mtplx/status` endpoint after startup or reviewing the server logs, which display the resolved `runtime.config.draft_block_size` during initialization.

### Can I change the draft depth without restarting the MTPLX server?

Yes. Send a POST request to `/_mtplx/settings` with a JSON payload containing `{"draft_block_size": N}`. This updates `state.args.draft_block_size` and the active configuration immediately without requiring a restart, as implemented in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (line 1994).

### Why does MTPLX raise a ValueError when I set draft_block_size to 1 for Gemma-4?

Gemma-4 models require a minimum draft depth of 2 tokens due to architectural constraints in their speculative decoding implementation. The validation occurs in [`mtplx/backends/gemma4_assistant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/gemma4_assistant.py) (line 512), which explicitly raises `ValueError("Gemma 4 assistant draft_block_size must be >= 2")` if the supplied value is insufficient.

### How does draft depth affect inference performance?

Higher draft depths increase the number of speculative tokens generated per step, potentially improving throughput when draft acceptance rates are high, but they consume additional memory and compute. Lower values reduce resource pressure but may diminish latency benefits. Tune this parameter based on your hardware constraints and the model's observed draft acceptance rate.