How to Retune the Draft Depth for a Model in MTPLX: 3 Configuration Methods

You can retune the draft depth in MTPLX using the --draft_block_size CLI flag, the MTPLX_DRAFT_BLOCK_SIZE environment variable, or by posting to the /_mtplx/settings HTTP endpoint at runtime.

MTPLX accelerates inference through speculative decoding, generating "draft" tokens before finalizing outputs. The number of tokens generated per speculative step—referred to as draft depth or draft block size—is a critical tuning parameter that balances latency against computational overhead. This guide explains how to adjust this setting in the youssofal/MTPLX repository using three distinct configuration mechanisms.

Understanding Draft Depth in MTPLX

During speculative decoding, MTPLX generates multiple candidate tokens in a single forward pass. The draft_block_size parameter controls how many tokens the model drafts before verification. Internally, this value populates runtime.config.draft_block_size and interfaces with backend.draft_semantics to manage the speculative generation loop, as defined in mtplx/server/openai.py.

Method 1: Configure via CLI Flag

The most direct approach is using the --draft_block_size argument (short form -d) when launching the server. According to the argument parser in mtplx/server/openai.py (line 34019), this flag directly updates args.draft_block_size, which then propagates to the runtime configuration.


# Start server with draft depth of 4 tokens per block

mtplx serve --model gemma4 --draft_block_size 4

# Or using the short flag

mtplx serve --model gemma4 -d 4

Method 2: Set via Environment Variable

For Docker containers or automated deployments, set the MTPLX_DRAFT_BLOCK_SIZE environment variable. As implemented in mtplx/app_settings.py (line 45), this variable is read during application startup and overrides the default when no CLI flag is supplied.

export MTPLX_DRAFT_BLOCK_SIZE=3
mtplx serve --model gemma4

Method 3: Adjust at Runtime via HTTP Endpoint

You can modify the draft depth without restarting the server by posting to the internal settings endpoint. The request handler in mtplx/server/openai.py (line 1994) processes these updates by modifying state.args.draft_block_size and synchronizing the active runtime.config.draft_block_size immediately.

import requests

response = requests.post(
    "http://localhost:8080/_mtplx/settings",
    json={"draft_block_size": 5}
)
response.raise_for_status()
print(f"Draft depth updated to {response.json()['draft_block_size']}")

Model-Specific Constraints and Validation

Certain model architectures enforce minimum draft depths. For Gemma-4 variants, mtplx/backends/gemma4_assistant.py (line 512) validates the configuration and rejects values below the threshold:

if self.draft_block_size < 2:
    raise ValueError("Gemma 4 assistant draft_block_size must be >= 2")

Attempting to set a non-compliant value raises a ValueError during either initialization or a runtime update.

Verifying the Active Draft Depth

To confirm your configuration is active, inspect the runtime telemetry or status endpoints. The test suite in mtplx/tests/test_server_openai.py (around line 1178) demonstrates that the server reports draft activity in telemetry envelopes under the key drafted_by_depth.

You can also query the current configuration directly:

import requests

status = requests.get("http://localhost:8080/_mtplx/status").json()
current_depth = status["runtime"]["config"]["draft_block_size"]
print(f"Current draft depth: {current_depth}")

Summary

  • Retune the draft depth using CLI flags (--draft_block_size or -d), environment variables (MTPLX_DRAFT_BLOCK_SIZE), or the HTTP settings endpoint (/_mtplx/settings).
  • The parameter maps to runtime.config.draft_block_size and controls speculative decoding batch size.
  • Gemma-4 models require a minimum draft depth of 2, enforced in mtplx/backends/gemma4_assistant.py.
  • Verify active settings through the /_mtplx/status endpoint or telemetry fields like drafted_by_depth.

Frequently Asked Questions

What is the default draft depth in MTPLX?

The default value is model-dependent and typically defined in mtplx/app_settings.py or by the specific backend implementation. You can determine the current default by querying the /_mtplx/status endpoint after startup or reviewing the server logs, which display the resolved runtime.config.draft_block_size during initialization.

Can I change the draft depth without restarting the MTPLX server?

Yes. Send a POST request to /_mtplx/settings with a JSON payload containing {"draft_block_size": N}. This updates state.args.draft_block_size and the active configuration immediately without requiring a restart, as implemented in mtplx/server/openai.py (line 1994).

Why does MTPLX raise a ValueError when I set draft_block_size to 1 for Gemma-4?

Gemma-4 models require a minimum draft depth of 2 tokens due to architectural constraints in their speculative decoding implementation. The validation occurs in mtplx/backends/gemma4_assistant.py (line 512), which explicitly raises ValueError("Gemma 4 assistant draft_block_size must be >= 2") if the supplied value is insufficient.

How does draft depth affect inference performance?

Higher draft depths increase the number of speculative tokens generated per step, potentially improving throughput when draft acceptance rates are high, but they consume additional memory and compute. Lower values reduce resource pressure but may diminish latency benefits. Tune this parameter based on your hardware constraints and the model's observed draft acceptance rate.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →