How to Auto-Tune Draft Depth in MTPLX for Specific Hardware: A Complete Optimization Guide
MTPLX automatically determines the optimal draft depth for your GPU by detecting hardware capabilities, consulting benchmark tables, and injecting the derived value into the runtime configuration when MTPLX_LONG_CONTEXT_MTP_DEPTH_POLICY is set to auto and no explicit block size is provided.
MTPLX is an open-source inference engine that accelerates large language model generation through speculative decoding. Learning how to auto-tune draft depth in MTPLX ensures you maximize token throughput without manual configuration across different hardware from consumer GPUs to data-center accelerators.
How Auto-Tuning Works in MTPLX
The draft depth (also called draft block size) controls how many speculative tokens MTPLX generates per inference block. When left to automatic selection, MTPLX executes a three-stage hardware-aware initialization sequence at server startup.
Hardware Capability Detection
First, MTPLX inspects GPU memory and model-specific context limits. In mtplx/server/openai.py (lines 33930-33935), the backend reads the device-specific capability (gemma_cap) for Gemma-4 models and clamps the candidate draft depth to ensure it does not exceed hardware limits. This prevents out-of-memory errors when deploying on constrained devices.
Benchmark-Driven Defaults
If the user has not specified a depth, MTPLX loads pre-computed performance benchmarks. The file mtplx/benchmarks/mtp_depth_sweep.py contains mappings between hardware classes and optimal block sizes. According to the source code in mtplx/backends/gemma4_assistant.py (lines 484-497), the system falls back to DEFAULT_DRAFT_BLOCK_SIZE only when no benchmark entry matches the detected hardware.
Runtime Configuration Injection
Finally, the derived value is written to runtime.config.draft_block_size. The UI metadata is simultaneously updated in mtplx/backends/descriptors.py (lines 986-1000), where the draft_control field's minimum, maximum, default, and value_labels properties are adjusted to reflect the auto-tuned range. This ensures client interfaces display valid options for the specific hardware.
Enabling and Configuring Auto-Tune Draft Depth
The auto-tuning behavior is controlled through environment variables and CLI arguments defined in mtplx/commands/public.py (line 1318).
Activate via Environment Variable
Set MTPLX_LONG_CONTEXT_MTP_DEPTH_POLICY to auto before starting the server:
export MTPLX_LONG_CONTEXT_MTP_DEPTH_POLICY=auto
mtplx server start
The server logs the selected block size during initialization (for example, draft_block_size=4 for an Nvidia A100 40GB).
Override via CLI
To disable auto-tuning and force a specific depth, pass the --draft-block-size argument:
mtplx server start --draft-block_size 2
This overrides the automatic selection and uses the specified value regardless of hardware capabilities or benchmark tables.
Access Through REST API and Python SDK
Once running, the auto-tuned configuration is exposed through the draft_control endpoint. Query the current settings to verify the hardware-derived limits:
import requests
resp = requests.get("http://localhost:8000/v1/model-controls")
print(resp.json()["draft_control"]["value_labels"])
# Output: ["D1", "D2", "D3", "D4"] indicating the auto-chosen maximum
Using the Python SDK provides structured access to the same metadata:
from mtplx import MTPLXClient
client = MTPLXClient()
info = client.model_controls()
print(info.draft_control)
# Output: DraftControl(maximum=4, default=1, ...)
Verifying Auto-Tuned Depth in Inference
To confirm the active draft depth during actual inference, inspect the usage statistics returned by the API:
response = client.chat_completion(
model="gemma4",
messages=[{"role": "user", "content": "Hello"}],
draft_block_size=None # Null value triggers auto-tune
)
print(response["usage"]["drafted_by_depth"])
# Output: [8, 8, 8] showing each block used the auto-tuned depth
When draft_block_size is explicitly set to None or omitted, MTPLX applies the hardware-specific optimum calculated during startup.
Key Implementation Files
Understanding the source structure helps with debugging and custom deployments:
mtplx/server/openai.py(lines 33930-33935): Clamps candidate depths to device-specific caps (gemma_cap) and injects values into runtime arguments.mtplx/backends/gemma4_assistant.py(lines 484-497): DefinesDEFAULT_DRAFT_BLOCK_SIZEand validates hardware-derived values against supported ranges.mtplx/runtime.py(lines 601-610): Applies the benchmark-derived best block size to the active runtime configuration.mtplx/backends/descriptors.py(lines 986-1000): Constructs thedraft_controlUI metadata that reflects auto-tuned ranges in client interfaces.mtplx/benchmarks/mtp_depth_sweep.py: Stores the benchmark table mapping hardware classes to optimal draft depths.mtplx/commands/public.py(line 1318): Exposes the--draft-block-sizeCLI argument and related configuration options.
Summary
- Auto-tuning triggers when
MTPLX_LONG_CONTEXT_MTP_DEPTH_POLICY=autois set and no--draft-block-sizeargument is passed. - Three-stage process: Hardware capability detection (
openai.py), benchmark lookup (gemma4_assistant.py), and runtime injection (descriptors.py,runtime.py). - Override capability: Explicit CLI arguments or API parameters disable auto-tuning and use the specified value instead.
- Verification: Check server logs at startup or inspect
usage.drafted_by_depthin API responses to confirm the active depth.
Frequently Asked Questions
What hardware does MTPLX draft depth auto-tuning support?
MTPLX auto-detection primarily targets CUDA-capable GPUs with varying memory tiers. The system reads device-specific caps such as gemma_cap in the Gemma-4 backend to determine safe maximums, and consults benchmark tables in mtp_depth_sweep.py that cover consumer cards through data-center accelerators like the A100.
Can I use auto-tuning for some requests but override for others?
Once the server starts with auto-tuning enabled, the derived draft_block_size becomes the runtime default. However, individual API requests can override this by passing the draft_block_size parameter in the request body. Setting this to null or omitting it uses the auto-tuned value, while providing an integer forces that specific depth for that inference call.
Where does MTPLX store the benchmark values for draft depth?
The benchmark mappings reside in mtplx/benchmarks/mtp_depth_sweep.py. This file contains pre-computed sweeps that map hardware classes to optimal block sizes. If your specific GPU is not present in the benchmark table, MTPLX falls back to the DEFAULT_DRAFT_BLOCK_SIZE defined in mtplx/backends/gemma4_assistant.py.
How do I debug which draft depth was auto-selected?
Check the server startup logs for entries like draft_block_size=4. Programmatically, query the /v1/model-controls endpoint and inspect the draft_control field, or check the drafted_by_depth array in completion response usage statistics to see the actual depth applied to each generation block.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →