How to Tune MTPLX for Optimal Draft Depth on Your Machine
Set your draft depth using the --draft-block-size CLI flag for static tuning or enable adaptive optimization with --adaptive-policy expected_value to let the engine automatically adjust speculation based on real-time acceptance rates.
MTPLX is an open-source speculative decoding engine that accelerates LLM inference by generating draft tokens ahead of the target model. The draft depth—controlled via block size parameters—determines how many tokens are speculated per verification cycle, directly impacting throughput and latency on your specific hardware.
Understanding Draft Depth Semantics
Draft depth represents the number of tokens the draft model generates before the target model verifies them. In mtplx/backends/descriptors.py, the draft_semantics_for_model function (around line 568) defines model-specific constraints including default, minimum, and maximum depths. For example, some model descriptors cap the maximum depth at 3 to prevent excessive speculation waste, as implemented in the descriptor logic source.
The server clamps incoming depth requests to these bounds in mtplx/server/openai.py (around line 35857), ensuring you cannot exceed hardware-safe limits source.
Setting Static Draft Depth
Via CLI Flags
The most direct method to tune MTPLX is passing the appropriate flag when starting the server. In mtplx/commands/public.py (around line 1347), the CLI maps --draft-block-size (for Gemma-4 models) and --depth (for other architectures) to the backend request fields source.
# For Gemma-4 based assistants
mtplx serve --model gemma4-assistant --draft-block-size 4
# For Qwen and other models
mtplx serve --model qwen3.5b --depth 3
Via Environment Variables
For containerized deployments or shell scripts, override CLI settings using MTPLX_DRAFT_BLOCK_SIZE. This variable is read during runtime initialization in mtplx/runtime.py (around line 573), taking precedence over configuration files source.
export MTPLX_DRAFT_BLOCK_SIZE=4
export MTPLX_ADAPTIVE_VERIFY_COST_FEEDBACK=0 # Disable adaptive feedback for fixed depth
mtplx serve --model gemma4-assistant
Enabling Adaptive Draft Depth (Recommended)
Rather than fixing depth manually, enable the adaptive policy to let MTPLX optimize dynamically. The expected_value policy observes break-even acceptance rates and verification costs, adjusting depth per-cycle to maximize throughput.
mtplx serve --model qwen3.5b \
--adaptive-policy expected_value \
--adaptive-ev-base-depth 2 \
--adaptive-ev-warmup-full-depth-cycles 5 \
--adaptive-ev-exploration-interval 17
This policy calculates the economic viability of speculation using the cost model defined in the test suite at tests/test_trace_diagnostics.py (around line 45), ensuring depth selection balances compute waste against verification overhead source.
According to the v2.11.1 release notes (line 84), this approach yields optimal results for Qwen-3.8, automatically settling on depth 3 after warmup cycles source.
Runtime Tuning via API
You can adjust draft depth without restarting the server by patching the settings endpoint:
curl -X PATCH http://localhost:8031/v1/mtplx/settings \
-H "Content-Type: application/json" \
-d '{"draft_block_size": 2}'
The server confirms the applied value in the response, as verified in tests/test_server_openai.py (around line 2104) source.
Benchmarking to Find Your Optimal Depth
To empirically determine the best depth for your GPU:
- Run the built-in depth sweep benchmark:
mtplx bench depth-grid --model qwen3.5b
-
Identify the inflection point where throughput gains drop below 5% per additional draft token.
-
Use that depth as your
--draft-block-sizevalue or let the adaptive policy converge to it automatically.
The benchmark runner evaluates acceptance rates and verification latency to recommend hardware-specific settings.
Summary
- Static tuning uses
--draft-block-sizeorMTPLX_DRAFT_BLOCK_SIZEto fix depth at startup, suitable for consistent workloads. - Adaptive tuning via
--adaptive-policy expected_valuedynamically optimizes depth based on real-time acceptance, ideal for variable prompt distributions. - Model constraints defined in
descriptors.pyautomatically clamp depth to safe maximums. - Runtime adjustment is possible through the
/v1/mtplx/settingsREST API without service restart. - Benchmarking with
mtplx bench depth-gridreveals the throughput-optimal depth for your specific GPU memory bandwidth and compute.
Frequently Asked Questions
What is the default draft depth in MTPLX?
The default is pulled from the model descriptor's draft_semantics.default field in mtplx/backends/descriptors.py. For most models, this defaults to 2 or 3 tokens depending on the architecture's speculation efficiency. You can override it via CLI, environment variables, or the settings API.
How do I know if my draft depth is too high?
If the acceptance rate drops below the break-even threshold (typically 40-50% for most models), excessive draft tokens are being rejected, wasting compute. Check the live dashboard introduced in v2.11.1 (line 84) to view acceptance curves by depth; flat or declining throughput at higher depths indicates you've exceeded the optimal range source.
Can I change draft depth without restarting the server?
Yes. Send a PATCH request to /v1/mtplx/settings with the draft_block_size field. The change takes effect immediately for subsequent inference requests, as confirmed by the server response echoing the applied value source.
What is the maximum draft depth supported?
Maximum depth is model-specific and enforced by the descriptor's draft_semantics.maximum value. For example, Qwen-3.8 caps at 3 tokens (line 734 in descriptors.py). Attempting to set a higher value results in automatic clamping to this maximum at the server level source.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →