MTPLX Release Notes: v2.9.2 Multi-Token Prediction Speedups and Vision Stability
MTPLX v2.9.2 introduces greedy decode speedups of up to 9.8% on Apple Silicon, stabilizes vision model handling, and adds experimental kernels for power users while maintaining exact rejection sampling compliance.
MTPLX is a native macOS runtime and CLI for executing large language models with multi-token prediction (MTP) on Apple Silicon. The latest v2.9.2 release, documented in docs/releases/v2.9.2.md, focuses on performance optimizations for greedy decoding, refined agent transcript handling, and critical fixes for vision model stability.
What's New in MTPLX v2.9.2
Greedy Decode Speedup for Temperature-Zero Inference
The headline improvement in MTPLX v2.9.2 targets chained greedy drafting, now enabled by default for temperature=0 requests under 12,000 tokens. According to the release documentation, this optimization delivers +2.5% to +9.8% speed gains on M5 Max hardware by leveraging deterministic token acceptance patterns. The implementation modifies the drafting strategy in the MTP engine to commit multiple tokens without full verification overhead when the temperature is zero, while maintaining mathematical correctness through the Leviathan & Chen rejection sampling theorem.
Agent Transcript Passthrough Controls
MTPLX v2.9.2 disables automatic transcript rewriting for agent workflows, giving developers explicit control over conversation history manipulation. Set the MTPLX_AGENT_REWRITES environment variable to enable specific rewrite strategies when needed. This change affects how mtplx/serve handles multi-turn contexts, preventing unwanted transformation of tool call sequences or system prompts.
Vision Model Stability Fixes
Images now survive canonicalization correctly across the qwen3_vl_tower.py pipeline, and vision rows persist through warm-cache restores. The mtplx/vision_graft.py module received targeted fixes to ensure that M-RoPE (multi-modal rotary position embeddings) calculations maintain image-aware token positions after session restoration. These corrections prevent embedding drift during long-running vision-language conversations.
Experimental Kernel Flags
Power users can now test two advanced optimizations via environment variables:
MTPLX_FUSE_PROJ: Enables projection fusion kernels for reduced memory bandwidthMTPLX_VK_CROSSROW: Activates cross-row verification patterns in the KV cache handling withinmtplx/kv_quant.py
Both flags default to off in v2.9.2 but provide additional throughput for quantized 27B and 9B models when enabled.
Forge and Quantization Correctness
The mtplx forge build pipeline now properly honors quantize: false overrides and fixes norm-convention handling during Hugging Face checkpoint conversion. When running mtplx forge build, the verification stage (mtplx forge verify) correctly validates that non-quantized models bypass the NAX kernel quantization paths.
Core Architecture Supporting v2.9.2
Multi-Token Prediction Engine
MTPLX implements speculative decoding using the model's own draft heads to generate several tokens ahead, then verifies each draft block in a single batched forward pass. This architecture, defined across mtplx/kv_quant.py and mtplx/vision_graft.py, enables up to 2× faster decoding compared to standard autoregressive generation on M-series Macs.
Turbo vs Sustained Operating Modes
The v2.9.2 release maintains support for dual execution modes controlled via MTPLX_MODE:
- Turbo: Uses compiled NAX verify kernels for quantized models, ideal for short-context interactions
- Sustained: Employs chunked pre-fill and request-sized KV caching for long-context sessions
Upgrading to MTPLX v2.9.2
Update via Homebrew to access the latest release:
brew upgrade youssofal/mtplx/mtplx
Verify the installation and check available experimental features:
mtplx --version
MTPLX_FUSE_PROJ=1 mtplx serve --port 8000
For vision workloads, ensure your client requests include proper base64 encoding to trigger the qwen3_vl_tower.py pipeline:
curl http://127.0.0.1:8000/v1/messages \
-H 'Content-Type: application/json' \
-d '{
"model":"mtplx",
"messages":[{"role":"user","content":[
{"type":"text","text":"Analyze this image"},
{"type":"image","source":{"type":"base64","media_type":"image/png","data":"..."}}
]}]
}'
Summary
- MTPLX v2.9.2 delivers up to 9.8% faster greedy decoding on Apple Silicon through chained drafting optimizations
- Agent transcript handling now requires explicit opt-in via
MTPLX_AGENT_REWRITESfor conversation rewrites - Vision models using
mtplx/vision/qwen3_vl_tower.pyretain image embeddings correctly across session restores - Experimental kernels
MTPLX_FUSE_PROJandMTPLX_VK_CROSSROWoffer additional performance headroom for advanced users - The
forgebuild system correctly processesquantize: falseoverrides and norm conventions
Frequently Asked Questions
What is multi-token prediction in MTPLX?
Multi-token prediction (MTP) is a speculative decoding technique where MTPLX uses the model's draft heads to predict several future tokens simultaneously, then verifies them in a single batched forward pass. According to the source code in mtplx/kv_quant.py, this process uses exact rejection sampling with residual correction to maintain sampling fidelity while achieving up to 2× speedup over standard autoregressive decoding on Apple Silicon.
How do I enable experimental kernels in MTPLX v2.9.2?
Set the environment variables MTPLX_FUSE_PROJ=1 or MTPLX_VK_CROSSROW=1 before launching the server. These flags, documented in docs/releases/v2.9.2.md, activate projection fusion and cross-row verification optimizations respectively. They are disabled by default and primarily benefit quantized 27B and 9B model deployments on high-memory Macs.
Does MTPLX v2.9.2 support vision-language models?
Yes. MTPLX v2.9.2 includes stability fixes for vision pipelines, specifically for Qwen-3-Next models using M-RoPE positioning. The mtplx/vision/qwen3_vl_tower.py module handles image embedding and sparse attention, while mtplx/vision_graft.py integrates vision towers into the MTP pipeline. Images now correctly survive canonicalization and warm-cache restores.
What is the difference between Turbo and Sustained modes in MTPLX?
Turbo mode uses compiled NAX verify kernels optimized for quantized models and short contexts, providing maximum throughput. Sustained mode employs chunked pre-fill and request-sized KV caching within mtplx/kv_quant.py for long-context sessions. Set MTPLX_MODE to select between them, or allow the runtime to auto-select based on context length heuristics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →