# MTPLX Release Notes: v2.9.2 Multi-Token Prediction Speedups and Vision Stability

> Explore MTPLX v2.9.2 release notes! Discover multi-token prediction speedups up to 9.8% on Apple Silicon, vision model stability improvements, and experimental kernels.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: release-notes
- Published: 2026-09-11

---

**MTPLX v2.9.2 introduces greedy decode speedups of up to 9.8% on Apple Silicon, stabilizes vision model handling, and adds experimental kernels for power users while maintaining exact rejection sampling compliance.**

**MTPLX** is a native macOS runtime and CLI for executing large language models with **multi-token prediction (MTP)** on Apple Silicon. The latest v2.9.2 release, documented in [`docs/releases/v2.9.2.md`](https://github.com/youssofal/MTPLX/blob/main/docs/releases/v2.9.2.md), focuses on performance optimizations for greedy decoding, refined agent transcript handling, and critical fixes for vision model stability.

## What's New in MTPLX v2.9.2

### Greedy Decode Speedup for Temperature-Zero Inference

The headline improvement in MTPLX v2.9.2 targets **chained greedy drafting**, now enabled by default for `temperature=0` requests under 12,000 tokens. According to the release documentation, this optimization delivers **+2.5% to +9.8%** speed gains on M5 Max hardware by leveraging deterministic token acceptance patterns. The implementation modifies the drafting strategy in the MTP engine to commit multiple tokens without full verification overhead when the temperature is zero, while maintaining mathematical correctness through the Leviathan & Chen rejection sampling theorem.

### Agent Transcript Passthrough Controls

MTPLX v2.9.2 disables automatic transcript rewriting for agent workflows, giving developers explicit control over conversation history manipulation. Set the `MTPLX_AGENT_REWRITES` environment variable to enable specific rewrite strategies when needed. This change affects how `mtplx/serve` handles multi-turn contexts, preventing unwanted transformation of tool call sequences or system prompts.

### Vision Model Stability Fixes

Images now survive canonicalization correctly across the [`qwen3_vl_tower.py`](https://github.com/youssofal/MTPLX/blob/main/qwen3_vl_tower.py) pipeline, and vision rows persist through warm-cache restores. The [`mtplx/vision_graft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py) module received targeted fixes to ensure that M-RoPE (multi-modal rotary position embeddings) calculations maintain image-aware token positions after session restoration. These corrections prevent embedding drift during long-running vision-language conversations.

### Experimental Kernel Flags

Power users can now test two advanced optimizations via environment variables:
- **`MTPLX_FUSE_PROJ`**: Enables projection fusion kernels for reduced memory bandwidth
- **`MTPLX_VK_CROSSROW`**: Activates cross-row verification patterns in the KV cache handling within [`mtplx/kv_quant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/kv_quant.py)

Both flags default to off in v2.9.2 but provide additional throughput for quantized 27B and 9B models when enabled.

### Forge and Quantization Correctness

The `mtplx forge` build pipeline now properly honors `quantize: false` overrides and fixes norm-convention handling during Hugging Face checkpoint conversion. When running `mtplx forge build`, the verification stage (`mtplx forge verify`) correctly validates that non-quantized models bypass the NAX kernel quantization paths.

## Core Architecture Supporting v2.9.2

### Multi-Token Prediction Engine

MTPLX implements speculative decoding using the model's own draft heads to generate several tokens ahead, then verifies each draft block in a single batched forward pass. This architecture, defined across [`mtplx/kv_quant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/kv_quant.py) and [`mtplx/vision_graft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py), enables up to **2× faster decoding** compared to standard autoregressive generation on M-series Macs.

### Turbo vs Sustained Operating Modes

The v2.9.2 release maintains support for dual execution modes controlled via `MTPLX_MODE`:
- **Turbo**: Uses compiled NAX verify kernels for quantized models, ideal for short-context interactions
- **Sustained**: Employs chunked pre-fill and request-sized KV caching for long-context sessions

## Upgrading to MTPLX v2.9.2

Update via Homebrew to access the latest release:

```bash
brew upgrade youssofal/mtplx/mtplx

```

Verify the installation and check available experimental features:

```bash
mtplx --version
MTPLX_FUSE_PROJ=1 mtplx serve --port 8000

```

For vision workloads, ensure your client requests include proper base64 encoding to trigger the [`qwen3_vl_tower.py`](https://github.com/youssofal/MTPLX/blob/main/qwen3_vl_tower.py) pipeline:

```bash
curl http://127.0.0.1:8000/v1/messages \
  -H 'Content-Type: application/json' \
  -d '{
        "model":"mtplx",
        "messages":[{"role":"user","content":[
            {"type":"text","text":"Analyze this image"},
            {"type":"image","source":{"type":"base64","media_type":"image/png","data":"..."}}
        ]}]
      }'

```

## Summary

- **MTPLX v2.9.2** delivers up to 9.8% faster greedy decoding on Apple Silicon through chained drafting optimizations
- Agent transcript handling now requires explicit opt-in via `MTPLX_AGENT_REWRITES` for conversation rewrites
- Vision models using [`mtplx/vision/qwen3_vl_tower.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision/qwen3_vl_tower.py) retain image embeddings correctly across session restores
- Experimental kernels `MTPLX_FUSE_PROJ` and `MTPLX_VK_CROSSROW` offer additional performance headroom for advanced users
- The `forge` build system correctly processes `quantize: false` overrides and norm conventions

## Frequently Asked Questions

### What is multi-token prediction in MTPLX?

**Multi-token prediction (MTP)** is a speculative decoding technique where MTPLX uses the model's draft heads to predict several future tokens simultaneously, then verifies them in a single batched forward pass. According to the source code in [`mtplx/kv_quant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/kv_quant.py), this process uses exact rejection sampling with residual correction to maintain sampling fidelity while achieving up to 2× speedup over standard autoregressive decoding on Apple Silicon.

### How do I enable experimental kernels in MTPLX v2.9.2?

Set the environment variables `MTPLX_FUSE_PROJ=1` or `MTPLX_VK_CROSSROW=1` before launching the server. These flags, documented in [`docs/releases/v2.9.2.md`](https://github.com/youssofal/MTPLX/blob/main/docs/releases/v2.9.2.md), activate projection fusion and cross-row verification optimizations respectively. They are disabled by default and primarily benefit quantized 27B and 9B model deployments on high-memory Macs.

### Does MTPLX v2.9.2 support vision-language models?

Yes. MTPLX v2.9.2 includes stability fixes for vision pipelines, specifically for Qwen-3-Next models using M-RoPE positioning. The [`mtplx/vision/qwen3_vl_tower.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision/qwen3_vl_tower.py) module handles image embedding and sparse attention, while [`mtplx/vision_graft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py) integrates vision towers into the MTP pipeline. Images now correctly survive canonicalization and warm-cache restores.

### What is the difference between Turbo and Sustained modes in MTPLX?

**Turbo mode** uses compiled NAX verify kernels optimized for quantized models and short contexts, providing maximum throughput. **Sustained mode** employs chunked pre-fill and request-sized KV caching within [`mtplx/kv_quant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/kv_quant.py) for long-context sessions. Set `MTPLX_MODE` to select between them, or allow the runtime to auto-select based on context length heuristics.