MTPLX Documentation: Multi-Token Prediction Runtime for Apple Silicon

MTPLX is a native macOS runtime that accelerates local large language models on Apple Silicon by leveraging Multi-Token Prediction (MTP) heads to draft and verify multiple tokens in parallel, achieving 1.5-2× higher throughput while maintaining exact probability distributions.

This MTPLX documentation covers the architecture and implementation details of the native inference engine built on top of MLX. The system exploits MTP heads included in modern models to perform speculative decoding without requiring a separate draft model, instead using the target model's own parallel prediction capability to generate candidate token blocks and verify them in a single batched forward pass.

Core Architecture

The MTPLX codebase organizes functionality into specialized modules that handle speculative decoding, hardware-optimized kernels, request scheduling, and API serving.

Speculative Engine

The mtplx/speculative.py module implements the draft-verify-accept loop that drives multi-token drafting. This component generates blocks of N candidate tokens in one forward pass, then evaluates the entire block using a single batched verification pass. The engine performs exact rejection sampling with residual correction based on the Leviathan-Chen theorem, ensuring the final output matches the exact probability distribution of standard autoregressive decoding while accepting valid token blocks in bulk.

TurboQuant Kernels

For quantized model execution, mtplx/turboquant.py provides optimized verification kernels used in Turbo mode. These NAX-compiled verify kernels target 27B and 9B flagship models, delivering maximal token-per-second throughput by minimizing verification latency on Apple Silicon GPUs.

Batching and Scheduling

The mtplx/batching/scheduler.py module handles request admission policies, dynamic KV-cache allocation, and chunked pre-fill operations. This scheduler supports Sustained mode for long-context decoding (16K-200K prompts) by managing request-sized KV caches and sustaining throughput during extended generation sessions.

Cache Bank and Persistence

mtplx/cache_bank/__init__.py implements an on-disk SSD cache that stores session prefixes. This enables near-instant warm-prefix restoration across application restarts, eliminating redundant computation for recurring conversation contexts.

OpenAI-Compatible Server

Located in mtplx/server/openai.py, the REST API layer exposes standard endpoints including /v1/chat/completions, /v1/completions, /v1/embeddings, and /v1/rerank. This allows any OpenAI or Anthropic client library to communicate with local MTPLX models without modification.

Execution Flow

MTPLX operates through a five-stage pipeline that runs entirely on Apple Silicon GPUs via the MLX Metal-based tensor library:

  1. Model Loading – The runtime loads an MLX-converted checkpoint containing an MTP head. If no MTP head is present, MTPLX raises a clear "no-MTP" error rather than silently degrading to pure autoregressive mode.
  2. Drafting – The model produces a block of candidate tokens in one forward pass through the MTP head.
  3. Verification – A single batched forward pass evaluates the drafted block, computing exact acceptance probabilities for each position.
  4. Exact Rejection Sampling – The system accepts valid token blocks or performs partial rejection with residual correction, preserving the original model distribution exactly.
  5. Cache Update – Accepted tokens append to the KV-cache; rejected tokens are discarded, and the loop repeats.

Operating Modes

MTPLX provides distinct execution modes tailored to different hardware constraints and workload characteristics.

Turbo Mode

Turbo mode targets quantized 27B and 9B flagship models. This configuration activates the NAX-verify kernels and compiled verification paths in mtplx/turboquant.py to maximize token generation speed for short-to-medium contexts.

Sustained Mode

Sustained mode optimizes for long-context workloads and extended inference sessions. It employs chunked pre-fill strategies and request-sized KV-cache management from mtplx/batching/scheduler.py, making it optimal for prompts ranging from 16K to 200K tokens while maintaining thermal stability.

Sustained Max Mode

Sustained Max applies the same scheduling logic as Sustained mode but pins system fans at 100% via the thermal sidecar (mtplx/thermal_sidecar.py), allowing maximum sustained performance when acoustic and thermal constraints permit.

Burst Mode

Burst mode serves as a legacy short-context lane for benchmarks and rapid prototyping. It prioritizes immediate throughput over thermal management, resulting in louder CPU/GPU utilization during short prompt processing.

Getting Started with the API

MTPLX exposes both CLI commands and REST endpoints for integration.

Starting the Server

Launch the complete stack including GUI and API:

mtplx start

For headless operation on port 8000:

mtplx serve --port 8000

Sending Chat Requests

Standard OpenAI-compatible requests work immediately. Using curl:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"mtplx","messages":[{"role":"user","content":"Hello, world!"}],"stream":true}'

Using the Python OpenAI client:

import openai

client = openai.OpenAI(base_url="http://127.0.0.1:8000/v1")
resp = client.chat.completions.create(
    model="mtplx",
    messages=[{"role": "user", "content": "Explain MTP in one sentence"}],
)
print(resp.choices[0].message.content)

Extensibility Features

Beyond core inference, MTPLX provides tools for model conversion and multimodal support.

Forge Conversion Tool

The mtplx forge command (implemented in mtplx/commands/forge.py) converts Hugging Face repositories into MTPLX-ready models. This pipeline performs MLX conversion, trains MTP adapters, verifies speed and accuracy benchmarks, and optionally publishes artifacts to model hubs.

Vision Tower Multimodal Support

mtplx/vision/qwen3_vl_tower.py implements the Qwen-VL visual encoder for image-enabled chat. Integration works by preprocessing images into token sequences:

from mtplx.vision.qwen3_vl_tower import VisionTower

tower = VisionTower(model_name="mlx-community/Qwen3-VL-Chat-7B")
msg = tower.prepare_message(
    user_prompt="Describe the picture",
    image_path="example.jpg"
)

# Send `msg` to the chat endpoint

Performance Benchmarking

Run standardized evaluations using the built-in benchmark tool:

mtplx bench aime --quick

This executes the AIME benchmark on the currently loaded model and outputs a detailed performance report.

Summary

  • MTPLX accelerates local LLMs on Apple Silicon using native Multi-Token Prediction heads rather than external draft models.
  • The speculative engine in mtplx/speculative.py guarantees exact probability distributions through rigorous rejection sampling.
  • Turbo mode leverages mtplx/turboquant.py for quantized kernel acceleration, while Sustained mode handles long contexts via intelligent KV-cache management.
  • The system exposes an OpenAI-compatible API through mtplx/server/openai.py, supporting chat, embeddings, and reranking endpoints.
  • Forge enables custom model conversion, and the Vision Tower adds multimodal capabilities without disrupting the MTP pipeline.

Frequently Asked Questions

What hardware does MTPLX support?

MTPLX runs exclusively on Apple Silicon Macs using the MLX framework. All computations execute on the Metal Performance Shaders GPU backend, eliminating CUDA dependencies and enabling optimized memory sharing between CPU and unified GPU memory.

How does MTPLX differ from standard speculative decoding?

Traditional speculative decoding requires a separate smaller "draft" model to propose tokens, which is then verified by the larger target model. MTPLX instead uses Multi-Token Prediction heads already present in modern architectures (like those in Qwen2.5 or similar models) to generate candidate token blocks directly, eliminating the memory overhead and synchronization complexity of maintaining two separate models.

Can MTPLX run models without MTP heads?

No. According to the source code in mtplx/speculative.py, if a loaded model lacks an MTP head, MTPLX raises an explicit error and refuses to fall back silently to pure autoregressive generation. This prevents performance degradation without user awareness. Users must convert models using mtplx forge to add MTP adapters if native support is absent.

What is the Cache Bank and when does it help?

The Cache Bank (mtplx/cache_bank/__init__.py) is an SSD-backed storage system that persists session prefixes. It accelerates repeated conversations with similar system prompts or document contexts by restoring the KV-cache state from disk rather than recomputing embeddings, enabling near-instant warm starts for applications like coding assistants or long-form writing tools.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →