MTPLX Features: Native macOS LLM Inference with Multi-Token Prediction

MTPLX is a native macOS application and command-line tool that accelerates local large language model inference by leveraging Multi-Token Prediction (MTP) to draft and verify multiple tokens in a single batched forward pass, delivering up to 2× speedups on Apple Silicon.

MTPLX features a complete speculative sampling pipeline implemented in the youssofal/MTPLX repository. Built specifically for Apple Silicon, it combines a thin CLI surface with backend-specific kernels, session caching, and an OpenAI-compatible server to enable efficient inference of models containing native MTP heads.

Multi-Token Prediction Engine

At the core of MTPLX is a speculative sampling pipeline that utilizes the model's own MTP heads to draft several tokens ahead and verify them simultaneously. This approach applies the Leviathan & Chen rejection sampling theorem to ensure mathematical exactness while reducing per-token latency.

The architecture delegates proposal (draft) and verification to backend-specific implementations tailored to model families such as Qwen-3-Next, DeepSeek, and Gemma. As documented in [docs/architecture.md](https://github.com/youssofal/MTPLX/blob/main/docs/architecture.md)#L23-L24, the speculative sampler exposes a backend-agnostic interface while executing optimized kernels underneath.

Execution Profiles and Modes

MTPLX features distinct execution profiles that map to specific hardware utilization strategies:

  • Turbo Mode – Employs NAX kernels for maximum throughput in short bursts
  • Sustained Mode – Uses chunked prefill to maintain consistent performance over long generations
  • Burst Mode – Optimized for sporadic, high-intensity inference tasks

These profiles interact with the compatibility registry, which inspects config.json and model.safetensors.index.json to determine whether a model contains native MTP heads, is autoregressive-only, or incompatible (README#L46-L50).

Session Management and Caching

The SessionBank feature caches KV-cache state across conversation turns, eliminating redundant computation for long-context interactions. For persistence across restarts, MTPLX provides an optional SSD-backed session cache that stores warm-prefix states on disk (README#L85-L86).

Server Layer and API Compatibility

MTPLX exposes an OpenAI-compatible HTTP server via mtplx serve or mtplx start, listening on configurable ports (default localhost:8000). The server implements:

  • /v1/chat/completions – Standard chat completions
  • /v1/completions – Text completions
  • /v1/models – Model listing
  • /v1/messages – Anthropic-compatible endpoint
  • Optional /v1/embeddings and /v1/rerank – Retrieval endpoints (README#L75-L84)

Notably, embedding and reranker models bypass the MTP path entirely, as they output vectors rather than tokens (README#L87-L104).

Multimodal Support

For vision-language models, MTPLX features vision tower grafting via mtplx.vision_graft. This utility restores vision tensors and side-cars from original checkpoints without modifying the language model or MTP weights, enabling multimodal capabilities while preserving speculative sampling performance ([mtplx/vision_graft.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py)#L1-L19).

Performance Optimization

Auto-Tune

On first run, MTPLX measures real-world throughput at each draft depth, identifies the optimal configuration for the specific hardware, and persists this setting for subsequent sessions (README#L55-L62).

Thermal Management

The fan control subsystem automatically pins cooling fans for sustained workloads and reliably restores automatic fan control upon process exit, preventing thermal throttling during intensive inference (README#L44-L46).

Installation and Setup

Install MTPLX via Homebrew or PyPI:


# Homebrew (recommended)

brew install youssofal/mtplx/mtplx

# PyPI

python3 -m pip install -U mtplx

Initialize the fan control helper (requires sudo once):

mtplx max --install

CLI Commands and Usage Examples

Health check and model management:


# Verify installation and hardware

mtplx doctor --summary

# Download optimized models

mtplx pull Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed

# Inspect model compatibility and MTP head configuration

mtplx inspect Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed --json

Interactive and server modes:


# Launch interactive chat UI with auto-selected draft depth

mtplx start

# Run OpenAI-compatible server only

mtplx serve --port 8000

Advanced operations:


# Benchmark and optimize draft depth for specific hardware

mtplx tune --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed --retune

# Convert Hugging Face repositories to MTP format

mtplx forge --repo huggingface.co/your-org/your-mtp-model \
  --output ./my-mtp-model

Enable retrieval-augmented generation endpoints:

mtplx serve \
  --embedding-model mlx-community/Qwen3-Embedding-8B-4bit-DWQ \
  --reranker-model vserifsaglam/Qwen3-Reranker-4B-4bit-MLX

Python Client Integration

Integrate with existing applications using the OpenAI SDK:

import openai

client = openai.OpenAI(base_url="http://127.0.0.1:8000/v1")
resp = client.chat.completions.create(
    model="mtplx",
    messages=[{"role": "user", "content": "What is MTPLX?"}],
    temperature=0.6,
    top_p=0.95,
)
print(resp.choices[0].message.content)

Key Source Files

File Purpose
[README.md](https://github.com/youssofal/MTPLX/blob/main/README.md) Installation guide, mode tables, and API endpoint documentation
[docs/architecture.md](https://github.com/youssofal/MTPLX/blob/main/docs/architecture.md) Visual architecture diagram and CLI-to-backend flow
[mtplx/vision_graft.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py) Vision tensor restoration for multimodal models
[docs/quickstart.md](https://github.com/youssofal/MTPLX/blob/main/docs/quickstart.md) Step-by-step setup and server launch instructions
[docs/concurrency.md](https://github.com/youssofal/MTPLX/blob/main/docs/concurrency.md) Scheduler modes (Turbo, Sustained, Burst) and concurrency details

Summary

  • MTPLX features native Multi-Token Prediction utilizing the model's speculative heads for exact draft-and-verify sampling
  • Execution profiles (Turbo, Sustained, Burst) map to specific kernel implementations optimized for Apple Silicon
  • SessionBank and SSD caching persist KV-cache states across conversation turns and restarts
  • OpenAI-compatible server exposes standard chat, completion, embedding, and rerank endpoints
  • Vision grafting in [mtplx/vision_graft.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py) enables multimodal support without compromising MTP performance
  • Auto-tune benchmarks draft depth on-device to maximize tokens per second
  • Thermal management automates fan control for sustained workloads

Frequently Asked Questions

What hardware is required to run MTPLX?

MTPLX requires Apple Silicon (M1, M2, M3, or M4 series) running macOS. The tool leverages the Neural Engine and unified memory architecture specific to Apple chips to execute the NAX kernels and chunked prefill operations described in the concurrency documentation.

How does MTPLX differ from standard autoregressive decoding?

Standard autoregressive decoding generates one token per forward pass. MTPLX features Multi-Token Prediction that drafts multiple future tokens using the model's native MTP heads, then verifies them in a single batched forward pass. This speculative approach follows the Leviathan & Chen rejection sampling theorem to ensure identical output distributions while achieving approximately 2× speedup compared to standard decoding.

Can MTPLX run models without native MTP heads?

Yes, but with fallback behavior. The compatibility registry inspects config.json and model.safetensors.index.json to detect MTP head presence. Models lacking these heads run in autoregressive-only mode without speculative acceleration. The mtplx inspect command reveals compatibility status before downloading weights.

Does MTPLX support function calling and structured output?

While the server exposes OpenAI-compatible endpoints including /v1/chat/completions, specific capabilities like function calling depend on the underlying model weights and tokenizer configuration. The HTTP API layer itself supports the standard request/response schema, allowing clients to send tools and tool_choice parameters compatible with the served model's capabilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →