# MTPLX Features: Native macOS LLM Inference with Multi-Token Prediction

> Explore MTPLX features: a native macOS app for faster LLM inference. Achieve up to 2x speedups on Apple Silicon using Multi-Token Prediction for local large language models.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: features
- Published: 2026-09-11

---

**MTPLX is a native macOS application and command-line tool that accelerates local large language model inference by leveraging Multi-Token Prediction (MTP) to draft and verify multiple tokens in a single batched forward pass, delivering up to 2× speedups on Apple Silicon.**

MTPLX features a complete speculative sampling pipeline implemented in the `youssofal/MTPLX` repository. Built specifically for Apple Silicon, it combines a thin CLI surface with backend-specific kernels, session caching, and an OpenAI-compatible server to enable efficient inference of models containing native MTP heads.

## Multi-Token Prediction Engine

At the core of MTPLX is a **speculative sampling pipeline** that utilizes the model's own MTP heads to draft several tokens ahead and verify them simultaneously. This approach applies the **Leviathan & Chen rejection sampling theorem** to ensure mathematical exactness while reducing per-token latency.

The architecture delegates proposal (draft) and verification to **backend-specific implementations** tailored to model families such as Qwen-3-Next, DeepSeek, and Gemma. As documented in [[`docs/architecture.md`](https://github.com/youssofal/MTPLX/blob/main/docs/architecture.md)](https://github.com/youssofal/MTPLX/blob/main/docs/architecture.md)#L23-L24, the speculative sampler exposes a backend-agnostic interface while executing optimized kernels underneath.

## Execution Profiles and Modes

MTPLX features distinct execution profiles that map to specific hardware utilization strategies:

- **Turbo Mode** – Employs NAX kernels for maximum throughput in short bursts
- **Sustained Mode** – Uses chunked prefill to maintain consistent performance over long generations
- **Burst Mode** – Optimized for sporadic, high-intensity inference tasks

These profiles interact with the **compatibility registry**, which inspects [`config.json`](https://github.com/youssofal/MTPLX/blob/main/config.json) and [`model.safetensors.index.json`](https://github.com/youssofal/MTPLX/blob/main/model.safetensors.index.json) to determine whether a model contains native MTP heads, is autoregressive-only, or incompatible ([README](https://github.com/youssofal/MTPLX/blob/main/README.md)#L46-L50).

## Session Management and Caching

The **SessionBank** feature caches KV-cache state across conversation turns, eliminating redundant computation for long-context interactions. For persistence across restarts, MTPLX provides an optional **SSD-backed session cache** that stores warm-prefix states on disk ([README](https://github.com/youssofal/MTPLX/blob/main/README.md)#L85-L86).

## Server Layer and API Compatibility

MTPLX exposes an OpenAI-compatible HTTP server via `mtplx serve` or `mtplx start`, listening on configurable ports (default localhost:8000). The server implements:

- `/v1/chat/completions` – Standard chat completions
- `/v1/completions` – Text completions  
- `/v1/models` – Model listing
- `/v1/messages` – Anthropic-compatible endpoint
- Optional `/v1/embeddings` and `/v1/rerank` – Retrieval endpoints ([README](https://github.com/youssofal/MTPLX/blob/main/README.md)#L75-L84)

Notably, embedding and reranker models bypass the MTP path entirely, as they output vectors rather than tokens ([README](https://github.com/youssofal/MTPLX/blob/main/README.md)#L87-L104).

## Multimodal Support

For vision-language models, MTPLX features **vision tower grafting** via `mtplx.vision_graft`. This utility restores vision tensors and side-cars from original checkpoints without modifying the language model or MTP weights, enabling multimodal capabilities while preserving speculative sampling performance ([[`mtplx/vision_graft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py)#L1-L19).

## Performance Optimization

### Auto-Tune

On first run, MTPLX measures real-world throughput at each draft depth, identifies the optimal configuration for the specific hardware, and persists this setting for subsequent sessions ([README](https://github.com/youssofal/MTPLX/blob/main/README.md)#L55-L62).

### Thermal Management

The **fan control subsystem** automatically pins cooling fans for sustained workloads and reliably restores automatic fan control upon process exit, preventing thermal throttling during intensive inference ([README](https://github.com/youssofal/MTPLX/blob/main/README.md)#L44-L46).

## Installation and Setup

Install MTPLX via Homebrew or PyPI:

```bash

# Homebrew (recommended)

brew install youssofal/mtplx/mtplx

# PyPI

python3 -m pip install -U mtplx

```

Initialize the fan control helper (requires sudo once):

```bash
mtplx max --install

```

## CLI Commands and Usage Examples

Health check and model management:

```bash

# Verify installation and hardware

mtplx doctor --summary

# Download optimized models

mtplx pull Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed

# Inspect model compatibility and MTP head configuration

mtplx inspect Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed --json

```

Interactive and server modes:

```bash

# Launch interactive chat UI with auto-selected draft depth

mtplx start

# Run OpenAI-compatible server only

mtplx serve --port 8000

```

Advanced operations:

```bash

# Benchmark and optimize draft depth for specific hardware

mtplx tune --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed --retune

# Convert Hugging Face repositories to MTP format

mtplx forge --repo huggingface.co/your-org/your-mtp-model \
  --output ./my-mtp-model

```

Enable retrieval-augmented generation endpoints:

```bash
mtplx serve \
  --embedding-model mlx-community/Qwen3-Embedding-8B-4bit-DWQ \
  --reranker-model vserifsaglam/Qwen3-Reranker-4B-4bit-MLX

```

## Python Client Integration

Integrate with existing applications using the OpenAI SDK:

```python
import openai

client = openai.OpenAI(base_url="http://127.0.0.1:8000/v1")
resp = client.chat.completions.create(
    model="mtplx",
    messages=[{"role": "user", "content": "What is MTPLX?"}],
    temperature=0.6,
    top_p=0.95,
)
print(resp.choices[0].message.content)

```

## Key Source Files

| File | Purpose |
|------|---------|
| [[`README.md`](https://github.com/youssofal/MTPLX/blob/main/README.md)](https://github.com/youssofal/MTPLX/blob/main/README.md) | Installation guide, mode tables, and API endpoint documentation |
| [[`docs/architecture.md`](https://github.com/youssofal/MTPLX/blob/main/docs/architecture.md)](https://github.com/youssofal/MTPLX/blob/main/docs/architecture.md) | Visual architecture diagram and CLI-to-backend flow |
| [[`mtplx/vision_graft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py) | Vision tensor restoration for multimodal models |
| [[`docs/quickstart.md`](https://github.com/youssofal/MTPLX/blob/main/docs/quickstart.md)](https://github.com/youssofal/MTPLX/blob/main/docs/quickstart.md) | Step-by-step setup and server launch instructions |
| [[`docs/concurrency.md`](https://github.com/youssofal/MTPLX/blob/main/docs/concurrency.md)](https://github.com/youssofal/MTPLX/blob/main/docs/concurrency.md) | Scheduler modes (Turbo, Sustained, Burst) and concurrency details |

## Summary

- **MTPLX features** native Multi-Token Prediction utilizing the model's speculative heads for exact draft-and-verify sampling
- **Execution profiles** (Turbo, Sustained, Burst) map to specific kernel implementations optimized for Apple Silicon
- **SessionBank and SSD caching** persist KV-cache states across conversation turns and restarts
- **OpenAI-compatible server** exposes standard chat, completion, embedding, and rerank endpoints
- **Vision grafting** in [[`mtplx/vision_graft.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision_graft.py) enables multimodal support without compromising MTP performance
- **Auto-tune** benchmarks draft depth on-device to maximize tokens per second
- **Thermal management** automates fan control for sustained workloads

## Frequently Asked Questions

### What hardware is required to run MTPLX?

MTPLX requires Apple Silicon (M1, M2, M3, or M4 series) running macOS. The tool leverages the Neural Engine and unified memory architecture specific to Apple chips to execute the NAX kernels and chunked prefill operations described in the [concurrency documentation](https://github.com/youssofal/MTPLX/blob/main/docs/concurrency.md).

### How does MTPLX differ from standard autoregressive decoding?

Standard autoregressive decoding generates one token per forward pass. MTPLX features **Multi-Token Prediction** that drafts multiple future tokens using the model's native MTP heads, then verifies them in a single batched forward pass. This speculative approach follows the Leviathan & Chen rejection sampling theorem to ensure identical output distributions while achieving approximately 2× speedup compared to standard decoding.

### Can MTPLX run models without native MTP heads?

Yes, but with fallback behavior. The **compatibility registry** inspects [`config.json`](https://github.com/youssofal/MTPLX/blob/main/config.json) and [`model.safetensors.index.json`](https://github.com/youssofal/MTPLX/blob/main/model.safetensors.index.json) to detect MTP head presence. Models lacking these heads run in autoregressive-only mode without speculative acceleration. The `mtplx inspect` command reveals compatibility status before downloading weights.

### Does MTPLX support function calling and structured output?

While the server exposes OpenAI-compatible endpoints including `/v1/chat/completions`, specific capabilities like function calling depend on the underlying model weights and tokenizer configuration. The HTTP API layer itself supports the standard request/response schema, allowing clients to send `tools` and `tool_choice` parameters compatible with the served model's capabilities.