# MTPLX Setup Guide: Install and Configure Multi-Token Prediction on macOS

> Install and configure MTPLX on macOS with this comprehensive setup guide. Learn how to leverage multi-token prediction for up to 2.2x faster local LLM inference on Apple Silicon.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: getting-started
- Published: 2026-09-11

---

**MTPLX is a native macOS-first runtime for local large language models that implements multi-token prediction (MTP) to achieve up to 2.2× speed-ups on Apple Silicon through batched draft verification and exact rejection sampling.**

This MTPLX setup guide walks you through installing the framework, launching the OpenAI-compatible daemon, and optimizing inference performance on Apple hardware. Whether you use the graphical macOS app or the Python CLI, the underlying architecture remains consistent: a background daemon handles heavy inference while the client manages model loading, tuning, and interaction.

## Installation Methods

MTPLX distributes two primary installation paths depending on your workflow preferences.

### Homebrew (macOS App + Daemon)

Install the full macOS application and background daemon using the official tap:

```bash
brew install youssofal/mtplx/mtplx

```

This installs the graphical wrapper that automatically manages daemon lifecycle, selects optimal models, and visualizes live decoding statistics through the UI components in `mtplx/ui/`.

### pip (Python CLI + Library)

For headless servers or development workflows, install the Python package directly:

```bash
python3 -m pip install mtplx

```

The pip installation provides the `mtplx` command-line tool defined in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py), which serves as the user-facing entry point for all operations including `mtplx start`, `mtplx serve`, and `mtplx forge`.

## Architecture Overview

Understanding the three-tier architecture helps troubleshoot connection issues and optimize resource usage.

### The Daemon-Client Model

At the core of MTPLX is a background OpenAI-compatible server defined in [`mtplx/daemon_client.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/daemon_client.py). This daemon exposes REST endpoints including `/v1/chat/completions` and `/health`, maintains a session cache, and executes the batched forward passes necessary for multi-token prediction. The CLI and macOS app function as clients that discover, attach to, or spawn this daemon process.

### Core Inference Components

The runtime relies on several specialized modules:

- **[`mtplx/verify_qmv.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/verify_qmv.py)** – Contains verification kernels that execute the single batched forward pass for draft token blocks
- **[`mtplx/batching/scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/scheduler.py)** – Implements dynamic request scheduling and determines optimal **draft depth** based on hardware profiling
- **[`mtplx/correctors/low_rank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/correctors/low_rank.py)** – Applies residual correction after verification to maintain exact output distributions
- **[`mtplx/vision/qwen3_vl_tower.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/vision/qwen3_vl_tower.py)** – Optional vision tower for multimodal Qwen 3 model support

## Starting the Daemon

MTPLX offers two primary launch modes depending on whether you need persistent background service or temporary testing.

### Interactive Launch

Start the daemon with automatic port detection and client attachment:

```bash
mtplx start

```

This command invokes `detect_attachable_daemon()` from [`mtplx/daemon_client.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/daemon_client.py) to check for existing instances. If found, it attaches the REPL; otherwise, it initializes a new server process and connects the interactive chat interface.

### Daemon-Only Mode

For API-only usage without the interactive REPL:

```bash
mtplx serve --port 8000

```

This launches the HTTP server on the specified port, enabling direct curl or client access:

```bash
curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"mtplx","messages":[{"role":"user","content":"hello"}],"stream":true}'

```

### Graceful Shutdown

To stop a running daemon programmatically, use the Python client:

```python
from mtplx.daemon_client import stop_daemon

result = stop_daemon(host="127.0.0.1", port=8000)
if result["ok"]:
    print(f"Daemon stopped (signal={result.get('signal')})")

```

The `stop_daemon` function handles port classification and ensures cleanup of the inference cache and Metal buffers.

## Client Interaction Patterns

Once the daemon runs, interact through the high-level Python API or CLI attachment.

### Python Client Library

Access the daemon directly without subprocess overhead:

```python
import mtplx

client = mtplx.ApiClient(base_url="http://127.0.0.1:8000")
response = client.chat_completions(
    model="mtplx",
    messages=[{"role": "user", "content": "Write a poem about apples"}],
    stream=False,
)
print(response["choices"][0]["message"]["content"])

```

The `ApiClient` class wraps the OpenAI-compatible schema exposed by the daemon's HTTP layer.

### REPL Attachment

Attach an interactive session to an existing daemon:

```python
from mtplx.daemon_client import detect_attachable_daemon, run_attach_chat

daemon = detect_attachable_daemon()
if daemon:
    run_attach_chat(daemon, prompt="Explain multi-token prediction.")
else:
    print("No running MTPLX daemon found.")

```

The `detect_attachable_daemon` function respects the `MTPLX_START_ATTACH_PROBE` environment variable and classifies port occupancy to distinguish between `PORT_FREE`, `PORT_APP_DAEMON`, and conflicting services.

## Configuration and Tuning

MTPLX stores persistent settings in `~/.mtplx/config.toml`, managed through [`mtplx/config.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/config.py).

### Runtime Settings

Query or modify configuration values:

```bash
mtplx settings get draft_depth
mtplx settings set draft_depth 4

```

These commands manipulate the TOML configuration file without requiring manual editing.

### Hardware Auto-Tuning

Determine the optimal draft depth for your specific Apple Silicon chip:

```bash
mtplx tune --model mlx-community/Qwen3-27B-OptimizedSpeed --retune

```

This executes a benchmark suite across multiple draft depths in [`mtplx/turboquant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/turboquant.py), measures tokens-per-second, and persists the fastest configuration. The scheduler in [`mtplx/batching/scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/scheduler.py) references these values when batching requests.

## Model Forging and Conversion

Convert Hugging Face checkpoints into MTPLX-optimized formats with built-in MTP heads:

```bash
mtplx forge <repository_path>

```

The forge command processes models like Qwen 3.5/3.6/3.8 and Gemma that support native multi-token prediction, preparing them for the verification kernels in [`mtplx/verify_qmv.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/verify_qmv.py) without requiring separate draft models.

## Summary

- **MTPLX** implements multi-token prediction on Apple Silicon through a daemon-client architecture, achieving up to 2.2× inference speed-ups via batched draft verification.
- Install via **Homebrew** for the full macOS app experience or **pip** for the CLI and Python library.
- The daemon in [`mtplx/daemon_client.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/daemon_client.py) exposes OpenAI-compatible endpoints while the CLI in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) handles lifecycle management.
- Use `detect_attachable_daemon()` to connect to running instances and `stop_daemon()` for graceful shutdowns.
- Configuration persists in `~/.mtplx/config.toml` and is queryable via `mtplx settings`.
- Run `mtplx tune` to automatically determine optimal draft depth for your hardware.

## Frequently Asked Questions

### What hardware requirements does MTPLX have?

MTPLX requires Apple Silicon (M1 or later) to leverage the optimized Metal kernels in [`mtplx/verify_qmv.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/verify_qmv.py) and the MLX-based inference engine. The framework uses unified memory architecture for efficient model loading and employs thermal profiling in [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py) to prevent throttling during extended inference sessions.

### How do I stop a running MTPLX daemon?

Execute `mtplx stop` from the CLI or call `stop_daemon(host, port)` from Python. The function sends a graceful shutdown signal to the HTTP server, clears the session cache, and releases GPU resources. If the daemon was started via `mtplx start`, the CLI automatically cleans up the background process when you exit the REPL.

### Can I use MTPLX with existing OpenAI-compatible clients?

Yes. The daemon implements the standard `/v1/chat/completions` and `/v1/embeddings` endpoints. Any client library that accepts a custom `base_url`—including the official OpenAI Python client, LangChain, or AutoGen—can connect to `http://127.0.0.1:8000` (or your configured port) and interact with MTPLX models using standard API calls.

### Where does MTPLX store its configuration and model cache?

Runtime configuration resides in `~/.mtplx/config.toml` as defined in [`mtplx/config.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/config.py). Downloaded models and forged checkpoints typically cache in the standard MLX directory (`~/.cache/mlx`) unless overridden by environment variables. The daemon respects `MTPLX_START_ATTACH_PROBE` for port discovery and stores tuning profiles specific to your hardware configuration.