MTPLX Setup Guide: Install and Configure Multi-Token Prediction on macOS

MTPLX is a native macOS-first runtime for local large language models that implements multi-token prediction (MTP) to achieve up to 2.2× speed-ups on Apple Silicon through batched draft verification and exact rejection sampling.

This MTPLX setup guide walks you through installing the framework, launching the OpenAI-compatible daemon, and optimizing inference performance on Apple hardware. Whether you use the graphical macOS app or the Python CLI, the underlying architecture remains consistent: a background daemon handles heavy inference while the client manages model loading, tuning, and interaction.

Installation Methods

MTPLX distributes two primary installation paths depending on your workflow preferences.

Homebrew (macOS App + Daemon)

Install the full macOS application and background daemon using the official tap:

brew install youssofal/mtplx/mtplx

This installs the graphical wrapper that automatically manages daemon lifecycle, selects optimal models, and visualizes live decoding statistics through the UI components in mtplx/ui/.

pip (Python CLI + Library)

For headless servers or development workflows, install the Python package directly:

python3 -m pip install mtplx

The pip installation provides the mtplx command-line tool defined in mtplx/cli.py, which serves as the user-facing entry point for all operations including mtplx start, mtplx serve, and mtplx forge.

Architecture Overview

Understanding the three-tier architecture helps troubleshoot connection issues and optimize resource usage.

The Daemon-Client Model

At the core of MTPLX is a background OpenAI-compatible server defined in mtplx/daemon_client.py. This daemon exposes REST endpoints including /v1/chat/completions and /health, maintains a session cache, and executes the batched forward passes necessary for multi-token prediction. The CLI and macOS app function as clients that discover, attach to, or spawn this daemon process.

Core Inference Components

The runtime relies on several specialized modules:

Starting the Daemon

MTPLX offers two primary launch modes depending on whether you need persistent background service or temporary testing.

Interactive Launch

Start the daemon with automatic port detection and client attachment:

mtplx start

This command invokes detect_attachable_daemon() from mtplx/daemon_client.py to check for existing instances. If found, it attaches the REPL; otherwise, it initializes a new server process and connects the interactive chat interface.

Daemon-Only Mode

For API-only usage without the interactive REPL:

mtplx serve --port 8000

This launches the HTTP server on the specified port, enabling direct curl or client access:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"mtplx","messages":[{"role":"user","content":"hello"}],"stream":true}'

Graceful Shutdown

To stop a running daemon programmatically, use the Python client:

from mtplx.daemon_client import stop_daemon

result = stop_daemon(host="127.0.0.1", port=8000)
if result["ok"]:
    print(f"Daemon stopped (signal={result.get('signal')})")

The stop_daemon function handles port classification and ensures cleanup of the inference cache and Metal buffers.

Client Interaction Patterns

Once the daemon runs, interact through the high-level Python API or CLI attachment.

Python Client Library

Access the daemon directly without subprocess overhead:

import mtplx

client = mtplx.ApiClient(base_url="http://127.0.0.1:8000")
response = client.chat_completions(
    model="mtplx",
    messages=[{"role": "user", "content": "Write a poem about apples"}],
    stream=False,
)
print(response["choices"][0]["message"]["content"])

The ApiClient class wraps the OpenAI-compatible schema exposed by the daemon's HTTP layer.

REPL Attachment

Attach an interactive session to an existing daemon:

from mtplx.daemon_client import detect_attachable_daemon, run_attach_chat

daemon = detect_attachable_daemon()
if daemon:
    run_attach_chat(daemon, prompt="Explain multi-token prediction.")
else:
    print("No running MTPLX daemon found.")

The detect_attachable_daemon function respects the MTPLX_START_ATTACH_PROBE environment variable and classifies port occupancy to distinguish between PORT_FREE, PORT_APP_DAEMON, and conflicting services.

Configuration and Tuning

MTPLX stores persistent settings in ~/.mtplx/config.toml, managed through mtplx/config.py.

Runtime Settings

Query or modify configuration values:

mtplx settings get draft_depth
mtplx settings set draft_depth 4

These commands manipulate the TOML configuration file without requiring manual editing.

Hardware Auto-Tuning

Determine the optimal draft depth for your specific Apple Silicon chip:

mtplx tune --model mlx-community/Qwen3-27B-OptimizedSpeed --retune

This executes a benchmark suite across multiple draft depths in mtplx/turboquant.py, measures tokens-per-second, and persists the fastest configuration. The scheduler in mtplx/batching/scheduler.py references these values when batching requests.

Model Forging and Conversion

Convert Hugging Face checkpoints into MTPLX-optimized formats with built-in MTP heads:

mtplx forge <repository_path>

The forge command processes models like Qwen 3.5/3.6/3.8 and Gemma that support native multi-token prediction, preparing them for the verification kernels in mtplx/verify_qmv.py without requiring separate draft models.

Summary

  • MTPLX implements multi-token prediction on Apple Silicon through a daemon-client architecture, achieving up to 2.2× inference speed-ups via batched draft verification.
  • Install via Homebrew for the full macOS app experience or pip for the CLI and Python library.
  • The daemon in mtplx/daemon_client.py exposes OpenAI-compatible endpoints while the CLI in mtplx/cli.py handles lifecycle management.
  • Use detect_attachable_daemon() to connect to running instances and stop_daemon() for graceful shutdowns.
  • Configuration persists in ~/.mtplx/config.toml and is queryable via mtplx settings.
  • Run mtplx tune to automatically determine optimal draft depth for your hardware.

Frequently Asked Questions

What hardware requirements does MTPLX have?

MTPLX requires Apple Silicon (M1 or later) to leverage the optimized Metal kernels in mtplx/verify_qmv.py and the MLX-based inference engine. The framework uses unified memory architecture for efficient model loading and employs thermal profiling in mtplx/thermal.py to prevent throttling during extended inference sessions.

How do I stop a running MTPLX daemon?

Execute mtplx stop from the CLI or call stop_daemon(host, port) from Python. The function sends a graceful shutdown signal to the HTTP server, clears the session cache, and releases GPU resources. If the daemon was started via mtplx start, the CLI automatically cleans up the background process when you exit the REPL.

Can I use MTPLX with existing OpenAI-compatible clients?

Yes. The daemon implements the standard /v1/chat/completions and /v1/embeddings endpoints. Any client library that accepts a custom base_url—including the official OpenAI Python client, LangChain, or AutoGen—can connect to http://127.0.0.1:8000 (or your configured port) and interact with MTPLX models using standard API calls.

Where does MTPLX store its configuration and model cache?

Runtime configuration resides in ~/.mtplx/config.toml as defined in mtplx/config.py. Downloaded models and forged checkpoints typically cache in the standard MLX directory (~/.cache/mlx) unless overridden by environment variables. The daemon respects MTPLX_START_ATTACH_PROBE for port discovery and stores tuning profiles specific to your hardware configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →