MTPLX Setup Guide: Install and Configure Multi-Token Prediction on macOS
MTPLX is a native macOS-first runtime for local large language models that implements multi-token prediction (MTP) to achieve up to 2.2× speed-ups on Apple Silicon through batched draft verification and exact rejection sampling.
This MTPLX setup guide walks you through installing the framework, launching the OpenAI-compatible daemon, and optimizing inference performance on Apple hardware. Whether you use the graphical macOS app or the Python CLI, the underlying architecture remains consistent: a background daemon handles heavy inference while the client manages model loading, tuning, and interaction.
Installation Methods
MTPLX distributes two primary installation paths depending on your workflow preferences.
Homebrew (macOS App + Daemon)
Install the full macOS application and background daemon using the official tap:
brew install youssofal/mtplx/mtplx
This installs the graphical wrapper that automatically manages daemon lifecycle, selects optimal models, and visualizes live decoding statistics through the UI components in mtplx/ui/.
pip (Python CLI + Library)
For headless servers or development workflows, install the Python package directly:
python3 -m pip install mtplx
The pip installation provides the mtplx command-line tool defined in mtplx/cli.py, which serves as the user-facing entry point for all operations including mtplx start, mtplx serve, and mtplx forge.
Architecture Overview
Understanding the three-tier architecture helps troubleshoot connection issues and optimize resource usage.
The Daemon-Client Model
At the core of MTPLX is a background OpenAI-compatible server defined in mtplx/daemon_client.py. This daemon exposes REST endpoints including /v1/chat/completions and /health, maintains a session cache, and executes the batched forward passes necessary for multi-token prediction. The CLI and macOS app function as clients that discover, attach to, or spawn this daemon process.
Core Inference Components
The runtime relies on several specialized modules:
mtplx/verify_qmv.py– Contains verification kernels that execute the single batched forward pass for draft token blocksmtplx/batching/scheduler.py– Implements dynamic request scheduling and determines optimal draft depth based on hardware profilingmtplx/correctors/low_rank.py– Applies residual correction after verification to maintain exact output distributionsmtplx/vision/qwen3_vl_tower.py– Optional vision tower for multimodal Qwen 3 model support
Starting the Daemon
MTPLX offers two primary launch modes depending on whether you need persistent background service or temporary testing.
Interactive Launch
Start the daemon with automatic port detection and client attachment:
mtplx start
This command invokes detect_attachable_daemon() from mtplx/daemon_client.py to check for existing instances. If found, it attaches the REPL; otherwise, it initializes a new server process and connects the interactive chat interface.
Daemon-Only Mode
For API-only usage without the interactive REPL:
mtplx serve --port 8000
This launches the HTTP server on the specified port, enabling direct curl or client access:
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"mtplx","messages":[{"role":"user","content":"hello"}],"stream":true}'
Graceful Shutdown
To stop a running daemon programmatically, use the Python client:
from mtplx.daemon_client import stop_daemon
result = stop_daemon(host="127.0.0.1", port=8000)
if result["ok"]:
print(f"Daemon stopped (signal={result.get('signal')})")
The stop_daemon function handles port classification and ensures cleanup of the inference cache and Metal buffers.
Client Interaction Patterns
Once the daemon runs, interact through the high-level Python API or CLI attachment.
Python Client Library
Access the daemon directly without subprocess overhead:
import mtplx
client = mtplx.ApiClient(base_url="http://127.0.0.1:8000")
response = client.chat_completions(
model="mtplx",
messages=[{"role": "user", "content": "Write a poem about apples"}],
stream=False,
)
print(response["choices"][0]["message"]["content"])
The ApiClient class wraps the OpenAI-compatible schema exposed by the daemon's HTTP layer.
REPL Attachment
Attach an interactive session to an existing daemon:
from mtplx.daemon_client import detect_attachable_daemon, run_attach_chat
daemon = detect_attachable_daemon()
if daemon:
run_attach_chat(daemon, prompt="Explain multi-token prediction.")
else:
print("No running MTPLX daemon found.")
The detect_attachable_daemon function respects the MTPLX_START_ATTACH_PROBE environment variable and classifies port occupancy to distinguish between PORT_FREE, PORT_APP_DAEMON, and conflicting services.
Configuration and Tuning
MTPLX stores persistent settings in ~/.mtplx/config.toml, managed through mtplx/config.py.
Runtime Settings
Query or modify configuration values:
mtplx settings get draft_depth
mtplx settings set draft_depth 4
These commands manipulate the TOML configuration file without requiring manual editing.
Hardware Auto-Tuning
Determine the optimal draft depth for your specific Apple Silicon chip:
mtplx tune --model mlx-community/Qwen3-27B-OptimizedSpeed --retune
This executes a benchmark suite across multiple draft depths in mtplx/turboquant.py, measures tokens-per-second, and persists the fastest configuration. The scheduler in mtplx/batching/scheduler.py references these values when batching requests.
Model Forging and Conversion
Convert Hugging Face checkpoints into MTPLX-optimized formats with built-in MTP heads:
mtplx forge <repository_path>
The forge command processes models like Qwen 3.5/3.6/3.8 and Gemma that support native multi-token prediction, preparing them for the verification kernels in mtplx/verify_qmv.py without requiring separate draft models.
Summary
- MTPLX implements multi-token prediction on Apple Silicon through a daemon-client architecture, achieving up to 2.2× inference speed-ups via batched draft verification.
- Install via Homebrew for the full macOS app experience or pip for the CLI and Python library.
- The daemon in
mtplx/daemon_client.pyexposes OpenAI-compatible endpoints while the CLI inmtplx/cli.pyhandles lifecycle management. - Use
detect_attachable_daemon()to connect to running instances andstop_daemon()for graceful shutdowns. - Configuration persists in
~/.mtplx/config.tomland is queryable viamtplx settings. - Run
mtplx tuneto automatically determine optimal draft depth for your hardware.
Frequently Asked Questions
What hardware requirements does MTPLX have?
MTPLX requires Apple Silicon (M1 or later) to leverage the optimized Metal kernels in mtplx/verify_qmv.py and the MLX-based inference engine. The framework uses unified memory architecture for efficient model loading and employs thermal profiling in mtplx/thermal.py to prevent throttling during extended inference sessions.
How do I stop a running MTPLX daemon?
Execute mtplx stop from the CLI or call stop_daemon(host, port) from Python. The function sends a graceful shutdown signal to the HTTP server, clears the session cache, and releases GPU resources. If the daemon was started via mtplx start, the CLI automatically cleans up the background process when you exit the REPL.
Can I use MTPLX with existing OpenAI-compatible clients?
Yes. The daemon implements the standard /v1/chat/completions and /v1/embeddings endpoints. Any client library that accepts a custom base_url—including the official OpenAI Python client, LangChain, or AutoGen—can connect to http://127.0.0.1:8000 (or your configured port) and interact with MTPLX models using standard API calls.
Where does MTPLX store its configuration and model cache?
Runtime configuration resides in ~/.mtplx/config.toml as defined in mtplx/config.py. Downloaded models and forged checkpoints typically cache in the standard MLX directory (~/.cache/mlx) unless overridden by environment variables. The daemon respects MTPLX_START_ATTACH_PROBE for port discovery and stores tuning profiles specific to your hardware configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →