MTPLX Dependencies: Complete Installation Guide for Core, Server, and Native Extensions

MTPLX requires platform-specific MLX libraries on Apple Silicon, core scientific packages like NumPy and Pydantic, and optional extras for server deployment and benchmarking, all declared in pyproject.toml and native_extensions/qsa_kernels/setup.py.

MTPLX is a high-performance speculative decoding engine maintained in the youssofal/MTPLX repository. To install and run the package correctly, you must understand its tiered dependency structure, which ranges from core scientific libraries to platform-specific Metal kernels and optional web service components.

Core Runtime and Scientific Stack

The essential packages required for MTPLX to function on any supported platform are declared in pyproject.toml lines [14‑52]. These dependencies form the backbone of the library’s tensor operations and API interfaces.

Key packages in this group include NumPy ≥ 2, Pydantic ≥ 2, Rich ≥ 14, Safetensors ≥ 0.6, and Uvicorn ≥ 0.46. Additionally, Pillow is required early in the import chain because mtplx.vision.processing needs immediate access to image decoding capabilities for vision model support.

Apple Silicon Platform Requirements

Running MTPLX on macOS Apple Silicon requires additional GPU-specific libraries declared in pyproject.toml lines [33‑45]. These dependencies must maintain strict version alignment with the compiled native extension.

The platform-specific requirements are:

  • MLX ≥ 0.32.2 < 0.33 — the low-level Metal backend that powers the speculative decoding engine
  • mlx-lm 0.31 – 0.32 — the language model integration layer for MLX
  • Transformers ≠ 5.13.0 < 5.15 — supplies tokenizer and model-configuration utilities while avoiding a known breaking change introduced in version 5.13.0

These versions must exactly match the ABI expected by the compiled kernels. The mlx version is particularly sensitive because the native extension is built against a specific Metal API surface.

Native Extension Build Requirements

The mtplx_qsa_kernels component provides a hand-tuned Metal kernel for Qwen4-QSA sparse-GQA decoding. This compiled extension is defined in native_extensions/qsa_kernels/setup.py at line [47].

The build script declares mlx==0.32.2 as a strict requirement in its install_requires. This pins the extension to a specific MLX ABI, ensuring that the Metal code generation remains compatible with the runtime libraries. When you install MTPLX from source or build the extension manually, pip resolves this constraint automatically.


# Build the native QSA kernel extension manually (rarely needed)

cd native_extensions/qsa_kernels
python setup.py install   # pulls mlx==0.32.2 as declared in install_requires

Optional Extras for Extended Functionality

MTPLX provides three optional dependency groups declared later in pyproject.toml for specialized use cases.

Server Deployment

The server extra, defined in lines [77‑81], adds FastAPI, llguidance, and Uvicorn to the environment. These enable MTPLX to expose a REST-style API capable of receiving image attachments, JSON-encoded prompts, and tool-call specifications.


# Install with the server extras – useful for running the HTTP API

pip install "mtplx[server]"

After installation, launch the service with:

python -m mtplx.server  # starts a FastAPI app on http://127.0.0.1:8000

Benchmarking Tools

The competitors extra, located in lines [73‑76], provides the dflash-mlx library for performance benchmarking against alternative decoding implementations.


# Install the competitive benchmark library

pip install "mtplx[competitors]"

Development Tools

The dev extra, defined in lines [82‑87], includes tools for building, testing, and releasing the package. These are not required for runtime operation but are necessary for contributors modifying the source code.


# Install the development extras (testing, linting, packaging)

pip install "mtplx[dev]"

Installation Commands

Platform-specific selectors in the package configuration ensure that macOS Apple Silicon users receive the correct MLX wheels while other platforms skip them automatically. Use the following commands to install the appropriate variant for your environment:


# Install the core library (includes Pillow, numpy, etc.)

pip install mtplx

# Install both server and dev extras for a full development environment

pip install "mtplx[server,dev]"

Verifying Your Installation

After installation, verify that all components imported correctly and that version constraints were respected:


# Quick sanity-check after installation

import mtplx
print(mtplx.__version__)          # → 2.11.2

print(mtplx.__all__)              # shows exported symbols, confirming imports succeeded

# Verify MLX availability on Apple Silicon

try:
    import mlx.core as mx
    print(f"MLX version: {mx.__version__}")  # should report 0.32.2

except ImportError:
    print("MLX not installed – expected on non-Apple platforms")

Summary

  • Core dependencies in pyproject.toml lines [14‑52] provide the scientific stack (NumPy, Pydantic, Rich) and vision support (Pillow) required on all platforms.
  • Apple Silicon users must install specific MLX versions (≥0.32.2 < 0.33) and avoid Transformers 5.13.0 to maintain compatibility with the native kernels.
  • The native extension (mtplx_qsa_kernels) is built from native_extensions/qsa_kernels/setup.py and pins mlx==0.32.2 for Metal ABI stability.
  • Optional extras add FastAPI for REST serving, dflash-mlx for benchmarking, and development tools for contributors.

Frequently Asked Questions

What Python version does MTPLX require?

The package is built against modern Python standards and requires Python 3.9 or newer to ensure compatibility with the type hints and async patterns used in the FastAPI server components.

Do I need Apple Silicon to run MTPLX?

No, MTPLX installs and runs on other platforms, but the speculative decoding acceleration relies on MLX, which is exclusive to Apple Silicon. On non-Apple hardware, the package will function but will not utilize GPU acceleration via the Metal kernels.

Why is Transformers version 5.13.0 excluded?

Version 5.13.0 of the Transformers library introduced a breaking change in tokenizer serialization that conflicts with MTPLX’s model configuration loading. The constraint !=5.13.0 <5.15 ensures MTPLX receives either a stable earlier version or the subsequent patch that resolves the incompatibility.

What is the purpose of the mtplx_qsa_kernels native extension?

This extension supplies a compiled Metal kernel (mtplx_qsa_kernels._ext) specifically optimized for Qwen4-QSA sparse-GQA decoding. It accelerates the attention mechanism in speculative decoding workloads by executing sparse grouped-query attention operations directly on the GPU, bypassing Python interpreter overhead for critical tensor operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →