# Key Files in the MTPLX Repository for Understanding Its Runtime: A Complete Guide

> Explore key files in the MTPLX repository like cli.py and runtime_options.py to understand its runtime. Learn about request scheduling and backend implementations.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-08

---

**The most important files for understanding the MTPLX runtime are [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) for the entry point, [`mtplx/runtime_options.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime_options.py) for configuration normalization, [`mtplx/batching/scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/scheduler.py) for request scheduling, the backend files in `mtplx/backends/` for model implementations, and `mtplx/kernels/` for low-level Apple Silicon optimizations.**

MTPLX is a high-performance, Apple-Silicon-first inference engine that implements **native-MTP speculative decoding** with **adaptive batch scheduling** and **hardware-aware quantization**. Understanding the **key files in the MTPLX repository** reveals exactly how the system transforms CLI arguments into optimized GPU kernel execution. This guide traces the complete execution pipeline from argument parsing through low-level Metal kernel invocation.

## The MTPLX Runtime Execution Flow

The MTPLX runtime operates as a deterministic pipeline that processes user requests through eight distinct stages. Execution begins at the command-line interface, flows through configuration normalization, passes into the batching scheduler, executes through model-specific backends, and terminates at the low-level kernel layer.

The pipeline progresses as follows:

1. **[`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py)** builds a `Namespace` of user flags through `build_parser()`.
2. **[`mtplx/runtime_options.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime_options.py)** normalizes flags via `canonicalize_flag_tokens()` and resolves environment overrides like `MTPLX_PAGED_KV_QUANT`.
3. **[`mtplx/batching/scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/scheduler.py)** selects a `SchedulerMode` (serial, ar-batch, or hyper) and determines draft depth.
4. **[`mtplx/adaptive.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/adaptive.py)** applies the `AdaptivePolicy` to tune speculative decoding steps dynamically.
5. **`mtplx/backends/*.py`** executes the model-specific `forward()` and `prefill()` implementations.
6. **`mtplx/kernels/*.py`** invokes Metal kernels for attention, MLP, and quantized KV operations.
7. **[`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py)** manages the KV-cache hierarchy with optional SSD spillover controlled by `MTPLX_SESSION_BLOCK_PREFIX_RESTORE`.
8. **[`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py)** monitors thermal state via `detect_thermal_control()` to maintain consistent performance.

## Entry Point and Command-Line Interface

[`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) serves as the primary entry point where the runtime begins execution. The `build_parser()` function constructs the argument parser and defines the command hierarchy, while the `_format_*_help` functions generate both public and advanced help documentation.

When you launch MTPLX, the CLI immediately dispatches to sub-commands based on the parsed `Namespace`. For interactive use, the `mtplx start` command initializes the chat interface:

```bash

# Start the interactive onboarding wizard and then drop into the chat UI

mtplx start

```

For one-shot generation without the UI, use the `ask` sub-command:

```bash

# Ask a single question and exit (use the default model & profile)

mtplx ask "Explain the core idea of native-MTP in 2 sentences"

```

## Configuration and Environment Handling

[`mtplx/runtime_options.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime_options.py) normalizes environment variables and expands abbreviated flags into canonical forms. This file contains `env_bool()` for parsing boolean environment variables, `canonicalize_flag_tokens()` for flag expansion, and `normalize_paged_kv_quantization()` for configuring KV-cache quantization modes.

The file also handles API key management for the local OpenAI-compatible server through `resolve_api_key()` and `generate_api_key_file()`. Configuration parameters flow through this module to ensure every component reads consistent values from environment variables like `MTPLX_SESSION_BLOCK_PREFIX_RESTORE`.

[`mtplx/version.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/version.py) supplies the package metadata through `DISPLAY_VERSION` and `__version__`, providing the user-facing version string reported at startup.

## Batch Scheduling and Adaptive Depth

[`mtplx/batching/scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/scheduler.py) controls how incoming requests are multiplexed using the `SchedulerMode` enum. The scheduler supports three modes: **serial** for single-request processing, **ar-batch** for automatic regression batching, and **hyper** for aggressive parallelization.

Draft depth for speculative decoding is determined dynamically by [`mtplx/adaptive.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/adaptive.py). The `AdaptivePolicy` class decides on-the-fly how many draft steps the MTP decoder should take based on current latency and throughput goals. This adaptive behavior ensures the system maximizes token generation speed without wasting compute on rejected draft tokens.

## Model Backends and Speculative Decoding

The `mtplx/backends/` directory contains the model-specific implementations that execute the forward pass. Each file defines a `Backend` subclass with `forward()` and `prefill()` hooks wired into the speculative MTP path:

- **[`mtplx/backends/qwen3_next.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/qwen3_next.py)** – Implements the Qwen-3 model family with native MTP support.
- **[`mtplx/backends/gemma4_assistant.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/gemma4_assistant.py)** – Handles the Gemma-4 architecture and assistant-mode optimizations.
- **[`mtplx/backends/deepseek_mtp.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/backends/deepseek_mtp.py)** – Provides DeepSeek-V4 implementation with integrated multi-token prediction.

These backends receive the runtime configuration from [`runtime_options.py`](https://github.com/youssofal/MTPLX/blob/main/runtime_options.py) and orchestrate the speculative decoding loop, calling into the kernel layer for actual matrix operations.

## Low-Level Kernel Execution

`mtplx/kernels/` supplies the Metal/MLX kernels that execute heavy matrix work on Apple Silicon. These files contain hardware-optimized implementations for attention mechanisms, position embeddings, and quantized operations:

- **[`mtplx/kernels/sdpa_nax_tile.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/kernels/sdpa_nax_tile.py)** – Scaled dot-product attention with tiled memory access patterns.
- **[`mtplx/kernels/qwen4_m4_rope.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/kernels/qwen4_m4_rope.py)** – Rotary Position Embedding (RoPE) optimized for Qwen4/M4 architectures.
- **[`mtplx/kernels/laguna_decode.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/kernels/laguna_decode.py)** – Quantized KV-cache decoding kernels.

The kernels execute the actual GPU work when invoked by the backend `forward()` methods, ensuring memory-bandwidth-efficient operations on the Neural Engine and GPU.

## Session Bank and Memory Management

[`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py) manages the KV-cache hierarchy through the `SessionBank` class. This component maintains the in-memory cache for active conversations and provides checkpoint/restore capabilities.

For very long contexts, the session bank supports an SSD cold tier. The `block_prefix_restore_enabled` helper checks the `MTPLX_SESSION_BLOCK_PREFIX_RESTORE` environment variable to determine whether to enable disk-based session caching:

```bash
mtplx start --ssd-session-cache=on --ssd-session-cache-dir=~/mtplx_ssd

```

The `SessionBank` constructor accepts parameters like `max_size_gb` to limit in-memory cache size before spilling to SSD.

## System Integration and Monitoring

[`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py) implements `detect_thermal_control()` and the `ThermalControl` helper class to manage fan speeds on Apple-Silicon Macs. This ensures sustained maximum performance by preventing thermal throttling during long inference runs:

```bash

# Override the fan policy for sustained-max performance

mtplx start --max

```

[`mtplx/version.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/version.py) exposes `DISPLAY_VERSION` for user-facing version reporting, allowing the CLI and API server to report consistent version strings.

## Programmatic Runtime Access

You can import and execute the same runtime components that the CLI uses for custom applications:

```python
from mtplx.runtime_options import env_bool, normalize_paged_kv_quantization
from mtplx.batching.scheduler import SchedulerMode
from mtplx.backends.qwen3_next import Qwen3Backend
from mtplx.session_bank import SessionBank

# Example: force a quantised KV cache and serial scheduler

kv_mode   = normalize_paged_kv_quantization("q8")
scheduler = SchedulerMode.SERIAL

# Initialise a backend (model path must point at a downloaded checkpoint)

backend = Qwen3Backend(model_path="~/.mtplx/cache/Qwen3.8-27B")
session = SessionBank(max_size_gb=64)      # in‑memory KV cache

# Perform a single forward pass (ignoring all the CLI plumbing)

output = backend.generate(prompt="Write a haiku", kv_quant=kv_mode, scheduler=scheduler)
print(output)

```

## Summary

- **Entry Point**: [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) parses arguments via `build_parser()` and dispatches to sub-commands.
- **Configuration**: [`mtplx/runtime_options.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime_options.py) normalizes flags through `canonicalize_flag_tokens()` and manages environment variables like `MTPLX_PAGED_KV_QUANT`.
- **Scheduling**: [`mtplx/batching/scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/scheduler.py) selects `SchedulerMode` (serial, ar-batch, hyper) while [`mtplx/adaptive.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/adaptive.py) tunes speculative draft depth.
- **Execution**: `mtplx/backends/*.py` (qwen3_next.py, gemma4_assistant.py, deepseek_mtp.py) implement model-specific `forward()` and `prefill()` methods with MTP support.
- **Kernels**: `mtplx/kernels/*.py` provide hardware-optimized Metal operations for attention, MLP, and quantized KV cache access.
- **Memory**: [`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py) manages the KV-cache hierarchy with optional SSD tiering via `MTPLX_SESSION_BLOCK_PREFIX_RESTORE`.
- **Monitoring**: [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py) implements `detect_thermal_control()` for thermal management on Apple Silicon.

## Frequently Asked Questions

### How does MTPLX handle KV-cache quantization configuration?

MTPLX configures KV-cache quantization through [`mtplx/runtime_options.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime_options.py), specifically via the `normalize_paged_kv_quantization()` function. This function parses user input (such as `"q8"` for 8-bit quantization) and translates it into internal representations used by the kernels in `mtplx/kernels/`.

### What determines the draft depth in MTPLX's speculative decoding?

Draft depth is determined by the `AdaptivePolicy` class in [`mtplx/adaptive.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/adaptive.py), which continuously evaluates latency and throughput metrics to decide how many tokens the MTP decoder should draft speculatively. This adaptive approach balances computational efficiency against acceptance rates.

### Can I use MTPLX programmatically without the CLI?

Yes, you can import the runtime components directly from [`mtplx/runtime_options.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/runtime_options.py), [`mtplx/batching/scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/scheduler.py), and `mtplx/backends/*.py` to initialize backends and schedulers in Python code. This allows integration into existing applications while bypassing the [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) argument parsing entirely.

### How does MTPLX manage thermal throttling on Apple Silicon?

MTPLX monitors system thermals through [`mtplx/thermal.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/thermal.py), which provides `detect_thermal_control()` and the `ThermalControl` helper class. These utilities can adjust fan speeds dynamically to maintain consistent performance during sustained inference workloads, configurable via flags like `--max` for maximum fan speed.