# MTPLX Performance Profiles: A Complete Guide to Runtime Optimization

> Explore MTPLX performance profiles like turbo sustained and stable to optimize inference for your hardware and model. Choose the right profile for quality and speed.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: performance
- Published: 2026-09-13

---

**MTPLX provides six distinct runtime profiles—`turbo`, `sustained`, `performance-cold`, `stable`, `exact`, and `max-diagnostic`—each configured via `RuntimeProfile` objects in [`mtplx/profiles.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py) to optimize inference for specific hardware constraints, model quantization levels, and quality-vs-speed trade-offs.**

The **MTPLX** inference engine (available at `youssofal/MTPLX`) uses runtime profiles to configure the execution pipeline without modifying source code. Each profile defined in [[`mtplx/profiles.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py) bundles a specific engine backend with a curated set of environment variables that control kernel selection, memory management, and hardware acceleration. Understanding these MTPLX performance profiles allows operators to maximize throughput for quantized flagship models or ensure exact correctness for testing scenarios.

## What Are MTPLX Performance Profiles?

At the architectural level, an MTPLX performance profile is a `RuntimeProfile` dataclass that specifies three critical components:

1. **Core Engine Selection** – The `runtime_profile` string (e.g., `"native_mtp_turbo"`) selects a concrete execution pipeline in [[`mtplx/engine_session.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py).
2. **Environment Variable Maps** – Each profile ships with a dictionary of default environment variables (the `env` field) that configure kernel paths, chunked pre-fill sizes, and verification modes.
3. **Dynamic Gating Logic** – The helper function `announce_runtime_gated_env` (lines 62-89 in [`mtplx/profiles.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py)) resolves flag dependencies and logs disabled features at launch time.

Profiles are applied via `apply_profile_env()` (around line 553), which reads the `MTPLX_PROFILE` environment variable or CLI arguments and injects the appropriate configuration into the process environment.

## The Six MTPLX Performance Profiles Explained

### turbo

The **`turbo`** profile activates `native_mtp_turbo`, the most aggressive optimization path in the MTPLX engine. According to the source code around line 795, this profile enables **NAX-verify kernels** and **compiled-verify steps** for quantized matrix multiplication acceleration.

- **Key Environment Variables**: `MTPLX_NAX_VERIFY=1`, `MTPLX_COMPILED_VERIFY=1`, `MTPLX_GQA_PACKED_SDPA=1`, `MTPLX_NAX_FLASH_ROUTE=1`
- **Best For**: Quantized flagship models (27B and 9B parameters, such as Qwen 3.8-Optimized-Speed/Quality variants) requiring maximum decode throughput with long-context memory safety.

### sustained

The **`sustained`** profile (`native_mtp_sustained`) serves as the **product default** for non-quantized models. Defined around line 517 via the `SUSTAINED_PREFILL_ENV` dictionary, it implements chunked pre-fill, final-token logits computation, and repaged KV cache management.

- **Key Environment Variables**: `MTPLX_SUSTAINED_PREFILL=1`, `MTPLX_PREFILL_CHUNK_SIZE=auto`, `MTPLX_LONG_CONTEXT_MTP_DEPTH=3`
- **Best For**: General-purpose long-context generation tasks such as coding assistants and agent workflows, preserving burst-mode TPS while maintaining stability.

### performance-cold

The **`performance-cold`** profile (`native_mtp_60_cold`) represents the original "burst" lane that delegates thermal management to the Apple fan controller for maximum short-term speed. The configuration resides in `NATIVE_MTP_60_FAST_PATH_ENV` (lines 2-30).

- **Key Characteristics**: Pure fast-path execution without sustained-mode safeguards.
- **Best For**: Short prompts under 8K tokens where raw throughput matters and fan-boost noise/thermal loads are acceptable.

### stable

The **`stable`** profile (`long_response_exact_staged`) provides a conservative fallback path documented in lines 497-514. It combines `EXACT_PAGED_ATTENTION_ENV` and `LONG_RESPONSE_STAGED_ENV` variables to disable fan-boost while guaranteeing exactness.

- **Characteristics**: Exact-staged long-reply path without fan control; slower than `turbo` or `sustained` but maximizes compatibility.
- **Best For**: Hidden fallback scenarios requiring maximum reliability when primary profiles are unavailable.

### exact

The **`exact`** profile implements `exact` paged attention with release gates for **QA-only verification**. It uses the same environment set as `EXACT_PAGED_ATTENTION_ENV` (lines 497-505) but disables all throughput optimizations.

- **Characteristics**: Strict correctness testing with exact paged attention verifier.
- **Best For**: Testing and validation workflows where mathematical exactness trumps speed; not recommended for production inference.

### max-diagnostic

The **`max-diagnostic`** profile (`max_diagnostic`) enables fan-control diagnostics without model-specific optimizations. It merges the environment dictionaries from `EXACT_PAGED_ATTENTION_ENV` and `LONG_RESPONSE_STAGED_ENV`.

- **Key Function**: Enables fan-speed monitoring and clock-anchor experimentation via the `--max` CLI flag.
- **Best For**: Hardware diagnostics and thermal behavior analysis only; provides no inference speed benefits.

## How to Select the Right MTPLX Performance Profile

Use this decision matrix to choose the optimal configuration:

| Scenario | Recommended Profile | Rationale |
|----------|-------------------|-----------|
| Quantized 27B/9B flagship models needing fastest decode | `turbo` | Activates NAX-verify kernels and compiled 4-bit matmul acceleration |
| General long-context models (non-quantized) | `sustained` | Balances chunked pre-fill with safe KV cache handling |
| Short prompts (<8K) where speed is critical | `performance-cold` | Maximizes TPS via burst-mode fan control |
| Strict correctness verification required | `exact` | Disables all approximations for QA validation |
| Fan-control hardware diagnostics | `max-diagnostic` | Enables monitoring without inference optimizations |
| Conservative fallback when primary profiles fail | `stable` | Exact-staged path without fan dependencies |

## How to Apply a Profile at Runtime

MTPLX supports both programmatic and CLI-based profile activation.

### Programmatic Usage

Import the profile utilities from `mtplx/profiles` to configure the runtime environment before initializing the engine session:

```python
from mtplx.profiles import apply_profile_env, get_profile

# Apply the turbo profile to the current process

apply_profile_env("turbo")  # See implementation at line 553

# Verify active configuration

print(get_profile().summary)

# Output: "Sustained Mode plus verify-specialized..."

```

### CLI Usage

Set the `MTPLX_PROFILE` environment variable or use the `--profile` flag. The CLI resolves aliases defined in `PROFILE_ALIASES` (around line 190):

```bash

# Explicit profile selection

MTPLX_PROFILE=turbo mtplx start --model qwen3.8-optimized-speed

# Using aliases ("default" resolves to "sustained")

mtplx start --profile default
mtplx start --profile safe  # Resolves to "stable"

```

The `--max` flag operates orthogonally to profiles, enabling fan-boost diagnostics on top of any selected configuration.

## Environment Variables and Dynamic Gating

Each profile defines default values for environment variables, but operators can override specific keys defined in `PROFILE_ENV_USER_OVERRIDE_KEYS` (lines 25-95). When conflicts occur between profile defaults and user settings, the user values take precedence.

The runtime also implements **dynamic gating** logic. For example, `MTPLX_BATCH_TARGET_ARRAYS` automatically disables when `MTPLX_LAZY_TARGET_DISTRIBUTIONS=1`. The `announce_runtime_gated_env` function logs these suppressed flags at startup, ensuring transparency about which optimizations are actually active.

Key configuration files referenced by the profile system:
- **[[`mtplx/profiles.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py)** – Profile definitions and `apply_profile_env()` implementation
- **[[`docs/profiles.md`](https://github.com/youssofal/MTPLX/blob/main/docs/profiles.md)](https://github.com/youssofal/MTPLX/blob/main/docs/profiles.md)** – Human-readable documentation of profile behaviors
- **[[`mtplx/engine_session.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py)** – Consumes `runtime_profile` names to instantiate execution pipelines
- **[[`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py)** – Parses `--profile` and `--max` arguments

## Summary

- **MTPLX performance profiles** are `RuntimeProfile` objects defined in [`mtplx/profiles.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py) that bundle engine selection with environment variable presets.
- The **`turbo`** profile provides maximum throughput for quantized 27B/9B models via NAX-verify kernels and compiled verification.
- **`sustained`** serves as the default for general-purpose models, offering chunked pre-fill and long-context stability.
- **`performance-cold`** delivers burst-mode speed for short prompts using the Apple fan controller.
- **`exact`** and **`stable`** provide conservative, exact-computation paths for QA and fallback scenarios.
- **`max-diagnostic`** enables hardware monitoring without inference optimizations.
- Profiles are activated via `apply_profile_env()` or the `--profile` CLI flag, with aliases like `"default"` (sustained) and `"safe"` (stable) available for convenience.

## Frequently Asked Questions

### What is the default MTPLX performance profile?

The default profile is **`sustained`**, mapped via the `DEFAULT_PROFILE_NAME` constant and the `"default"`/`"auto"` aliases in `PROFILE_ALIASES` (around line 190). This profile optimizes for long-context generation across general-purpose models while maintaining stable throughput.

### Can I use environment variables to override profile settings?

Yes. While each profile provides default environment variables (such as those in `SUSTAINED_PREFILL_ENV` for the sustained profile), operators can override specific keys listed in `PROFILE_ENV_USER_OVERRIDE_KEYS` (lines 25-95). User-defined values take precedence over profile defaults when both are present.

### When should I use `turbo` instead of `sustained`?

Use **`turbo`** when running quantized flagship models (specifically 27B or 9B parameter variants like Qwen 3.8-Optimized) where you need the fastest possible decode speed and can utilize NAX-verify kernels. Use **`sustained`** for non-quantized models or when maximum hardware-specific optimization is less critical than general long-context stability.

### How does the `--max` flag interact with performance profiles?

The **`--max` flag is orthogonal** to profile selection. It enables fan-boost and diagnostic monitoring (activating the `max-diagnostic` behaviors) regardless of which base profile is active. For example, `mtplx start --profile sustained --max` runs the sustained engine with maximum fan speed enabled for short bursts, while `mtplx start --max` alone defaults to sustained mode with diagnostics.