MTPLX Performance Profiles: A Complete Guide to Runtime Optimization

MTPLX provides six distinct runtime profiles—turbo, sustained, performance-cold, stable, exact, and max-diagnostic—each configured via RuntimeProfile objects in mtplx/profiles.py to optimize inference for specific hardware constraints, model quantization levels, and quality-vs-speed trade-offs.

The MTPLX inference engine (available at youssofal/MTPLX) uses runtime profiles to configure the execution pipeline without modifying source code. Each profile defined in [mtplx/profiles.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py) bundles a specific engine backend with a curated set of environment variables that control kernel selection, memory management, and hardware acceleration. Understanding these MTPLX performance profiles allows operators to maximize throughput for quantized flagship models or ensure exact correctness for testing scenarios.

What Are MTPLX Performance Profiles?

At the architectural level, an MTPLX performance profile is a RuntimeProfile dataclass that specifies three critical components:

  1. Core Engine Selection – The runtime_profile string (e.g., "native_mtp_turbo") selects a concrete execution pipeline in [mtplx/engine_session.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py).
  2. Environment Variable Maps – Each profile ships with a dictionary of default environment variables (the env field) that configure kernel paths, chunked pre-fill sizes, and verification modes.
  3. Dynamic Gating Logic – The helper function announce_runtime_gated_env (lines 62-89 in mtplx/profiles.py) resolves flag dependencies and logs disabled features at launch time.

Profiles are applied via apply_profile_env() (around line 553), which reads the MTPLX_PROFILE environment variable or CLI arguments and injects the appropriate configuration into the process environment.

The Six MTPLX Performance Profiles Explained

turbo

The turbo profile activates native_mtp_turbo, the most aggressive optimization path in the MTPLX engine. According to the source code around line 795, this profile enables NAX-verify kernels and compiled-verify steps for quantized matrix multiplication acceleration.

  • Key Environment Variables: MTPLX_NAX_VERIFY=1, MTPLX_COMPILED_VERIFY=1, MTPLX_GQA_PACKED_SDPA=1, MTPLX_NAX_FLASH_ROUTE=1
  • Best For: Quantized flagship models (27B and 9B parameters, such as Qwen 3.8-Optimized-Speed/Quality variants) requiring maximum decode throughput with long-context memory safety.

sustained

The sustained profile (native_mtp_sustained) serves as the product default for non-quantized models. Defined around line 517 via the SUSTAINED_PREFILL_ENV dictionary, it implements chunked pre-fill, final-token logits computation, and repaged KV cache management.

  • Key Environment Variables: MTPLX_SUSTAINED_PREFILL=1, MTPLX_PREFILL_CHUNK_SIZE=auto, MTPLX_LONG_CONTEXT_MTP_DEPTH=3
  • Best For: General-purpose long-context generation tasks such as coding assistants and agent workflows, preserving burst-mode TPS while maintaining stability.

performance-cold

The performance-cold profile (native_mtp_60_cold) represents the original "burst" lane that delegates thermal management to the Apple fan controller for maximum short-term speed. The configuration resides in NATIVE_MTP_60_FAST_PATH_ENV (lines 2-30).

  • Key Characteristics: Pure fast-path execution without sustained-mode safeguards.
  • Best For: Short prompts under 8K tokens where raw throughput matters and fan-boost noise/thermal loads are acceptable.

stable

The stable profile (long_response_exact_staged) provides a conservative fallback path documented in lines 497-514. It combines EXACT_PAGED_ATTENTION_ENV and LONG_RESPONSE_STAGED_ENV variables to disable fan-boost while guaranteeing exactness.

  • Characteristics: Exact-staged long-reply path without fan control; slower than turbo or sustained but maximizes compatibility.
  • Best For: Hidden fallback scenarios requiring maximum reliability when primary profiles are unavailable.

exact

The exact profile implements exact paged attention with release gates for QA-only verification. It uses the same environment set as EXACT_PAGED_ATTENTION_ENV (lines 497-505) but disables all throughput optimizations.

  • Characteristics: Strict correctness testing with exact paged attention verifier.
  • Best For: Testing and validation workflows where mathematical exactness trumps speed; not recommended for production inference.

max-diagnostic

The max-diagnostic profile (max_diagnostic) enables fan-control diagnostics without model-specific optimizations. It merges the environment dictionaries from EXACT_PAGED_ATTENTION_ENV and LONG_RESPONSE_STAGED_ENV.

  • Key Function: Enables fan-speed monitoring and clock-anchor experimentation via the --max CLI flag.
  • Best For: Hardware diagnostics and thermal behavior analysis only; provides no inference speed benefits.

How to Select the Right MTPLX Performance Profile

Use this decision matrix to choose the optimal configuration:

Scenario Recommended Profile Rationale
Quantized 27B/9B flagship models needing fastest decode turbo Activates NAX-verify kernels and compiled 4-bit matmul acceleration
General long-context models (non-quantized) sustained Balances chunked pre-fill with safe KV cache handling
Short prompts (<8K) where speed is critical performance-cold Maximizes TPS via burst-mode fan control
Strict correctness verification required exact Disables all approximations for QA validation
Fan-control hardware diagnostics max-diagnostic Enables monitoring without inference optimizations
Conservative fallback when primary profiles fail stable Exact-staged path without fan dependencies

How to Apply a Profile at Runtime

MTPLX supports both programmatic and CLI-based profile activation.

Programmatic Usage

Import the profile utilities from mtplx/profiles to configure the runtime environment before initializing the engine session:

from mtplx.profiles import apply_profile_env, get_profile

# Apply the turbo profile to the current process

apply_profile_env("turbo")  # See implementation at line 553

# Verify active configuration

print(get_profile().summary)

# Output: "Sustained Mode plus verify-specialized..."

CLI Usage

Set the MTPLX_PROFILE environment variable or use the --profile flag. The CLI resolves aliases defined in PROFILE_ALIASES (around line 190):


# Explicit profile selection

MTPLX_PROFILE=turbo mtplx start --model qwen3.8-optimized-speed

# Using aliases ("default" resolves to "sustained")

mtplx start --profile default
mtplx start --profile safe  # Resolves to "stable"

The --max flag operates orthogonally to profiles, enabling fan-boost diagnostics on top of any selected configuration.

Environment Variables and Dynamic Gating

Each profile defines default values for environment variables, but operators can override specific keys defined in PROFILE_ENV_USER_OVERRIDE_KEYS (lines 25-95). When conflicts occur between profile defaults and user settings, the user values take precedence.

The runtime also implements dynamic gating logic. For example, MTPLX_BATCH_TARGET_ARRAYS automatically disables when MTPLX_LAZY_TARGET_DISTRIBUTIONS=1. The announce_runtime_gated_env function logs these suppressed flags at startup, ensuring transparency about which optimizations are actually active.

Key configuration files referenced by the profile system:

Summary

  • MTPLX performance profiles are RuntimeProfile objects defined in mtplx/profiles.py that bundle engine selection with environment variable presets.
  • The turbo profile provides maximum throughput for quantized 27B/9B models via NAX-verify kernels and compiled verification.
  • sustained serves as the default for general-purpose models, offering chunked pre-fill and long-context stability.
  • performance-cold delivers burst-mode speed for short prompts using the Apple fan controller.
  • exact and stable provide conservative, exact-computation paths for QA and fallback scenarios.
  • max-diagnostic enables hardware monitoring without inference optimizations.
  • Profiles are activated via apply_profile_env() or the --profile CLI flag, with aliases like "default" (sustained) and "safe" (stable) available for convenience.

Frequently Asked Questions

What is the default MTPLX performance profile?

The default profile is sustained, mapped via the DEFAULT_PROFILE_NAME constant and the "default"/"auto" aliases in PROFILE_ALIASES (around line 190). This profile optimizes for long-context generation across general-purpose models while maintaining stable throughput.

Can I use environment variables to override profile settings?

Yes. While each profile provides default environment variables (such as those in SUSTAINED_PREFILL_ENV for the sustained profile), operators can override specific keys listed in PROFILE_ENV_USER_OVERRIDE_KEYS (lines 25-95). User-defined values take precedence over profile defaults when both are present.

When should I use turbo instead of sustained?

Use turbo when running quantized flagship models (specifically 27B or 9B parameter variants like Qwen 3.8-Optimized) where you need the fastest possible decode speed and can utilize NAX-verify kernels. Use sustained for non-quantized models or when maximum hardware-specific optimization is less critical than general long-context stability.

How does the --max flag interact with performance profiles?

The --max flag is orthogonal to profile selection. It enables fan-boost and diagnostic monitoring (activating the max-diagnostic behaviors) regardless of which base profile is active. For example, mtplx start --profile sustained --max runs the sustained engine with maximum fan speed enabled for short bursts, while mtplx start --max alone defaults to sustained mode with diagnostics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →