MTPLX Performance Profiles: A Complete Guide to Runtime Optimization
MTPLX provides six distinct runtime profiles—turbo, sustained, performance-cold, stable, exact, and max-diagnostic—each configured via RuntimeProfile objects in mtplx/profiles.py to optimize inference for specific hardware constraints, model quantization levels, and quality-vs-speed trade-offs.
The MTPLX inference engine (available at youssofal/MTPLX) uses runtime profiles to configure the execution pipeline without modifying source code. Each profile defined in [mtplx/profiles.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py) bundles a specific engine backend with a curated set of environment variables that control kernel selection, memory management, and hardware acceleration. Understanding these MTPLX performance profiles allows operators to maximize throughput for quantized flagship models or ensure exact correctness for testing scenarios.
What Are MTPLX Performance Profiles?
At the architectural level, an MTPLX performance profile is a RuntimeProfile dataclass that specifies three critical components:
- Core Engine Selection – The
runtime_profilestring (e.g.,"native_mtp_turbo") selects a concrete execution pipeline in [mtplx/engine_session.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py). - Environment Variable Maps – Each profile ships with a dictionary of default environment variables (the
envfield) that configure kernel paths, chunked pre-fill sizes, and verification modes. - Dynamic Gating Logic – The helper function
announce_runtime_gated_env(lines 62-89 inmtplx/profiles.py) resolves flag dependencies and logs disabled features at launch time.
Profiles are applied via apply_profile_env() (around line 553), which reads the MTPLX_PROFILE environment variable or CLI arguments and injects the appropriate configuration into the process environment.
The Six MTPLX Performance Profiles Explained
turbo
The turbo profile activates native_mtp_turbo, the most aggressive optimization path in the MTPLX engine. According to the source code around line 795, this profile enables NAX-verify kernels and compiled-verify steps for quantized matrix multiplication acceleration.
- Key Environment Variables:
MTPLX_NAX_VERIFY=1,MTPLX_COMPILED_VERIFY=1,MTPLX_GQA_PACKED_SDPA=1,MTPLX_NAX_FLASH_ROUTE=1 - Best For: Quantized flagship models (27B and 9B parameters, such as Qwen 3.8-Optimized-Speed/Quality variants) requiring maximum decode throughput with long-context memory safety.
sustained
The sustained profile (native_mtp_sustained) serves as the product default for non-quantized models. Defined around line 517 via the SUSTAINED_PREFILL_ENV dictionary, it implements chunked pre-fill, final-token logits computation, and repaged KV cache management.
- Key Environment Variables:
MTPLX_SUSTAINED_PREFILL=1,MTPLX_PREFILL_CHUNK_SIZE=auto,MTPLX_LONG_CONTEXT_MTP_DEPTH=3 - Best For: General-purpose long-context generation tasks such as coding assistants and agent workflows, preserving burst-mode TPS while maintaining stability.
performance-cold
The performance-cold profile (native_mtp_60_cold) represents the original "burst" lane that delegates thermal management to the Apple fan controller for maximum short-term speed. The configuration resides in NATIVE_MTP_60_FAST_PATH_ENV (lines 2-30).
- Key Characteristics: Pure fast-path execution without sustained-mode safeguards.
- Best For: Short prompts under 8K tokens where raw throughput matters and fan-boost noise/thermal loads are acceptable.
stable
The stable profile (long_response_exact_staged) provides a conservative fallback path documented in lines 497-514. It combines EXACT_PAGED_ATTENTION_ENV and LONG_RESPONSE_STAGED_ENV variables to disable fan-boost while guaranteeing exactness.
- Characteristics: Exact-staged long-reply path without fan control; slower than
turboorsustainedbut maximizes compatibility. - Best For: Hidden fallback scenarios requiring maximum reliability when primary profiles are unavailable.
exact
The exact profile implements exact paged attention with release gates for QA-only verification. It uses the same environment set as EXACT_PAGED_ATTENTION_ENV (lines 497-505) but disables all throughput optimizations.
- Characteristics: Strict correctness testing with exact paged attention verifier.
- Best For: Testing and validation workflows where mathematical exactness trumps speed; not recommended for production inference.
max-diagnostic
The max-diagnostic profile (max_diagnostic) enables fan-control diagnostics without model-specific optimizations. It merges the environment dictionaries from EXACT_PAGED_ATTENTION_ENV and LONG_RESPONSE_STAGED_ENV.
- Key Function: Enables fan-speed monitoring and clock-anchor experimentation via the
--maxCLI flag. - Best For: Hardware diagnostics and thermal behavior analysis only; provides no inference speed benefits.
How to Select the Right MTPLX Performance Profile
Use this decision matrix to choose the optimal configuration:
| Scenario | Recommended Profile | Rationale |
|---|---|---|
| Quantized 27B/9B flagship models needing fastest decode | turbo |
Activates NAX-verify kernels and compiled 4-bit matmul acceleration |
| General long-context models (non-quantized) | sustained |
Balances chunked pre-fill with safe KV cache handling |
| Short prompts (<8K) where speed is critical | performance-cold |
Maximizes TPS via burst-mode fan control |
| Strict correctness verification required | exact |
Disables all approximations for QA validation |
| Fan-control hardware diagnostics | max-diagnostic |
Enables monitoring without inference optimizations |
| Conservative fallback when primary profiles fail | stable |
Exact-staged path without fan dependencies |
How to Apply a Profile at Runtime
MTPLX supports both programmatic and CLI-based profile activation.
Programmatic Usage
Import the profile utilities from mtplx/profiles to configure the runtime environment before initializing the engine session:
from mtplx.profiles import apply_profile_env, get_profile
# Apply the turbo profile to the current process
apply_profile_env("turbo") # See implementation at line 553
# Verify active configuration
print(get_profile().summary)
# Output: "Sustained Mode plus verify-specialized..."
CLI Usage
Set the MTPLX_PROFILE environment variable or use the --profile flag. The CLI resolves aliases defined in PROFILE_ALIASES (around line 190):
# Explicit profile selection
MTPLX_PROFILE=turbo mtplx start --model qwen3.8-optimized-speed
# Using aliases ("default" resolves to "sustained")
mtplx start --profile default
mtplx start --profile safe # Resolves to "stable"
The --max flag operates orthogonally to profiles, enabling fan-boost diagnostics on top of any selected configuration.
Environment Variables and Dynamic Gating
Each profile defines default values for environment variables, but operators can override specific keys defined in PROFILE_ENV_USER_OVERRIDE_KEYS (lines 25-95). When conflicts occur between profile defaults and user settings, the user values take precedence.
The runtime also implements dynamic gating logic. For example, MTPLX_BATCH_TARGET_ARRAYS automatically disables when MTPLX_LAZY_TARGET_DISTRIBUTIONS=1. The announce_runtime_gated_env function logs these suppressed flags at startup, ensuring transparency about which optimizations are actually active.
Key configuration files referenced by the profile system:
- [
mtplx/profiles.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py) – Profile definitions andapply_profile_env()implementation - [
docs/profiles.md](https://github.com/youssofal/MTPLX/blob/main/docs/profiles.md) – Human-readable documentation of profile behaviors - [
mtplx/engine_session.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py) – Consumesruntime_profilenames to instantiate execution pipelines - [
mtplx/cli.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) – Parses--profileand--maxarguments
Summary
- MTPLX performance profiles are
RuntimeProfileobjects defined inmtplx/profiles.pythat bundle engine selection with environment variable presets. - The
turboprofile provides maximum throughput for quantized 27B/9B models via NAX-verify kernels and compiled verification. sustainedserves as the default for general-purpose models, offering chunked pre-fill and long-context stability.performance-colddelivers burst-mode speed for short prompts using the Apple fan controller.exactandstableprovide conservative, exact-computation paths for QA and fallback scenarios.max-diagnosticenables hardware monitoring without inference optimizations.- Profiles are activated via
apply_profile_env()or the--profileCLI flag, with aliases like"default"(sustained) and"safe"(stable) available for convenience.
Frequently Asked Questions
What is the default MTPLX performance profile?
The default profile is sustained, mapped via the DEFAULT_PROFILE_NAME constant and the "default"/"auto" aliases in PROFILE_ALIASES (around line 190). This profile optimizes for long-context generation across general-purpose models while maintaining stable throughput.
Can I use environment variables to override profile settings?
Yes. While each profile provides default environment variables (such as those in SUSTAINED_PREFILL_ENV for the sustained profile), operators can override specific keys listed in PROFILE_ENV_USER_OVERRIDE_KEYS (lines 25-95). User-defined values take precedence over profile defaults when both are present.
When should I use turbo instead of sustained?
Use turbo when running quantized flagship models (specifically 27B or 9B parameter variants like Qwen 3.8-Optimized) where you need the fastest possible decode speed and can utilize NAX-verify kernels. Use sustained for non-quantized models or when maximum hardware-specific optimization is less critical than general long-context stability.
How does the --max flag interact with performance profiles?
The --max flag is orthogonal to profile selection. It enables fan-boost and diagnostic monitoring (activating the max-diagnostic behaviors) regardless of which base profile is active. For example, mtplx start --profile sustained --max runs the sustained engine with maximum fan speed enabled for short bursts, while mtplx start --max alone defaults to sustained mode with diagnostics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →