Understanding MTPLX Execution Profiles: When to Use Each Runtime Configuration

MTPLX execution profiles are pre-configured runtime environments defined in mtplx/profiles.py that map friendly names like "turbo" or "sustained" to specific environment variable sets, controlling performance, memory usage, and fan behavior for different inference scenarios.

MTPLX ships with a sophisticated profile system that allows developers to switch between optimized runtime configurations without manual environment tuning. These execution profiles encode predetermined environment variables for various performance and thermal scenarios, from high-throughput decoding to conservative memory-safe operation. Understanding when to apply each profile ensures optimal token-per-second throughput while respecting hardware thermal constraints.

What Are MTPLX Execution Profiles?

At their core, MTPLX execution profiles are concrete RuntimeProfile objects stored in the PROFILES dictionary within mtplx/profiles.py (lines 33‑40). Each profile maps a human-readable name to a configuration that includes the internal generation engine name, descriptive summary, and a tuple of environment variable key-value pairs injected at process launch.

When the CLI initializes, it invokes resolve_profile_name() to normalize user input, map aliases (for example, "default" → "sustained" or "safe" → "stable"), and validate against available profiles (lines 336‑342). The selected profile then passes to apply_profile_env() (lines 553‑560), which merges the profile’s environment with any user-provided overrides defined in PROFILE_ENV_USER_OVERRIDE_KEYS.

Available Execution Profiles

The following execution profiles are defined in the source code and summarized in docs/profiles.md:

  • turbo — Default for flagship quantized models such as Qwen 3.8 Optimized-Speed/Quality and 9B Speed variants. This profile enables sustained-mode base with NAX verify kernels and compiled-verify for 4-bit models, delivering the fastest decode throughput. Choose this when you need maximum tokens-per-second on supported hardware, especially for long response generation.

  • sustained — General-purpose profile for long-context workloads. It activates the native-MTP long-context path with chunked pre-fill, final-token logits, and paged KV caching, with the Apple fan controller enabled. This is the preferred choice for most coding or chat scenarios requiring memory safety and moderate throughput.

  • sustained + --max — Identical to the standard sustained profile but adds ThermalForge/TG Pro fan-pinning capabilities. Use this when you need extra thermal headroom during heavy generation and your hardware supports fan control.

  • performance-cold + --max — "Burst" mode designed for short-context workloads (≤ 8K tokens). This configuration delivers the highest raw TPS but should only be used where speed outweighs fan-control considerations.

  • performance-cold — Legacy burst mode without fan boost. This profile is maintained for backward compatibility but hidden from onboarding flows. Use only when legacy scripts explicitly request it.

  • stable — Conservative fallback implementing exact attention with staged processing for deterministic long-reply paths. This profile is hidden from first-run onboarding but useful when you need reproducible, safe operation.

  • exact — QA-only verification mode employing exact-paged verifiers with release gates enabled. Employ this when strict verification of model outputs is required for testing or debugging.

  • max-diagnostic — Diagnostic fan-control profile enabling fan control and clock-anchor features. This is intended solely for diagnosing fan-control behavior and is not recommended for production inference.

Programmatic Profile Management

You can interact with execution profiles directly in Python scripts to inspect configurations or manage environments:


# List all available profiles with summaries

from mtplx.profiles import list_profiles

for p in list_profiles():
    print(f"{p['name']}: {p['summary']}")

# Retrieve and inspect a specific profile

from mtplx.profiles import get_profile

profile = get_profile("turbo")
print(profile.runtime_profile)          # → native_mtp_turbo

print(profile.env_dict()["MTPLX_NAX_VERIFY"])  # → "1"

# Apply a profile temporarily and restore the original environment

from mtplx.profiles import apply_profile_env, restore_profile_env

previous_env = apply_profile_env("sustained")

# ... run inference ...

restore_profile_env(previous_env)

# Resolve aliases programmatically

from mtplx.profiles import resolve_profile_name

assert resolve_profile_name("default") == "sustained"
assert resolve_profile_name("safe") == "stable"

Summary

  • MTPLX execution profiles are predefined RuntimeProfile objects in mtplx/profiles.py that bundle environment variables for specific performance scenarios.
  • turbo delivers maximum throughput for quantized models, while sustained balances performance with thermal safety for general workloads.
  • The --max flag augments compatible profiles with aggressive fan control for thermal headroom during intensive tasks.
  • resolve_profile_name() handles alias normalization and validation, while apply_profile_env() injects configuration at runtime.
  • Programmatic access via get_profile() and list_profiles() enables dynamic inspection and environment manipulation in custom scripts.

Frequently Asked Questions

What is the difference between the turbo and sustained execution profiles?

The turbo profile targets flagship quantized models with NAX verify kernels and compiled-verify for 4-bit implementations, prioritizing raw decode speed. In contrast, sustained activates chunked pre-fill, paged KV caching, and native-MTP long-context paths, trading absolute speed for memory safety and moderate throughput during extended generation sessions.

When should I use performance-cold instead of sustained?

Choose performance-cold (preferably with --max) exclusively for short-context workloads under 8K tokens where you require the highest possible raw TPS and can tolerate louder fan operation. For contexts exceeding 8K tokens or when fan noise and thermal stability matter, use sustained or sustained --max instead.

How do I temporarily apply a profile in a Python script without affecting the global environment?

Import apply_profile_env() and restore_profile_env() from mtplx/profiles. Call apply_profile_env("profile_name") to capture and return the previous environment state, run your inference code, then pass the returned state to restore_profile_env() to revert changes. This pattern ensures clean isolation between different execution contexts.

What is the purpose of the max-diagnostic profile?

The max-diagnostic profile enables diagnostic fan-control behavior and clock-anchor features specifically for troubleshooting thermal management issues. According to the source code in mtplx/profiles.py, this profile is intended for diagnosing hardware fan-control behavior and should never be used for production inference workloads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →