# What Is the performance-cold --max Profile in MTPLX?

> Discover the performance-cold --max profile in MTPLX. This high-speed burst mode maximizes throughput for benchmarks and short prompts by running system fans at maximum speed.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: performance
- Published: 2026-09-02

---

**The `performance-cold --max` profile is a high-speed, short-context execution mode (labeled "Burst" in the UI) that pins system fans at maximum speed to deliver maximum throughput for brief benchmark runs and small prompts.**

The MTPLX inference engine provides specialized execution profiles optimized for different thermal and context-length scenarios. The `performance-cold --max` profile represents the most aggressive configuration, designed specifically for scenarios requiring absolute maximum throughput without thermal throttling concerns, though it carries strict limitations regarding context window size.

## Technical Overview of the performance-cold --max Profile

The `performance-cold --max` configuration activates what the codebase internally refers to as **Burst mode**. According to the profile documentation in [`docs/profiles.md`](https://github.com/youssofal/MTPLX/blob/main/docs/profiles.md) (lines 8-10), this setting runs the "old max-fan performance-cold lane," which physically pins system cooling fans at their maximum RPM for the duration of the inference run.

This aggressive cooling strategy allows the engine to maintain peak clock speeds without thermal throttling, but it creates significant mechanical and thermal constraints. The documentation explicitly warns that this mode **is not safe for long contexts** and recommends avoiding usage beyond approximately **8 KB of context** (roughly 2,000 tokens).

## Source Code Implementation

The implementation spans multiple critical files within the `youssofal/MTPLX` repository:

- **[`docs/profiles.md`](https://github.com/youssofal/MTPLX/blob/main/docs/profiles.md)** (lines 8-10): Defines the profile table entry marking this configuration as "Burst: old max-fan headline lane, not recommended beyond 8K context."

- **[`mtplx/ui/onboarding.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/ui/onboarding.py)** (lines 38-40): Contains the UI description labeling it as the "Old max-fan performance-cold lane" and restricting it to "Fastest headline burst for short prompts and benchmarks only; avoid for long documents or coding contexts."

- **[`mtplx/profiles.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py)** (lines 16-22): Declares the `"performance-cold"` profile name within `PROFILE_CHOICES` and handles internal profile validation.

- **[`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py)** (lines 1527-1529): Implements the fan-boost logic that engages maximum cooling when the `performance-cold` profile receives the `--max` modifier.

## When to Use (and Avoid) performance-cold --max

Use this profile exclusively for:

- **Short prompts** requiring minimal context (under 8 KB)
- **Benchmark runs** where maximum tokens-per-second is the priority
- **Headline generation** or single-turn queries without conversation history

Avoid this profile for:

- **Long-form document processing**
- **Code generation** with extensive file contexts
- **Multi-turn conversations** where accumulated context exceeds the 8 KB threshold
- **Sustained production workloads** due to maximum fan noise and mechanical wear

The profile remains hidden from MTPLX's default onboarding flow and requires explicit selection via command-line flags or API parameters.

## How to Enable performance-cold --max

### Command-Line Interface

Invoke the profile using the `--profile` and `--max` flags:

```bash

# Run a quick benchmark with the Burst profile (max-fan)

mtplx start --profile performance-cold --max --prompt "Explain quantum entanglement in 100 words."

```

### Python API

Access the profile programmatically through the MTPLX client:

```python
import mtplx

client = mtplx.Client()
response = client.chat(
    prompt="Explain quantum entanglement in 100 words.",
    profile="performance-cold",
    max=True,          # enables the max-fan mode

)
print(response.text)

```

## Summary

- The `performance-cold --max` profile activates **Burst mode**, pinning cooling fans at maximum RPM.
- According to [`docs/profiles.md`](https://github.com/youssofal/MTPLX/blob/main/docs/profiles.md), the safe context limit is approximately **8 KB**; exceeding this risks thermal throttling or hardware stress.
- The profile is **excluded from default onboarding** and must be explicitly enabled via `--max` flag or API parameter.
- Best suited for **short prompts and benchmarking**, not sustained or long-context workloads.
- Implementation resides in [`mtplx/profiles.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/profiles.py) (profile definition) and [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (fan-boost logic).

## Frequently Asked Questions

### What exactly does the `--max` flag do when combined with `performance-cold`?

The `--max` flag invokes the "Burst" lane within the performance-cold profile, commanding the system to maintain cooling fans at absolute maximum speed throughout the inference job. As implemented in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (lines 1527-1529), this prevents thermal throttling during peak computational load but generates significant noise and mechanical stress.

### Why is the performance-cold --max profile limited to 8 KB of context?

The 8 KB restriction exists because maximum fan speed provides finite cooling capacity. According to the documentation in [`docs/profiles.md`](https://github.com/youssofal/MTPLX/blob/main/docs/profiles.md), sustained inference on contexts larger than 8KB generates heat faster than the maximum airflow can dissipate, risking thermal throttling or hardware degradation during long-generation tasks.

### How does Burst mode differ from standard performance-cold without --max?

Standard `performance-cold` operates with dynamic fan curves that adjust to thermal loads, supporting longer contexts and sustained operation. The `--max` variant (Burst mode) locks fans at 100% speed immediately, optimized exclusively for short, high-intensity bursts where rapid heat dissipation is prioritized over acoustic comfort or mechanical longevity.

### Is the performance-cold --max profile suitable for production APIs?

No. The [`mtplx/ui/onboarding.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/ui/onboarding.py) source (lines 38-40) explicitly restricts this profile to "benchmarks only" and warns against using it for "long documents or coding contexts." Production deployments should use standard thermal profiles that balance throughput with sustainable cooling and acoustic levels.