# How Temperature and top_p Settings Impact MTPLX's MTP Exactness

> Discover how temperature and top_p settings affect MTPLX's MTP exactness. Learn how these parameters influence token probability without compromising mathematical guarantees.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: performance
- Published: 2026-09-04

---

**Temperature and top_p settings shape the target probability distribution used by MTPLX's MTP module, but they do not violate the mathematical exactness guarantee—the marginal distribution of accepted tokens always matches the filtered target distribution defined by those parameters.**

MTPLX is a speculative decoding framework that accelerates inference through its Memory-Token-Projection (MTP) module. The impact of temperature and top_p settings on MTPLX's MTP exactness depends on how these parameters transform the raw logits into the effective target distribution that the acceptance mechanism evaluates against the draft model's output.

## How MTPLX Constructs the Target Distribution

The MTP exactness guarantee assumes the target distribution is created through a two-stage pipeline in [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py). This distribution represents what the full model would sample, and MTP ensures the final output matches it exactly—even when using aggressive filtering.

### Temperature Scaling in Softmax

Temperature scaling occurs at lines 80-86 in [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py), where the `softmax` function divides logits by the `temperature` parameter before exponentiation:

```python
softmax(logits, temperature=config.temperature)

```

- **Lower temperature** (< 1) sharpens the distribution, concentrating probability mass on the highest-logit tokens and moving toward deterministic output.
- **Higher temperature** (> 1) flattens the distribution, increasing entropy and giving weight to lower-probability tokens.

### Top-p and Top-k Filtering

After temperature scaling, `apply_top_p_top_k` (lines 15-24 in [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py)) truncates the distribution:

```python
apply_top_p_top_k(probs, top_p=config.top_p, top_k=config.top_k)

```

- **Top-p** (nucleus filtering) retains only the highest-probability tokens whose cumulative probability reaches the `top_p` threshold, then renormalizes the remaining mass.
- **Top-k** restricts the candidate set to the `k` highest-probability tokens regardless of cumulative probability.

Together, these stages define the **effective target distribution** that MTP treats as ground truth.

## The Acceptance Mechanism and Exactness Proof

MTP maintains exactness through a rejection sampling scheme implemented in `acceptance_probability` (lines 44-49 in [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py)). The algorithm compares the target probability `p` against the draft probability `q` for each proposed token:

```python
p = _probability(target_p, token_id)
q = _probability(draft_q, token_id)
if q <= 0:
    return 1.0 if p > 0 else 0.0
return min(1.0, p / q)

```

When a draft token is rejected, MTP falls back to `residual_distribution`, which samples from the renormalized remainder of the target distribution. The `speculative_output_marginal` function contains the formal proof that this acceptance/rejection cycle preserves the exact marginal distribution of the target—**including all temperature and top_p modifications applied beforehand**.

## Practical Impact on MTP Behavior

While exactness is mathematically preserved regardless of settings, the practical behavior of speculative decoding changes significantly based on distribution shape:

### Low Temperature (→ 0) and Deterministic Output

When temperature approaches zero, the target distribution collapses to a one-hot vector where the top token has probability 1.0. In [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py), this means:
- Acceptance probability becomes 1.0 for the top token (if the draft proposes it) or 0.0 (if not)
- The residual distribution trivially returns the top token
- Exactness holds but speculative decoding gains minimal speedup since rejections are absolute

### High Temperature with Full Top-p

With temperature > 1.0 and top_p = 1.0, the distribution spreads across many tokens:
- The draft model's approximations more frequently mismatch the target distribution
- `acceptance_probability` returns values < 1.0 more often, triggering the residual path
- Throughput decreases but the output remains mathematically exact per the proof in `speculative_output_marginal`

### Aggressive Top-p Truncation (e.g., 0.5)

Setting `top_p=0.5` removes the low-probability tail before MTP evaluation:
- Exactness is guaranteed **with respect to the truncated distribution**, not the original full distribution
- The acceptance calculation in [`mtplx/mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_patch.py) uses this filtered distribution as its target reference
- This effectively changes the "ground truth" the model commits to matching

## Configuration Examples

These examples from [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py) demonstrate how `SamplerConfig` parameters alter the effective target distribution while preserving MTP exactness:

```python

# Example 1: Deterministic output (greedy decoding)

cfg = SamplerConfig(temperature=0.0, top_p=1.0, top_k=0)
probs = distribution_from_logits(logits, cfg)

# Target distribution is one-hot; MTP acceptance always succeeds for the top token

```

```python

# Example 2: Standard sampling with nucleus filtering

cfg = SamplerConfig(temperature=0.8, top_p=0.95, top_k=0)
probs = distribution_from_logits(logits, cfg)

# Softened distribution with rare tokens truncated; 

# rejections trigger residual_distribution to maintain exact marginal

```

```python

# Example 3: Aggressive truncation with moderate randomness

cfg = SamplerConfig(temperature=1.0, top_p=0.5, top_k=0)
probs = distribution_from_logits(logits, cfg)

# Exactness holds for this heavily filtered distribution only

```

## Summary

- **Temperature and top_p modify the target distribution** in [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py) before MTP evaluation, but do not break the exactness guarantee.
- **Exactness is distribution-relative**: MTP guarantees the output marginal matches whatever distribution results from your temperature, top_p, and top_k settings—not necessarily the untruncated base distribution.
- **Low temperature** reduces rejection rates but limits diversity; **high temperature** increases rejections but preserves exactness through `residual_distribution`.
- **Top-p < 1** establishes a new "ground truth" by truncating tails; MTP's proof in `speculative_output_marginal` applies to this truncated reference.

## Frequently Asked Questions

### Does increasing temperature break MTP exactness?

No. Higher temperature flattens the probability distribution in `softmax`, but the `acceptance_probability` function in [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py) still ensures the final sequence follows that flattened distribution exactly. The proof in `speculative_output_marginal` holds for any valid probability distribution, regardless of entropy.

### What happens when top_p is set to 1.0 versus 0.9?

At `top_p=1.0`, MTP targets the full distribution produced after temperature scaling. At `top_p=0.9`, `apply_top_p_top_k` removes the bottom 10% of probability mass and renormalizes. MTP then guarantees exactness relative to this truncated 90% nucleus distribution, not the original full distribution.

### How does the residual distribution maintain exactness after rejections?

When the draft token is rejected (probability `min(1.0, p/q)` < 1), `residual_distribution` samples from the remaining probability mass of the target distribution, excluding the accepted portion. This rejection sampling technique ensures the marginal distribution of the accepted token—or the fallback sample—exactly equals the target distribution, as formalized in `speculative_output_marginal`.

### Where is the MTP exactness guarantee implemented?

The mathematical guarantee is implemented in `speculative_output_marginal` within [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py), while the runtime acceptance logic resides in [`mtplx/mtp_patch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_patch.py). The `acceptance_probability` function (lines 44-49) provides the core acceptance ratio calculation that makes the proof operational.