How Temperature and top_p Settings Impact MTPLX's MTP Exactness
Temperature and top_p settings shape the target probability distribution used by MTPLX's MTP module, but they do not violate the mathematical exactness guarantee—the marginal distribution of accepted tokens always matches the filtered target distribution defined by those parameters.
MTPLX is a speculative decoding framework that accelerates inference through its Memory-Token-Projection (MTP) module. The impact of temperature and top_p settings on MTPLX's MTP exactness depends on how these parameters transform the raw logits into the effective target distribution that the acceptance mechanism evaluates against the draft model's output.
How MTPLX Constructs the Target Distribution
The MTP exactness guarantee assumes the target distribution is created through a two-stage pipeline in mtplx/sampling.py. This distribution represents what the full model would sample, and MTP ensures the final output matches it exactly—even when using aggressive filtering.
Temperature Scaling in Softmax
Temperature scaling occurs at lines 80-86 in mtplx/sampling.py, where the softmax function divides logits by the temperature parameter before exponentiation:
softmax(logits, temperature=config.temperature)
- Lower temperature (< 1) sharpens the distribution, concentrating probability mass on the highest-logit tokens and moving toward deterministic output.
- Higher temperature (> 1) flattens the distribution, increasing entropy and giving weight to lower-probability tokens.
Top-p and Top-k Filtering
After temperature scaling, apply_top_p_top_k (lines 15-24 in mtplx/sampling.py) truncates the distribution:
apply_top_p_top_k(probs, top_p=config.top_p, top_k=config.top_k)
- Top-p (nucleus filtering) retains only the highest-probability tokens whose cumulative probability reaches the
top_pthreshold, then renormalizes the remaining mass. - Top-k restricts the candidate set to the
khighest-probability tokens regardless of cumulative probability.
Together, these stages define the effective target distribution that MTP treats as ground truth.
The Acceptance Mechanism and Exactness Proof
MTP maintains exactness through a rejection sampling scheme implemented in acceptance_probability (lines 44-49 in mtplx/sampling.py). The algorithm compares the target probability p against the draft probability q for each proposed token:
p = _probability(target_p, token_id)
q = _probability(draft_q, token_id)
if q <= 0:
return 1.0 if p > 0 else 0.0
return min(1.0, p / q)
When a draft token is rejected, MTP falls back to residual_distribution, which samples from the renormalized remainder of the target distribution. The speculative_output_marginal function contains the formal proof that this acceptance/rejection cycle preserves the exact marginal distribution of the target—including all temperature and top_p modifications applied beforehand.
Practical Impact on MTP Behavior
While exactness is mathematically preserved regardless of settings, the practical behavior of speculative decoding changes significantly based on distribution shape:
Low Temperature (→ 0) and Deterministic Output
When temperature approaches zero, the target distribution collapses to a one-hot vector where the top token has probability 1.0. In mtplx/sampling.py, this means:
- Acceptance probability becomes 1.0 for the top token (if the draft proposes it) or 0.0 (if not)
- The residual distribution trivially returns the top token
- Exactness holds but speculative decoding gains minimal speedup since rejections are absolute
High Temperature with Full Top-p
With temperature > 1.0 and top_p = 1.0, the distribution spreads across many tokens:
- The draft model's approximations more frequently mismatch the target distribution
acceptance_probabilityreturns values < 1.0 more often, triggering the residual path- Throughput decreases but the output remains mathematically exact per the proof in
speculative_output_marginal
Aggressive Top-p Truncation (e.g., 0.5)
Setting top_p=0.5 removes the low-probability tail before MTP evaluation:
- Exactness is guaranteed with respect to the truncated distribution, not the original full distribution
- The acceptance calculation in
mtplx/mtp_patch.pyuses this filtered distribution as its target reference - This effectively changes the "ground truth" the model commits to matching
Configuration Examples
These examples from mtplx/sampling.py demonstrate how SamplerConfig parameters alter the effective target distribution while preserving MTP exactness:
# Example 1: Deterministic output (greedy decoding)
cfg = SamplerConfig(temperature=0.0, top_p=1.0, top_k=0)
probs = distribution_from_logits(logits, cfg)
# Target distribution is one-hot; MTP acceptance always succeeds for the top token
# Example 2: Standard sampling with nucleus filtering
cfg = SamplerConfig(temperature=0.8, top_p=0.95, top_k=0)
probs = distribution_from_logits(logits, cfg)
# Softened distribution with rare tokens truncated;
# rejections trigger residual_distribution to maintain exact marginal
# Example 3: Aggressive truncation with moderate randomness
cfg = SamplerConfig(temperature=1.0, top_p=0.5, top_k=0)
probs = distribution_from_logits(logits, cfg)
# Exactness holds for this heavily filtered distribution only
Summary
- Temperature and top_p modify the target distribution in
mtplx/sampling.pybefore MTP evaluation, but do not break the exactness guarantee. - Exactness is distribution-relative: MTP guarantees the output marginal matches whatever distribution results from your temperature, top_p, and top_k settings—not necessarily the untruncated base distribution.
- Low temperature reduces rejection rates but limits diversity; high temperature increases rejections but preserves exactness through
residual_distribution. - Top-p < 1 establishes a new "ground truth" by truncating tails; MTP's proof in
speculative_output_marginalapplies to this truncated reference.
Frequently Asked Questions
Does increasing temperature break MTP exactness?
No. Higher temperature flattens the probability distribution in softmax, but the acceptance_probability function in mtplx/sampling.py still ensures the final sequence follows that flattened distribution exactly. The proof in speculative_output_marginal holds for any valid probability distribution, regardless of entropy.
What happens when top_p is set to 1.0 versus 0.9?
At top_p=1.0, MTP targets the full distribution produced after temperature scaling. At top_p=0.9, apply_top_p_top_k removes the bottom 10% of probability mass and renormalizes. MTP then guarantees exactness relative to this truncated 90% nucleus distribution, not the original full distribution.
How does the residual distribution maintain exactness after rejections?
When the draft token is rejected (probability min(1.0, p/q) < 1), residual_distribution samples from the remaining probability mass of the target distribution, excluding the accepted portion. This rejection sampling technique ensures the marginal distribution of the accepted token—or the fallback sample—exactly equals the target distribution, as formalized in speculative_output_marginal.
Where is the MTP exactness guarantee implemented?
The mathematical guarantee is implemented in speculative_output_marginal within mtplx/sampling.py, while the runtime acceptance logic resides in mtplx/mtp_patch.py. The acceptance_probability function (lines 44-49) provides the core acceptance ratio calculation that makes the proof operational.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →