Are presence_penalty and frequency_penalty Supported in MTPLX's MTP? Implementation Guide

Yes, MTPLX fully supports OpenAI-style presence_penalty and frequency_penalty in its Multi-Token Prediction (MTP) pipeline, applying these penalties to both draft and target token distributions before speculative decoding begins.

MTPLX is an open-source inference engine designed for high-performance LLM serving with Multi-Token Prediction capabilities. The repository maintains strict compatibility with OpenAI's API specification, ensuring that repetition control parameters work seamlessly across both standard and speculative generation modes. This article examines the implementation details of presence_penalty and frequency_penalty support throughout the MTPLX sampling stack, from HTTP request parsing to low-level logit manipulation.

How MTPLX Implements presence_penalty and frequency_penalty

The support for repetition penalties in MTPLX's MTP is built into four distinct architectural layers. Each layer ensures that penalties propagate correctly from the API request down to the final token selection.

Sampler Configuration Layer

At the configuration level, the SamplerConfig dataclass in mtplx/sampling.py defines the penalty fields with safe defaults. Lines 19-27 declare both presence_penalty and frequency_penalty as floating-point values defaulting to 0.0, ensuring backward compatibility when clients omit these parameters.


# Conceptual representation based on mtplx/sampling.py structure

@dataclass
class SamplerConfig:
    temperature: float = 1.0
    top_p: float = 1.0
    top_k: int = 0
    presence_penalty: float = 0.0  # Defaults to no-op

    frequency_penalty: float = 0.0  # Defaults to no-op

Logit Adjustment Layer

The actual penalty computation occurs in apply_penalties() within mtplx/sampling.py (lines 61-67). This function clips incoming penalty values to the allowed range of -2.0 to 2.0, then subtracts the calculated penalties from raw logits before temperature scaling occurs. This clipping prevents extreme values from destabilizing the probability distribution.

Distribution Construction Pipeline

The distribution_from_logits() function (lines 14-23 in mtplx/sampling.py) orchestrates the sampling preparation. It first invokes apply_penalties() to modify logits based on token occurrence counts, then proceeds through the standard softmax, top-p, and top-k filtering pipeline. Crucially, this happens before the MTP speculative decoder processes the distributions, ensuring penalties affect both draft and target predictions consistently.

Request Policy and OpenAI Compatibility

Incoming HTTP requests are handled by mtplx/server/request_policy.py (lines 200-219), where the policy extractor parses presence_penalty and frequency_penalty from client payloads. If omitted, these fall back to the global defaults defined in the configuration. The OpenAI-compatible endpoint in mtplx/server/openai.py (lines 14972-15032) validates and forwards these fields to the sampler, maintaining full API parity with OpenAI's behavior.

Using presence_penalty and frequency_penalty with MTPLX MTP

You can configure these penalties programmatically through the Python API or via HTTP requests to the OpenAI-compatible endpoint.

Python API Implementation

Import the sampling utilities and instantiate SamplerConfig with non-zero penalty values:

from mtplx.sampling import SamplerConfig, distribution_from_logits
import numpy as np

# Configure sampler with repetition penalties

cfg = SamplerConfig(
    temperature=0.7,
    top_p=0.9,
    top_k=20,
    presence_penalty=1.2,   # Penalize tokens that have already appeared

    frequency_penalty=0.5,  # Stronger penalty for repeated occurrences

)

# Apply to model logits (50257-dim vocabulary example)

logits = np.random.randn(50257)
token_counts = {1234: 2, 5678: 1}  # Token ID: occurrence count

# Generate distribution respecting penalties

dist = distribution_from_logits(logits, cfg, token_counts=token_counts)

# dist can now be fed to the MTP speculative decoder

HTTP API Usage

Send penalties through the OpenAI-compatible endpoint:

curl -X POST http://localhost:8000/v1/mtplx/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4-mtp",
    "prompt": "Write a haiku about sunrise.",
    "presence_penalty": 1.0,
    "frequency_penalty": 0.3,
    "max_tokens": 20
  }'

The server extracts these fields in mtplx/server/request_policy.py and passes them to the underlying MTP sampler.

Technical Architecture of Penalty Application in MTP Batches

Unlike some speculative decoding implementations that apply penalties only to final token selection, MTPLX applies presence_penalty and frequency_penalty before logits enter the MTP batch processor. This architectural decision guarantees that:

  • Draft tokens generated by the MTP speculative model respect the same repetition constraints as the target model
  • Target verification uses penalty-adjusted distributions, preventing the acceptance of tokens that violate the user's repetition preferences
  • Batch consistency ensures that all tokens in an MTP batch face identical repetition constraints

The penalties are calculated based on token occurrence counts tracked across the sequence, with presence_penalty applying a fixed subtraction for any previously seen token, while frequency_penalty scales the subtraction by the token's occurrence frequency.

Summary

  • MTPLX fully supports presence_penalty and frequency_penalty in MTP mode, maintaining OpenAI API compatibility from configuration through generation.
  • Penalty application occurs in apply_penalties() within mtplx/sampling.py, clipping values to the -2.0 to 2.0 range before modifying logits.
  • Configuration is handled by SamplerConfig (lines 19-27) with safe zero defaults that prevent unintended behavior.
  • Request pipeline extracts penalties in mtplx/server/request_policy.py (lines 200-219) and validates them in mtplx/server/openai.py (lines 14972-15032).
  • MTP integration ensures penalties affect both draft and target distributions before speculative decoding begins, maintaining consistency across the entire batch.

Frequently Asked Questions

What is the valid range for presence_penalty and frequency_penalty in MTPLX?

MTPLX clips both penalty parameters to the range -2.0 to 2.0 in the apply_penalties() function (lines 61-67 of mtplx/sampling.py). Values outside this range are truncated to the nearest boundary, preventing distribution collapse while allowing negative values (which encourage repetition) and positive values (which discourage repetition).

Do presence_penalty and frequency_penalty affect speculative draft tokens in MTP?

Yes. The penalties are applied in distribution_from_logits() before the MTP speculative decoder processes the distributions. This ensures that both draft tokens generated by the speculative model and target tokens verified by the main model respect the same repetition constraints, maintaining consistency across the entire MTP batch.

How does MTPLX handle missing penalty parameters in API requests?

If a client omits presence_penalty or frequency_penalty from the request payload, MTPLX falls back to the global defaults defined in SamplerConfig (0.0 for both penalties). The mtplx/server/request_policy.py file (lines 200-219) handles this extraction logic, defaulting to no penalty when the fields are unspecified.

What is the difference between presence_penalty and frequency_penalty in MTPLX's implementation?

presence_penalty applies a fixed penalty to any token that has appeared at least once in the prior context, regardless of how many times it occurred. frequency_penalty scales the penalty linearly with the token's occurrence count, applying stronger penalties to tokens that appear frequently. Both are calculated in apply_penalties() and subtracted from the raw logits before temperature scaling and softmax normalization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →