# Are presence_penalty and frequency_penalty Supported in MTPLX's MTP? Implementation Guide

> Discover if MTPLX supports presence_penalty and frequency_penalty in MTP. This guide explains how these OpenAI-style penalties enhance draft and target token distributions before decoding.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-04

---

**Yes, MTPLX fully supports OpenAI-style presence_penalty and frequency_penalty in its Multi-Token Prediction (MTP) pipeline, applying these penalties to both draft and target token distributions before speculative decoding begins.**

MTPLX is an open-source inference engine designed for high-performance LLM serving with Multi-Token Prediction capabilities. The repository maintains strict compatibility with OpenAI's API specification, ensuring that repetition control parameters work seamlessly across both standard and speculative generation modes. This article examines the implementation details of `presence_penalty` and `frequency_penalty` support throughout the MTPLX sampling stack, from HTTP request parsing to low-level logit manipulation.

## How MTPLX Implements presence_penalty and frequency_penalty

The support for repetition penalties in MTPLX's MTP is built into four distinct architectural layers. Each layer ensures that penalties propagate correctly from the API request down to the final token selection.

### Sampler Configuration Layer

At the configuration level, the `SamplerConfig` dataclass in [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py) defines the penalty fields with safe defaults. Lines 19-27 declare both `presence_penalty` and `frequency_penalty` as floating-point values defaulting to `0.0`, ensuring backward compatibility when clients omit these parameters.

```python

# Conceptual representation based on mtplx/sampling.py structure

@dataclass
class SamplerConfig:
    temperature: float = 1.0
    top_p: float = 1.0
    top_k: int = 0
    presence_penalty: float = 0.0  # Defaults to no-op

    frequency_penalty: float = 0.0  # Defaults to no-op

```

### Logit Adjustment Layer

The actual penalty computation occurs in `apply_penalties()` within [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py) (lines 61-67). This function clips incoming penalty values to the allowed range of **-2.0 to 2.0**, then subtracts the calculated penalties from raw logits before temperature scaling occurs. This clipping prevents extreme values from destabilizing the probability distribution.

### Distribution Construction Pipeline

The `distribution_from_logits()` function (lines 14-23 in [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py)) orchestrates the sampling preparation. It first invokes `apply_penalties()` to modify logits based on token occurrence counts, then proceeds through the standard softmax, top-p, and top-k filtering pipeline. Crucially, this happens **before** the MTP speculative decoder processes the distributions, ensuring penalties affect both draft and target predictions consistently.

### Request Policy and OpenAI Compatibility

Incoming HTTP requests are handled by [`mtplx/server/request_policy.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/request_policy.py) (lines 200-219), where the policy extractor parses `presence_penalty` and `frequency_penalty` from client payloads. If omitted, these fall back to the global defaults defined in the configuration. The OpenAI-compatible endpoint in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (lines 14972-15032) validates and forwards these fields to the sampler, maintaining full API parity with OpenAI's behavior.

## Using presence_penalty and frequency_penalty with MTPLX MTP

You can configure these penalties programmatically through the Python API or via HTTP requests to the OpenAI-compatible endpoint.

### Python API Implementation

Import the sampling utilities and instantiate `SamplerConfig` with non-zero penalty values:

```python
from mtplx.sampling import SamplerConfig, distribution_from_logits
import numpy as np

# Configure sampler with repetition penalties

cfg = SamplerConfig(
    temperature=0.7,
    top_p=0.9,
    top_k=20,
    presence_penalty=1.2,   # Penalize tokens that have already appeared

    frequency_penalty=0.5,  # Stronger penalty for repeated occurrences

)

# Apply to model logits (50257-dim vocabulary example)

logits = np.random.randn(50257)
token_counts = {1234: 2, 5678: 1}  # Token ID: occurrence count

# Generate distribution respecting penalties

dist = distribution_from_logits(logits, cfg, token_counts=token_counts)

# dist can now be fed to the MTP speculative decoder

```

### HTTP API Usage

Send penalties through the OpenAI-compatible endpoint:

```bash
curl -X POST http://localhost:8000/v1/mtplx/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4-mtp",
    "prompt": "Write a haiku about sunrise.",
    "presence_penalty": 1.0,
    "frequency_penalty": 0.3,
    "max_tokens": 20
  }'

```

The server extracts these fields in [`mtplx/server/request_policy.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/request_policy.py) and passes them to the underlying MTP sampler.

## Technical Architecture of Penalty Application in MTP Batches

Unlike some speculative decoding implementations that apply penalties only to final token selection, MTPLX applies `presence_penalty` and `frequency_penalty` **before** logits enter the MTP batch processor. This architectural decision guarantees that:

- **Draft tokens** generated by the MTP speculative model respect the same repetition constraints as the target model
- **Target verification** uses penalty-adjusted distributions, preventing the acceptance of tokens that violate the user's repetition preferences
- **Batch consistency** ensures that all tokens in an MTP batch face identical repetition constraints

The penalties are calculated based on token occurrence counts tracked across the sequence, with `presence_penalty` applying a fixed subtraction for any previously seen token, while `frequency_penalty` scales the subtraction by the token's occurrence frequency.

## Summary

- **MTPLX fully supports** `presence_penalty` and `frequency_penalty` in MTP mode, maintaining OpenAI API compatibility from configuration through generation.
- **Penalty application** occurs in `apply_penalties()` within [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py), clipping values to the **-2.0 to 2.0** range before modifying logits.
- **Configuration** is handled by `SamplerConfig` (lines 19-27) with safe zero defaults that prevent unintended behavior.
- **Request pipeline** extracts penalties in [`mtplx/server/request_policy.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/request_policy.py) (lines 200-219) and validates them in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (lines 14972-15032).
- **MTP integration** ensures penalties affect both draft and target distributions before speculative decoding begins, maintaining consistency across the entire batch.

## Frequently Asked Questions

### What is the valid range for presence_penalty and frequency_penalty in MTPLX?

MTPLX clips both penalty parameters to the range **-2.0 to 2.0** in the `apply_penalties()` function (lines 61-67 of [`mtplx/sampling.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/sampling.py)). Values outside this range are truncated to the nearest boundary, preventing distribution collapse while allowing negative values (which encourage repetition) and positive values (which discourage repetition).

### Do presence_penalty and frequency_penalty affect speculative draft tokens in MTP?

Yes. The penalties are applied in `distribution_from_logits()` before the MTP speculative decoder processes the distributions. This ensures that both draft tokens generated by the speculative model and target tokens verified by the main model respect the same repetition constraints, maintaining consistency across the entire MTP batch.

### How does MTPLX handle missing penalty parameters in API requests?

If a client omits `presence_penalty` or `frequency_penalty` from the request payload, MTPLX falls back to the global defaults defined in `SamplerConfig` (0.0 for both penalties). The [`mtplx/server/request_policy.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/request_policy.py) file (lines 200-219) handles this extraction logic, defaulting to no penalty when the fields are unspecified.

### What is the difference between presence_penalty and frequency_penalty in MTPLX's implementation?

**`presence_penalty`** applies a fixed penalty to any token that has appeared at least once in the prior context, regardless of how many times it occurred. **`frequency_penalty`** scales the penalty linearly with the token's occurrence count, applying stronger penalties to tokens that appear frequently. Both are calculated in `apply_penalties()` and subtracted from the raw logits before temperature scaling and softmax normalization.