# Leviathan and Chen Rejection Sampling in MTPLX: Exact Token Generation and Unbiased Evaluation

> Explore Leviathan-Chen rejection sampling in MTPLX for exact token generation and unbiased pass@k evaluation. Improve code-generation benchmarks with this powerful technique.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: deep-dive
- Published: 2026-09-06

---

**TLDR:** MTPLX uses **Leviathan‑Chen law** for exact token-level generation with improved draft acceptance rates, and **Chen rejection sampling** for unbiased pass@k evaluation in code-generation benchmarks.

Large language model inference frameworks constantly balance **sampling fidelity** against **computational efficiency**. MTPLX, an open-source speculative decoding engine, implements two complementary techniques from the research literature—Leviathan‑Chen law and Chen rejection sampling—to address both generation quality and evaluation accuracy. These mechanisms operate in distinct parts of the codebase but share a common mathematical foundation in rejection sampling theory.

## What Is Leviathan‑Chen Law in MTPLX?

The **Leviathan‑Chen law** is MTPLX's implementation of an exact sampling algorithm that modifies how draft tokens are accepted during speculative decoding. Instead of capping each acceptance factor individually, the law caps the **running reach product at 1**, then **water‑fills** the remaining probability budget across the next-depth draft support.

### How Leviathan‑Chen Clipping Works

Traditional speculative decoding clips each factor separately, which can waste probability budget and reduce draft acceptance rates. The Leviathan‑Chen approach:

- Applies a **single clip** to the cumulative product
- Distributes residual probability through **water‑filling**
- Corrects using the scaled residual formula `(c·p − q)+`

This yields mathematically equivalent samples to the baseline method while drawing the same number of uniform random numbers. Empirical testing shows approximately **+1.85% tokens per window** improvement in offline benchmarks.

### Implementation Location

The Leviathan‑Chen logic resides in [`mtplx/generation.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py) at lines 331‑336:

```python

# Pseudo-structure based on source reference

# The actual implementation performs water-filling after clipping

# the running reach product to 1.0, then applies residual correction

def leviathan_chen_accept(draft_probs, target_probs, random_draws):
    # Cumulative product clipped at 1.0

    # Water-fill across next-depth support

    # Scaled residual: (c * p - q)^+

    pass  # See mtplx/generation.py#L331-L336 for full implementation

```

The algorithm is also referenced in [`mtplx/loop_guard.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/loop_guard.py) at line 31, where acceptance logic is documented as following **"Leviathan‑Chen residual math"** to maintain exact‑sampler guarantees.

### Enabling Leviathan‑Chen in MTPLX

Block verification mode activates the Leviathan‑Chen law via environment variable:

```python
import os

# Enable Leviathan-Chen block verification

os.environ["MTPLX_QWEN4_BLOCK_VERIFY"] = "1"

from mtplx import MTPLXClient

client = MTPLXClient()
response = client.generate(prompt="Explain quantum entanglement.")
print(response.text)

```

When `MTPLX_QWEN4_BLOCK_VERIFY` is set, the generation loop substitutes Leviathan‑Chen clipping for the older per‑factor clipping method. The randomness budget remains unchanged—only the acceptance efficiency improves.

## What Is Chen Rejection Sampling in MTPLX?

**Chen rejection sampling** in MTPLX refers to an unbiased estimator for the **pass@k** metric, implemented according to Chen et al. 2021. This statistical technique enables reliable evaluation of code-generation models without the bias inherent in naive success-rate estimation.

### The pass@k Estimation Problem

Evaluating code generation presents a challenge: for any single problem, you generate k candidate solutions, but success is binary (passes tests or fails). Simple averaging over problems with varying difficulty introduces bias. The Chen estimator solves this by providing:

- **Unbiased expectation** for true pass@k
- **Variance reduction** through proper conditioning
- **Cross-model comparability** regardless of sampling strategy

### Implementation Location

The Chen pass@k estimator lives in [`mtplx/benchmarks/code_eval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/code_eval.py) at line 342:

```python
from mtplx.benchmarks.code_eval import pass_at_k

# Example usage for unbiased evaluation

solutions = [
    "def sort_list(l): return sorted(l)",  # candidate 1

    "def sort_list(l): l.sort(); return l",  # candidate 2

    # ... more generated solutions

]

ground_truth = "def sort_list(l): return sorted(l)"

# Unbiased pass@k estimator following Chen et al. 2021

score = pass_at_k(solutions, ground_truth, k=10)
print(f"pass@10 = {score:.2%}")

```

## Comparing Leviathan‑Chen and Chen Rejection Sampling

| Aspect | Leviathan‑Chen Law | Chen Rejection Sampling |
|--------|-------------------|------------------------|
| **Primary purpose** | Exact token generation with higher acceptance | Unbiased pass@k estimation |
| **When applied** | During inference (generation loop) | During evaluation (benchmarking) |
| **Key operation** | Water‑fill residual correction | Unbiased statistical estimator |
| **Source file** | [`mtplx/generation.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py#L331)‑336) | [`mtplx/benchmarks/code_eval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/code_eval.py#L342) |
| **Mathematical basis** | Running reach product clipping | Chen et al. 2021 rejection sampling |

Despite sharing "Chen" in their names and both involving rejection sampling theory, these mechanisms serve **non‑overlapping purposes** in MTPLX's architecture. Leviathan‑Chen optimizes the generation path; Chen rejection sampling validates the results.

## How the Components Integrate

MTPLX's pipeline leverages both techniques at different stages:

1. **Generation phase** — With `MTPLX_QWEN4_BLOCK_VERIFY=1`, the Leviathan‑Chen law governs draft token acceptance, improving throughput while maintaining exact sampling guarantees.

2. **Verification phase** — The [`loop_guard.py`](https://github.com/youssofal/MTPLX/blob/main/loop_guard.py) module references Leviathan‑Chen residual math to ensure acceptance logic remains mathematically sound across depth transitions.

3. **Evaluation phase** — The Chen pass@k estimator in [`code_eval.py`](https://github.com/youssofal/MTPLX/blob/main/code_eval.py) provides unbiased performance metrics, enabling fair comparison between models using different generation strategies.

This separation of concerns allows MTPLX to optimize inference efficiency without compromising evaluation integrity.

## Key Files for Leviathan‑Chen and Chen Rejection Sampling

| File Path | Lines | Purpose |
|-----------|-------|---------|
| [`mtplx/generation.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py) | 331‑336 | Leviathan‑Chen clipping and water‑filling implementation |
| [`mtplx/loop_guard.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/loop_guard.py) | 31 | Reference to Leviathan‑Chen residual math in acceptance logic |
| [`mtplx/benchmarks/code_eval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/code_eval.py) | 342 | Chen pass@k estimator for unbiased evaluation |

## Summary

- **Leviathan‑Chen law** enables exact speculative decoding with ~1.85% improved token acceptance through single‑clip water‑filling, implemented in [`mtplx/generation.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py).
- **Chen rejection sampling** provides unbiased pass@k estimation for code‑generation benchmarks, located in [`mtplx/benchmarks/code_eval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/code_eval.py).
- Enable Leviathan‑Chen generation via `MTPLX_QWEN4_BLOCK_VERIFY=1` environment variable.
- Both techniques preserve mathematical correctness—exact sampling distribution and unbiased evaluation—while improving practical efficiency.

## Frequently Asked Questions

### What is the difference between Leviathan‑Chen law and Chen rejection sampling in MTPLX?

Leviathan‑Chen law is a **generation‑time optimization** for speculative decoding that improves draft acceptance rates through water‑filling, while Chen rejection sampling is an **evaluation‑time statistical estimator** for computing unbiased pass@k scores. They operate in different pipeline stages and solve different problems.

### How do I enable Leviathan‑Chen sampling in MTPLX?

Set the environment variable `MTPLX_QWEN4_BLOCK_VERIFY="1"` before initializing the MTPLX client. This toggles the generation loop to use Leviathan‑Chen clipping instead of per‑factor clipping.

### Where is the Chen pass@k estimator implemented?

The Chen pass@k estimator is implemented at line 342 of [`mtplx/benchmarks/code_eval.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/code_eval.py), following the formulation from Chen et al. 2021. Import `pass_at_k` from this module to evaluate code generation performance without estimator bias.

### Does Leviathan‑Chen sampling change the output distribution?

No. The Leviathan‑Chen law is an **exact sampler**—it produces samples from the identical target distribution as standard rejection sampling. The improvement comes from higher acceptance probabilities (enabling more tokens per verification step), not from distribution shift.