# What Is the Compiled Verify Bank and How Does Chunked Prefill Work in MTPLX?

> Discover the compiled verify bank and chunked prefill in MTPLX. Accelerate speculative decoding with GPU verification and optimize pipeline efficiency by splitting long prompts.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: deep-dive
- Published: 2026-09-13

---

**The compiled verify bank is a JIT-compiled verification kernel that accelerates speculative decoding by running the verify step entirely on the GPU, while chunked prefill splits long prompts into smaller token spans to sustain pipeline efficiency and avoid memory pressure.**

MTPLX is an open-source inference engine that optimizes large language model serving through aggressive kernel fusion and memory management. Two of its most impactful optimizations are the **compiled verify bank** for accelerating speculative verification and **chunked prefill** for handling long-context prompts without stalling the GPU. Understanding these subsystems is essential for operators tuning high-throughput generation workloads.

## The Compiled Verify Bank in MTPLX

The `CompiledVerifyBank` class provides a high-performance execution path for speculative verification, replacing the standard eager (CPU-side) verification loop with a fused GPU kernel.

### Architecture and Initialization

The bank is instantiated in [`mtplx/generation.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py) whenever compiled-verify mode is active. Its constructor configures a shadow cache, determines the maximum verify window length, and selects the appropriate capture backend.

According to the source code, the class is defined at [line 1665 of [`mtplx/graphbank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/graphbank.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/graphbank.py#L1665). During initialization, the bank:

- Allocates a temporary verify buffer sized by `max_verify_len` (typically 4 tokens)
- Configures the runtime capture backend (e.g., `"gdn"`)
- Reserves a growth budget controlled by `MTPLX_COMPILED_VERIFY_MAX_LEN` to handle variable-length generations without reallocating

When a generation request begins, the bank dispatches the compiled `verify_step` kernel via `forward_ar_capture`, then mirrors results back into the real KV cache without triggering the standard `rollback_state` path. This eliminates CPU-GPU synchronization overhead during the verify phase.

### Parity Modes and Correctness

To ensure numerical equivalence between compiled and eager paths, MTPLX implements **parity checking**. When `_compiled_verify_mode` is set to `"parity"` or `"parity2"`, the bank executes both kernels and compares outputs.

If a mismatch occurs, the code raises `CompiledVerifyParityError` at [lines 1555-1562 of [`graphbank.py`](https://github.com/youssofal/MTPLX/blob/main/graphbank.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/graphbank.py#L1555). Statistics including `compiled_calls`, `fallback_calls`, and `parity_failures` are tracked on the internal `stats` dictionary (see [lines 1830-1845](https://github.com/youssofal/MTPLX/blob/main/mtplx/graphbank.py#L1830)).

### Fallback Mechanisms

The compiled verify bank automatically degrades to eager verification when hardware constraints are detected. At [lines 1649-1657](https://github.com/youssofal/MTPLX/blob/main/mtplx/graphbank.py#L1649), the constructor checks if the model uses unsupported quantization widths (e.g., 6-bit). If so, it sets `permanent_eager=True`, ensuring the request uses standard CPU verification for its entire lifetime.

```python
from mtplx.graphbank import CompiledVerifyBank

# Instantiate for a 4-token verify window with parity checking

bank = CompiledVerifyBank(
    runtime=rt,
    max_verify_len=4,
    request_max_tokens=8192,
    capture_backend="gdn",
    parity=True,  # Enable correctness checking against eager path

    restored_tokens=prompt_state.cached_tokens
)

```

## Chunked Prefill Execution

Prefill—the process of loading a prompt into the key-value cache—can become a bottleneck for long contexts. MTPLX addresses this by splitting prompts into discrete chunks that better utilize the GPU's parallel execution units.

### Span Generation Logic

The core function `_iter_prefill_chunk_spans` at [line 1409 of [`mtplx/generation.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py#L1409) generates `(start, end)` tuples representing token ranges to process. The helper `_split_spans_at` (defined at [line 1381](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py#L1381)) inserts mandatory split points wherever stable-prefix edges occur, ensuring that chunk boundaries align with semantic prompt boundaries.

When `MTPLX_SUSTAINED_PREFILL_LAYOUT` is enabled, the iterator respects these constraints while maximizing chunk throughput. If sustained prefill is disabled, the function returns a single span covering the entire prompt ([line 1417](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py#L1417)).

### Chunk Size Configuration

The `_prefill_chunk_size()` function at [line 1336](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py#L1336) determines chunk dimensions by first checking the `MTPLX_PREFILL_CHUNK_SIZE` environment variable, then falling back to a heuristic based on the model's memory budget. This allows operators to tune prefill granularity without recompiling the engine.

```python
from mtplx.generation import _iter_prefill_chunk_spans

# Split a 3000-token prompt into 512-token chunks

# while preserving a stable prefix at token 1024

spans = _iter_prefill_chunk_spans(
    token_count=3000,
    chunk_size=512,
    mandatory_edges=(1024,)
)

for start, end in spans:
    process_chunk(prompt_ids[start:end])  # Execute prefill kernel

```

## Integration in Generation

In production runs, these systems work in tandem. The generation driver at [line 8640 of [`generation.py`](https://github.com/youssofal/MTPLX/blob/main/generation.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/generation.py#L8640) conditionally instantiates `CompiledVerifyBank` when speculative decoding is active, while the prefill phase uses the chunked iterator to stream prompt tokens into the cache. This separation allows the verify bank to focus on short-window validation (typically 4 tokens) while the prefill system handles arbitrary context lengths.

## Summary

- **CompiledVerifyBank** is a JIT-compiled verification accelerator that runs speculative validation entirely on GPU, falling back to eager mode for unsupported quantization formats.
- **Parity modes** (`parity` and `parity2`) enable deterministic validation of compiled kernels by comparing against reference implementations.
- **Chunked prefill** uses `_iter_prefill_chunk_spans` to break long prompts into manageable GPU kernels, configured via `MTPLX_PREFILL_CHUNK_SIZE`.
- Both subsystems are gated by environment variables and runtime checks, allowing safe deployment across heterogeneous hardware.

## Frequently Asked Questions

### What triggers the compiled verify bank to fall back to eager verification?

The bank detects unsupported model configurations—specifically 6-bit quantization or incompatible tensor layouts—during initialization at [`graphbank.py`](https://github.com/youssofal/MTPLX/blob/main/graphbank.py) lines 1649-1657. When detected, it sets `permanent_eager=True`, forcing all subsequent verification steps to use the standard CPU-side kernel for that request.

### How does chunked prefill affect prompt caching?

Chunked prefill preserves prompt caching semantics by respecting `mandatory_edges` parameters. When a stable prefix boundary is provided, `_split_spans_at` ensures chunk boundaries align exactly with cached positions, preventing redundant computation while still allowing the remaining prompt to be processed in parallel chunks.

### Can I use parity mode in production environments?

Parity mode is designed for validation and debugging, not production. When enabled, every verify step runs twice (compiled plus eager) and compares outputs, doubling compute overhead. The `CompiledVerifyParityError` exception will halt generation on mismatch, making it suitable for testing new compiled kernels but inappropriate for latency-sensitive serving.

### What is the relationship between `max_verify_len` and the compiled verify bank?

`max_verify_len` determines the size of the verify window that gets JIT-compiled. The bank allocates internal buffers and capture graphs sized for this specific length—commonly 4 tokens. If a generation request requires verifying a different number of tokens, the bank either handles it within the pre-allocated growth reserve (up to `MTPLX_COMPILED_VERIFY_MAX_LEN`) or falls back to eager verification for that step.