# How the Hyper Scheduler Mode in MTPLX Differs from Serial and Batch Modes

> Discover how MTPLX hyper scheduler mode enforces strict singleton admission for serial inference paths and enables speculative width expansion unlike batch or serial modes.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: deep-dive
- Published: 2026-09-05

---

**The hyper scheduler mode in MTPLX enforces a strict singleton admission policy that preserves the exact serial inference path for single requests while optionally enabling speculative width expansion through a dedicated width seam.**

MTPLX is an open-source inference server that implements multiple scheduler modes to control how model work is admitted and executed. Unlike the standard serial mode or the multi-row mtp_batch mode, the hyper scheduler mode operates as a specialized "H0" chassis designed to maintain performance parity with serial execution while providing hooks for speculative generation.

## Admission Control Architecture

The primary distinction of the hyper scheduler mode lies in its admission gate implementation and concurrency model.

### Fixed Admission Cap

In [`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py), the hyper mode enforces `HYPER_ADMISSION_CAP = 1`, ensuring only one request enters the model-owner thread at a time. This cap is enforced by the model-owner FIFO without additional locks or semaphores. This differs fundamentally from **mtp_batch** mode, which allows concurrent requests up to the configured batch width, and from **serial** mode, which admits one request but lacks the explicit hyper admission gate architecture.

The admission mechanism uses the `HyperAdmissionGate` class rather than simple FIFO ordering. When a request arrives, the gate creates an admission ticket through `gate.admit()` that binds the request to the serial closure.

### Request Lifecycle

Once admitted, requests in hyper mode begin with **width = 1**, meaning they execute as singleton passthroughs. Only when a `HyperWidthExecutor` is installed can the width increase, and exclusively for self-generated speculative rows belonging to that same request. This contrasts with **mtp_batch** mode, where width may exceed one immediately based on external row aggregation, and **serial** mode, where width is permanently fixed at one with no speculative capability.

## Execution Path and Kernel Optimization

The hyper scheduler mode preserves critical optimizations by routing requests differently than batched modes.

### Serial Closure Preservation

When operating at width = 1, the hyper mode hands requests directly to the untouched serial closure (labeled `singleton` in the codebase). According to the implementation in [`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py), this path builds no batch objects, introduces no wait windows, and keeps all custom kernels active—including paged KV cache, Metal fast paths, and CompiledVerifyBank.

In contrast, **mtp_batch** mode routes requests through a batched driver that disables these optimizations because it assumes multi-row execution. The **serial** mode also uses the serial closure but lacks the hyper gate's ability to transition to wider speculative execution via `plan_cohort` and `record_cohort_width` methods.

### Width Expansion Mechanics

For width > 1 scenarios, hyper mode requires a user-provided `HyperWidthExecutor` to run the cohort through its `run_cohort` method. This executor must still call the serial closure for the singleton portion of the work, maintaining the byte-identical trajectory guarantee. The width seam is implemented through the gate's `plan_cohort` logic, which is absent in both serial and mtp_batch implementations.

## Telemetry and Performance Contracts

The hyper scheduler mode exposes distinct observability metrics and maintains strict performance requirements.

### Dedicated Telemetry Blocks

As documented in [`docs/dashboard.md`](https://github.com/youssofal/MTPLX/blob/main/docs/dashboard.md) and implemented in [`hyper.py`](https://github.com/youssofal/MTPLX/blob/main/hyper.py), hyper mode surfaces a `scheduler.hyper` stats block that reports stage `h0`, live width, the `serial_b1` passthrough label, and a histogram of requests-by-width. This histogram contains only width = 1 entries unless a width executor is actively generating speculative rows.

Serial mode exposes `scheduler.serial` counters without width-seam fields, while mtp_batch mode reports `scheduler.mtp_batch` with `batch_histogram` and fork-related statistics.

### Performance Guarantees

The hyper mode must achieve **≥ 0.99× serial token-per-second** at width = 1 and produce byte-identical output trajectories compared to the serial path. This parity gate is enforced through comments at the top of [`hyper.py`](https://github.com/youssofal/MTPLX/blob/main/hyper.py) and validated at runtime.

Mtp_batch mode optimizes for aggregate throughput across multiple requests but incurs an "eager-verify tax" (~380 ms versus 137 ms) when forced through the width-1 path, making hyper mode preferable for latency-sensitive single-request workloads.

## Configuration and Validation

Enabling hyper mode requires specific CLI flags and passes through rigorous validation.

### CLI Usage

To activate the hyper scheduler, launch MTPLX with:

```bash
mtplx serve --scheduler-mode hyper

```

This triggers `openai._validate_hyper_settings` in [`mtplx/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/openai.py), which rejects configurations violating the single-request contract. For example, the validator raises an error if `--max-active-requests` exceeds 1, whereas serial mode allows most flags and mtp_batch mode validates batch-specific parameters like `--batch-wait-ms`.

### Programmatic Integration

Developers can interact with the hyper gate programmatically:

```python
from mtplx.server.hyper import HyperAdmissionGate, HyperRequestMeta

gate = HyperAdmissionGate()
ticket = gate.admit(request_id="req-42", prompt_tokens=128)

def serial_closure():
    return {"output": "hello"}

bound = gate.bind(ticket, serial_closure, generation_mode="mtp")
result = bound()

```

### Custom Width Executors

To enable speculative generation within the hyper framework, install a custom width executor:

```python
class MyWidthExecutor:
    def plan(self, request: HyperRequestMeta):
        return HyperCohortPlan(width=4, reason="my_fork")
    
    def run_cohort(self, plan, singleton):
        speculative = {"spec_rows": plan.width - 1}
        return {**singleton(), **speculative}

gate.install_width_executor(MyWidthExecutor())

```

This pattern allows width expansion only for self-generated speculative rows while maintaining the serial path for the primary request, a capability unavailable in standard serial or mtp_batch modes.

## Summary

- **Hyper mode** enforces a singleton admission cap of 1 via `HyperAdmissionGate`, unlike **mtp_batch** which allows concurrent requests or **serial** which lacks the gate architecture.
- Requests start at width = 1 and execute through the untouched serial closure, preserving paged KV cache and custom kernels that batched modes disable.
- Optional speculative width expansion is supported only through user-installed `HyperWidthExecutor` instances that manage cohort planning via `plan_cohort` methods.
- The mode guarantees ≥ 0.99× serial performance and byte-identical trajectories at width = 1.
- Configuration requires `--scheduler-mode hyper` and passes through `openai._validate_hyper_settings` to ensure single-request semantics.
- Telemetry is exposed through the `scheduler.hyper` dashboard block, distinct from `scheduler.serial` and `scheduler.mtp_batch` implementations.

## Frequently Asked Questions

### Can I run multiple concurrent requests in hyper scheduler mode?

No. The hyper scheduler mode strictly enforces `HYPER_ADMISSION_CAP = 1` in [`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py), allowing only one active request at a time. The `openai._validate_hyper_settings` function explicitly rejects launch configurations where `--max-active-requests` exceeds 1, ensuring the singleton contract is maintained.

### How does hyper mode maintain performance parity with serial execution?

At width = 1, hyper mode routes requests directly to the serial closure without constructing batch objects or introducing wait windows, keeping all custom kernels like paged KV cache and Metal fast paths active. The implementation guarantees ≥ 0.99× serial token-per-second and byte-identical trajectories, verified by parity gates in [`hyper.py`](https://github.com/youssofal/MTPLX/blob/main/hyper.py).

### What is the difference between hyper mode and mtp_batch mode for single requests?

While both can theoretically process single requests, hyper mode preserves the optimized serial path and kernel configurations, whereas mtp_batch routes through a batched driver that disables optimizations like paged KV cache. Additionally, hyper mode allows dynamic width expansion for speculative rows within the same request via `HyperWidthExecutor`, while mtp_batch aggregates multiple external rows into fixed-width batches.

### When should I install a custom HyperWidthExecutor?

Install a custom `HyperWidthExecutor` when you need to generate speculative tokens for the currently executing request (self-generated rows). The executor's `plan` method determines the cohort width, and `run_cohort` manages execution, but all results must still funnel through the serial closure for the singleton portion. This is unnecessary for standard inference but essential for speculative decoding or fork-style generation patterns in hyper mode.