How the Hyper Scheduler Mode in MTPLX Differs from Serial and Batch Modes
The hyper scheduler mode in MTPLX enforces a strict singleton admission policy that preserves the exact serial inference path for single requests while optionally enabling speculative width expansion through a dedicated width seam.
MTPLX is an open-source inference server that implements multiple scheduler modes to control how model work is admitted and executed. Unlike the standard serial mode or the multi-row mtp_batch mode, the hyper scheduler mode operates as a specialized "H0" chassis designed to maintain performance parity with serial execution while providing hooks for speculative generation.
Admission Control Architecture
The primary distinction of the hyper scheduler mode lies in its admission gate implementation and concurrency model.
Fixed Admission Cap
In mtplx/server/hyper.py, the hyper mode enforces HYPER_ADMISSION_CAP = 1, ensuring only one request enters the model-owner thread at a time. This cap is enforced by the model-owner FIFO without additional locks or semaphores. This differs fundamentally from mtp_batch mode, which allows concurrent requests up to the configured batch width, and from serial mode, which admits one request but lacks the explicit hyper admission gate architecture.
The admission mechanism uses the HyperAdmissionGate class rather than simple FIFO ordering. When a request arrives, the gate creates an admission ticket through gate.admit() that binds the request to the serial closure.
Request Lifecycle
Once admitted, requests in hyper mode begin with width = 1, meaning they execute as singleton passthroughs. Only when a HyperWidthExecutor is installed can the width increase, and exclusively for self-generated speculative rows belonging to that same request. This contrasts with mtp_batch mode, where width may exceed one immediately based on external row aggregation, and serial mode, where width is permanently fixed at one with no speculative capability.
Execution Path and Kernel Optimization
The hyper scheduler mode preserves critical optimizations by routing requests differently than batched modes.
Serial Closure Preservation
When operating at width = 1, the hyper mode hands requests directly to the untouched serial closure (labeled singleton in the codebase). According to the implementation in mtplx/server/hyper.py, this path builds no batch objects, introduces no wait windows, and keeps all custom kernels active—including paged KV cache, Metal fast paths, and CompiledVerifyBank.
In contrast, mtp_batch mode routes requests through a batched driver that disables these optimizations because it assumes multi-row execution. The serial mode also uses the serial closure but lacks the hyper gate's ability to transition to wider speculative execution via plan_cohort and record_cohort_width methods.
Width Expansion Mechanics
For width > 1 scenarios, hyper mode requires a user-provided HyperWidthExecutor to run the cohort through its run_cohort method. This executor must still call the serial closure for the singleton portion of the work, maintaining the byte-identical trajectory guarantee. The width seam is implemented through the gate's plan_cohort logic, which is absent in both serial and mtp_batch implementations.
Telemetry and Performance Contracts
The hyper scheduler mode exposes distinct observability metrics and maintains strict performance requirements.
Dedicated Telemetry Blocks
As documented in docs/dashboard.md and implemented in hyper.py, hyper mode surfaces a scheduler.hyper stats block that reports stage h0, live width, the serial_b1 passthrough label, and a histogram of requests-by-width. This histogram contains only width = 1 entries unless a width executor is actively generating speculative rows.
Serial mode exposes scheduler.serial counters without width-seam fields, while mtp_batch mode reports scheduler.mtp_batch with batch_histogram and fork-related statistics.
Performance Guarantees
The hyper mode must achieve ≥ 0.99× serial token-per-second at width = 1 and produce byte-identical output trajectories compared to the serial path. This parity gate is enforced through comments at the top of hyper.py and validated at runtime.
Mtp_batch mode optimizes for aggregate throughput across multiple requests but incurs an "eager-verify tax" (~380 ms versus 137 ms) when forced through the width-1 path, making hyper mode preferable for latency-sensitive single-request workloads.
Configuration and Validation
Enabling hyper mode requires specific CLI flags and passes through rigorous validation.
CLI Usage
To activate the hyper scheduler, launch MTPLX with:
mtplx serve --scheduler-mode hyper
This triggers openai._validate_hyper_settings in mtplx/openai.py, which rejects configurations violating the single-request contract. For example, the validator raises an error if --max-active-requests exceeds 1, whereas serial mode allows most flags and mtp_batch mode validates batch-specific parameters like --batch-wait-ms.
Programmatic Integration
Developers can interact with the hyper gate programmatically:
from mtplx.server.hyper import HyperAdmissionGate, HyperRequestMeta
gate = HyperAdmissionGate()
ticket = gate.admit(request_id="req-42", prompt_tokens=128)
def serial_closure():
return {"output": "hello"}
bound = gate.bind(ticket, serial_closure, generation_mode="mtp")
result = bound()
Custom Width Executors
To enable speculative generation within the hyper framework, install a custom width executor:
class MyWidthExecutor:
def plan(self, request: HyperRequestMeta):
return HyperCohortPlan(width=4, reason="my_fork")
def run_cohort(self, plan, singleton):
speculative = {"spec_rows": plan.width - 1}
return {**singleton(), **speculative}
gate.install_width_executor(MyWidthExecutor())
This pattern allows width expansion only for self-generated speculative rows while maintaining the serial path for the primary request, a capability unavailable in standard serial or mtp_batch modes.
Summary
- Hyper mode enforces a singleton admission cap of 1 via
HyperAdmissionGate, unlike mtp_batch which allows concurrent requests or serial which lacks the gate architecture. - Requests start at width = 1 and execute through the untouched serial closure, preserving paged KV cache and custom kernels that batched modes disable.
- Optional speculative width expansion is supported only through user-installed
HyperWidthExecutorinstances that manage cohort planning viaplan_cohortmethods. - The mode guarantees ≥ 0.99× serial performance and byte-identical trajectories at width = 1.
- Configuration requires
--scheduler-mode hyperand passes throughopenai._validate_hyper_settingsto ensure single-request semantics. - Telemetry is exposed through the
scheduler.hyperdashboard block, distinct fromscheduler.serialandscheduler.mtp_batchimplementations.
Frequently Asked Questions
Can I run multiple concurrent requests in hyper scheduler mode?
No. The hyper scheduler mode strictly enforces HYPER_ADMISSION_CAP = 1 in mtplx/server/hyper.py, allowing only one active request at a time. The openai._validate_hyper_settings function explicitly rejects launch configurations where --max-active-requests exceeds 1, ensuring the singleton contract is maintained.
How does hyper mode maintain performance parity with serial execution?
At width = 1, hyper mode routes requests directly to the serial closure without constructing batch objects or introducing wait windows, keeping all custom kernels like paged KV cache and Metal fast paths active. The implementation guarantees ≥ 0.99× serial token-per-second and byte-identical trajectories, verified by parity gates in hyper.py.
What is the difference between hyper mode and mtp_batch mode for single requests?
While both can theoretically process single requests, hyper mode preserves the optimized serial path and kernel configurations, whereas mtp_batch routes through a batched driver that disables optimizations like paged KV cache. Additionally, hyper mode allows dynamic width expansion for speculative rows within the same request via HyperWidthExecutor, while mtp_batch aggregates multiple external rows into fixed-width batches.
When should I install a custom HyperWidthExecutor?
Install a custom HyperWidthExecutor when you need to generate speculative tokens for the currently executing request (self-generated rows). The executor's plan method determines the cohort width, and run_cohort manages execution, but all results must still funnel through the serial closure for the singleton portion. This is unnecessary for standard inference but essential for speculative decoding or fork-style generation patterns in hyper mode.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →