# Understanding the Hyper Scheduler Mode in MTPLX

> Discover MTPLX hyper scheduler mode, a CPU-safe singleton option for deterministic latency. Learn how it limits requests to one and disables batching for predictable performance.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: internals
- Published: 2026-09-06

---

**The hyper scheduler mode is a first-class, CPU-safe singleton scheduling option in MTPLX that enforces single-request admission for deterministic latency by limiting active requests to one and disabling batching heuristics.**

MTPLX is a high-performance inference server that supports multiple scheduling strategies for managing request throughput. The **hyper scheduler mode** provides a specialized execution path designed for deterministic latency and simplified resource management in single-request scenarios. This mode operates as a singleton chassis defined in the `SchedulerMode` enumeration, making it ideal for CPU-bound deployments or debugging scenarios where predictable performance is prioritized over throughput.

## What Is the Hyper Scheduler Mode?

The hyper scheduler mode is exposed through the `SchedulerMode` enumeration in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py), where it is defined as the string value `"hyper"`. Unlike batch-oriented schedulers, this mode implements a strict **singleton admission policy** (`hyper_singleton_admission_1`) that permits only one active request at a time, effectively eliminating concurrent multi-admission interference.

### Core Characteristics

- **Single Admission Enforcement**: The scheduler strictly limits concurrent execution to a single request through the `hyper_singleton_admission_1` policy, removing queue contention.
- **CPU-Safe Execution**: The mode is explicitly designed to operate without GPU-specific batching logic, ensuring safe and efficient operation in CPU-only environments.
- **Strict Configuration Validation**: When activated, MTPLX enforces three specific constraints via internal validation logic:
  - `--max-active-requests` must be set to `1`
  - `--batching-preset` must be set to `solo`
  - `--batch-wait-ms` must be set to `0`
  
  Any deviation triggers a `SystemExit` with a descriptive error message, as verified in [`tests/test_hyper_scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_hyper_scheduler.py).

## Configuration and Activation

You can activate the hyper scheduler mode via command-line arguments or environment variables. The parser validates that all required constraints are satisfied before initializing the singleton chassis.

### Command-Line Interface

Pass the `--scheduler-mode` flag with the required constraint flags when starting the server:

```bash
python -m mtplx.server serve \
    --scheduler-mode hyper \
    --max-active-requests 1 \
    --batching-preset solo \
    --batch-wait-ms 0

```

### Environment Variables

Alternatively, set the `MTPLX_SCHEDULER_MODE` environment variable to automatically select hyper mode without explicitly passing CLI flags:

```bash
export MTPLX_SCHEDULER_MODE=hyper
python -m mtplx.server serve

```

When this environment variable is set, the parsed arguments automatically adopt the hyper mode configuration, though the validation constraints still apply.

## Internal Implementation Details

The hyper scheduler implementation spans several critical files within the MTPLX repository, coordinating enum definitions, validation logic, and the singleton execution chassis.

### Enum Definition and Argument Parsing

In [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py), the `SchedulerMode` enum registers `"hyper"` as a valid scheduling option. The server's argument parser, accessible via `openai.parse_args()`, recognizes this value and constructs a configuration object that the server uses to instantiate the hyper scheduler.

### Validation Logic

The `_validate_hyper_settings` function enforces the strict configuration requirements. According to [`tests/test_hyper_scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_hyper_scheduler.py), this validation ensures that:
1. Maximum active requests equals 1
2. Batching preset is configured for solo execution
3. Batch wait time is zero milliseconds

Failure to meet any of these conditions results in immediate termination with a clear error message indicating which constraint was violated.

### Singleton Chassis

The core scheduling logic resides in [`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py). This module implements the singleton chassis that manages the single-admission policy. When the server reports health status, the scheduler identifies itself via the `_scheduler_policy_label` function, which returns the string `"hyper_singleton_admission_1"` for telemetry and monitoring purposes.

## When to Use the Hyper Scheduler Mode

Choose the hyper scheduler mode when your deployment requires specific operational characteristics that prioritize predictability over throughput.

**Ideal Use Cases:**

- **Deterministic Latency**: Eliminating concurrent requests removes queueing variability, providing predictable response times essential for debugging or latency-sensitive applications.
- **CPU-Only Deployments**: Operating without GPU batching heuristics makes this mode suitable for pure-CPU inference workloads where GPU acceleration is unavailable or unnecessary.
- **Simplified Resource Management**: The absence of complex queueing logic reduces overhead and configuration complexity for low-throughput scenarios or single-stream inference.

**Avoid This Mode When:**

You require high throughput, parallel request processing, or GPU-accelerated batching. The single-admission constraint inherently limits concurrency, making alternative scheduler modes such as `mtp_batch` or `solo` more appropriate for high-volume production environments.

## Programmatic Configuration

For testing or custom server builds, you can configure the hyper mode programmatically using the internal API:

```python
from mtplx.batching.state import SchedulerMode
from mtplx.server import openai

# Construct arguments programmatically as done in the test suite

args = openai.parse_args([
    "--warmup-tokens", "0",
    "--scheduler-mode", "hyper"
])

# Generate scheduler configuration from parsed args

config = openai._scheduler_config_from_args(args)

# Verify the policy label reflects singleton admission

assert openai._scheduler_policy_label(config) == "hyper_singleton_admission_1"

```

This programmatic approach is useful for unit testing scheduler behavior or integrating MTPLX into larger Python applications that manage their own configuration lifecycle.

## Summary

- The **hyper scheduler mode** in MTPLX provides a CPU-safe, singleton scheduling option defined in the `SchedulerMode` enum within [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py).
- It enforces a single-admission policy (`hyper_singleton_admission_1`) that limits active requests to one and disables batching heuristics for deterministic execution.
- Configuration requires strict adherence to `--max-active-requests 1`, `--batching-preset solo`, and `--batch-wait-ms 0`, validated by `_validate_hyper_settings`.
- The implementation lives in [`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py), with comprehensive validation and policy labeling tested in [`tests/test_hyper_scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/tests/test_hyper_scheduler.py).
- This mode is ideal for deterministic latency and CPU-only workloads but should not be used for high-throughput GPU batching scenarios.

## Frequently Asked Questions

### What is the hyper scheduler mode in MTPLX?

The hyper scheduler mode is a first-class scheduling option that operates as a singleton chassis for CPU-safe execution. According to the MTPLX source code in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py), it is defined as the `"hyper"` value in the `SchedulerMode` enumeration and limits the server to processing exactly one request at a time, bypassing GPU batching logic to provide deterministic latency.

### How do I enable the hyper scheduler mode?

You can enable it by passing `--scheduler-mode hyper` to the CLI along with the required constraint flags (`--max-active-requests 1`, `--batching-preset solo`, `--batch-wait-ms 0`), or by setting the `MTPLX_SCHEDULER_MODE=hyper` environment variable before starting the server. Both methods trigger the same validation logic in the argument parser.

### Why does the hyper mode require max-active-requests to be 1?

The hyper scheduler enforces a singleton admission policy (`hyper_singleton_admission_1`) to prevent concurrent multi-admission interference. Setting `--max-active-requests` to 1 ensures strict single-request execution, which is fundamental to the mode's deterministic latency guarantees and CPU-safe operation without batching overhead.

### Can I use the hyper scheduler with GPU acceleration?

No, the hyper scheduler mode is specifically designed for CPU-safe execution and does not utilize GPU batching heuristics. For GPU-accelerated inference with higher throughput, use alternative scheduler modes such as `mtp_batch` or `solo` with appropriate batching presets configured in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py).