# How to Configure the Scheduler Mode in MTPLX

> Learn how to configure the scheduler mode in MTPLX using the config file, CLI flag, or ServerArgs. Explore options like serial, cooperative, and batch modes.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-06

---

**You can configure the scheduler mode in MTPLX via the `scheduler_mode` key in `~/.mtplx/config.toml`, the `--scheduler-mode` CLI flag, or the `scheduler_mode` parameter in `ServerArgs` objects, with valid options including `serial`, `cooperative`, `ar_batch`, `mtp_batch`, `mtp_cohort_experimental`, and `hyper`.**

MTPLX is an open-source inference engine that optimizes request execution through a pluggable runtime scheduler. Configuring the scheduler mode determines how the engine batches and executes inference requests, directly impacting throughput and latency for workloads ranging from single-request scenarios to high-concurrency batches. This guide covers the configuration mechanisms and implementation details based on the youssofal/MTPLX source code.

## Available Scheduler Modes in MTPLX

The available scheduling strategies are defined as a `StrEnum` in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py) (lines 18-33). Each mode corresponds to a concrete `ModelWorkScheduler` implementation:

- **`serial`**: Solo MTP oracle processing one request at a time (default).
- **`cooperative`**: Cooperative batching for mixed foreground and background workloads.
- **`ar_batch`**: Batched autoregressive decode lane optimized for concurrent, decode-heavy requests.
- **`mtp_batch`**: Batched MTP decode lane designed for prefill-heavy inference loads.
- **`mtp_cohort_experimental`**: Experimental cohort-aware MTP scheduling for research scenarios.
- **`hyper`**: Width-1 "hyper" chassis for single-request speculative decoding.

According to the implementation in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (lines 792-794), the runtime enforces mode-specific invariants, such as requiring `mtp_batch` for balanced MTP configurations.

## Configuration Methods

You can configure the scheduler mode through three interfaces, all reading from the same `UserConfig` schema.

### Config File

Create or edit `~/.mtplx/config.toml` and set the `scheduler_mode` key:

```toml

# ~/.mtplx/config.toml

scheduler_mode = "mtp_batch"

```

The configuration loader in [`mtplx/config.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/config.py) (lines 88-90) normalizes this value using `_str_or_none` and stores it in `UserConfig.scheduler_mode`.

### CLI Flag

Override the config file setting using the `--scheduler-mode` flag available in commands like `mtplx serve` and `mtplx start`:

```bash
mtplx serve --scheduler-mode ar_batch

```

The argument is registered in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) (lines 65-78) with `"serial"` as the default value and choices limited to the `SchedulerMode` enum values.

### Programmatic API

When building server arguments programmatically, pass `scheduler_mode` as a string parameter to `ServerArgs` or `OpenAIArgs`:

```python
from mtplx.commands.public import ServerArgs

args = ServerArgs(
    model="models/example",
    scheduler_mode="hyper"
)

```

The `UserConfig` object processes this parameter identically to CLI and file-based configuration.

## Technical Implementation Details

### Configuration Propagation

When MTPLX builds the backend server command in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) (lines 11523-11524), it conditionally appends `--scheduler-mode <mode>` to the generated command unless the mode matches the default `serial`. This ensures the scheduler selection propagates to the runtime environment.

### Runtime Enforcement

Inside the server implementation, specifically in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) and [`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py), the scheduler mode string determines the `ModelWorkScheduler` configuration. The code validates that specific features, such as balanced MTP batching, require the corresponding `scheduler_mode` setting.

## Selecting the Right Scheduler Mode

Choose your scheduler mode based on traffic patterns and latency requirements:

- **`serial`**: Use for lowest-latency single-request workloads or when debugging inference behavior.
- **`cooperative`**: Best for mixed workloads requiring cooperative batching without AR-specific optimizations.
- **`ar_batch`**: Ideal for high-throughput, decode-heavy traffic with concurrent requests.
- **`mtp_batch`**: Optimized for prefill-heavy loads where batching MTP decode steps improves throughput.
- **`mtp_cohort_experimental`**: Deploy for research into cohort-aware MTP scheduling strategies.
- **`hyper`**: Suitable for single-request speculative decoding scenarios requiring width-1 chassis.

## Summary

- MTPLX supports six scheduler modes defined in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py): `serial`, `cooperative`, `ar_batch`, `mtp_batch`, `mtp_cohort_experimental`, and `hyper`.
- Configure via `~/.mtplx/config.toml` using the `scheduler_mode` key, via CLI with `--scheduler-mode`, or programmatically via `ServerArgs`.
- The default mode is `serial`, enforced through [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) and propagated via [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py).
- Runtime validation in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) ensures mode-specific constraints are satisfied.

## Frequently Asked Questions

### What is the default scheduler mode in MTPLX?

The default scheduler mode is `serial`, which processes one request at a time using the solo MTP oracle. This default is defined in the CLI argument registration in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) and applies when no configuration override is provided.

### Can I change the scheduler mode without restarting the server?

No, the scheduler mode is determined at server initialization when the `ServerArgs` are processed and the backend command is constructed in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py). Changing modes requires restarting the MTPLX server with the new configuration.

### Why does my configuration fail when using balanced MTP?

The runtime enforces that balanced MTP configurations require `scheduler_mode="mtp_batch"`. If you receive validation errors in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) (lines 792-794), ensure your config file or CLI flag explicitly sets the mode to `mtp_batch` rather than the default `serial`.

### What is the difference between `ar_batch` and `mtp_batch`?

According to the `SchedulerMode` enum in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py), `ar_batch` configures a batched autoregressive decode lane for decode-heavy workloads, while `mtp_batch` configures a batched MTP decode lane optimized for prefill-heavy inference loads.