# MTPLX Scheduler Modes for Concurrent Batched Decoding: A Complete Guide

> Explore MTPLX scheduler modes for concurrent batched decoding. Learn about serial, cooperative, ar_batch, mtp_batch, and more to optimize high throughput processing.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: how-to-guide
- Published: 2026-09-08

---

**MTPLX defines six distinct scheduler modes—serial, cooperative, ar_batch, mtp_batch, mtp_cohort_experimental, and hyper—that control how requests are admitted, batched, and decoded, with ar_batch and mtp_batch specifically optimizing for high-throughput concurrent processing.**

MTPLX is an open-source inference engine designed to optimize large language model serving through advanced batching strategies. Understanding the different **scheduler modes for concurrent batched decoding** is essential for maximizing throughput and minimizing latency in production deployments. The runtime behavior is governed by the `SchedulerMode` enum defined in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py), which exposes granular control over how multiple requests share GPU resources during the decoding phase.

## Overview of Scheduler Modes

The `SchedulerMode` enum in [[`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py) defines six distinct concurrency strategies. These modes determine whether requests are processed sequentially, interleaved, or grouped into batches for parallel token generation.

### Serial Mode

**Serial** mode processes one request at a time with no batching overhead. According to the source code at line 21, `SchedulerMode.SERIAL` disables all batching logic, making it ideal for debugging scenarios or latency-critical workloads where deterministic, single-stream performance is required. This mode eliminates contention between requests but cannot leverage GPU parallelism for throughput gains.

### Cooperative Mode

**Cooperative** admission allows the scheduler to interleave work across heterogeneous workloads without creating fixed-size batches. As implemented at line 22 in [`state.py`](https://github.com/youssofal/MTPLX/blob/main/state.py), `SchedulerMode.COOPERATIVE` provides flexible scheduling that admits requests based on resource availability rather than strict batching constraints. This mode suits deployments with mixed request sizes where rigid batching would cause head-of-line blocking.

### AR Batch Mode

**AR Batch** (Autoregressive Batching) groups multiple requests into a single decoding step where each request advances exactly one token. Defined at line 23 as `SchedulerMode.AR_BATCH`, this mode excels when serving many short generations simultaneously. The underlying implementation batches one token per request across the request pool, allowing the GPU kernel to process multiple sequences in parallel while maintaining simple autoregressive semantics.

### MTP Batch Mode

**MTP Batch** (Multi-Token Prediction) represents the primary high-throughput configuration for concurrent batched decoding. Located at line 24 in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py), `SchedulerMode.MTP_BATCH` enables a single request to generate multiple speculative tokens in a batch while still admitting other requests concurrently. This approach yields the highest throughput for long generations by amortizing kernel launch overhead across multiple tokens and requests simultaneously, leveraging underlying kernels like Flash-Attention and MQA.

### MTP Cohort Experimental

**MTP Cohort Experimental** is an unstable variant that groups requests into “cohorts” to share speculative work. Defined at line 25 as `SchedulerMode.MTP_COHORT_EXPERIMENTAL`, this mode requires explicit opt-in via experimental flags and is not recommended for production use. It attempts to further increase throughput by allowing cross-request speculation, though the implementation remains under active development.

### Hyper Mode

**Hyper** mode dedicates the entire batching machinery to a single external request (width = 1). Implemented at line 31 as `SchedulerMode.HYPER`, this mode behaves like serial admission but routes through the full batch pipeline, enabling precise profiling and benchmarking of solitary requests without interference from concurrent traffic.

## Configuring Scheduler Modes

MTPLX exposes scheduler configuration through command-line flags parsed in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) and [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py). You can specify the desired mode when launching the server to optimize for your specific throughput requirements.

```python

# Launch with AR Batch mode for many short requests

import subprocess, shlex

cmd_ar = shlex.split(
    "python -m mtplx.server.openai serve --scheduler-mode ar_batch"
)
subprocess.Popen(cmd_ar)

# Launch with MTP Batch mode for high-throughput long generation

cmd_mtp = shlex.split(
    "python -m mtplx.server.openai serve --scheduler-mode mtp_batch"
)
subprocess.Popen(cmd_mtp)

```

To verify the active configuration programmatically, query the server status endpoint:

```python
from mtplx.client import OpenAIClient

client = OpenAIClient(base_url="http://localhost:8000/v1")
resp = client.get("/status")
print("Active scheduler mode:", resp["scheduler_mode"])

```

## Core Implementation Files

The scheduler architecture spans several key files that handle mode parsing, state management, and execution:

- **[`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py)** – Contains the `SchedulerMode` enum definition and related scheduling primitives (lines 21-31).

- **[`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py)** – Implements CLI flag parsing for `--scheduler-mode` and injects the selected mode into the runtime state.

- **[`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py)** – Defines the public CLI interface (`mtplx serve`) and propagates scheduler arguments to the server initialization.

- **[`mtplx/server/mtp_batch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/mtp_batch.py)** – Houses the core implementation of the MTP batch scheduler, managing speculative token generation for `mtp_batch` mode.

- **[`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py)** – Contains specialized logic for the `hyper` mode, handling single-request dedicated batching for profiling scenarios.

## Summary

- **MTPLX offers six scheduler modes** defined in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py): serial, cooperative, ar_batch, mtp_batch, mtp_cohort_experimental, and hyper.
- **AR Batch** optimizes throughput for many short requests by batching one token per request across the request pool.
- **MTP Batch** provides the highest throughput for long generations through speculative multi-token prediction while maintaining concurrent request admission.
- **Serial and Hyper** modes isolate single requests for debugging or profiling, while **Cooperative** offers flexible interleaving without fixed batch sizes.
- Configuration occurs via the `--scheduler-mode` CLI flag parsed in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) and [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py).

## Frequently Asked Questions

### What is the difference between ar_batch and mtp_batch in MTPLX?

**AR Batch** processes multiple requests simultaneously but generates exactly one token per request per decoding step, making it efficient for short sequences. **MTP Batch** allows individual requests to generate multiple speculative tokens in parallel while still admitting other requests, maximizing throughput for longer generations by better utilizing GPU compute units.

### Which scheduler mode should I use for production workloads?

For most production scenarios requiring **concurrent batched decoding**, `mtp_batch` provides the optimal balance of throughput and latency, particularly for long-form content generation. If your workload consists primarily of short queries (fewer than 50 tokens), `ar_batch` may offer more predictable latency characteristics.

### How do I enable experimental cohort batching?

The `mtp_cohort_experimental` mode is gated behind experimental flags and not enabled by default. You must explicitly opt-in through environment variables or configuration flags parsed by [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py), as the implementation in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py) (line 25) remains unstable and subject to breaking changes.

### Where is the scheduler mode state managed in the codebase?

The canonical definition resides in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py) where the `SchedulerMode` enum is declared. The active mode propagates through the system starting from CLI parsing in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py), through server initialization in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), and finally to the specific executor implementations such as [`mtplx/server/mtp_batch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/mtp_batch.py) or [`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py).