MTPLX Scheduler Modes for Concurrent Batched Decoding: A Complete Guide

MTPLX defines six distinct scheduler modes—serial, cooperative, ar_batch, mtp_batch, mtp_cohort_experimental, and hyper—that control how requests are admitted, batched, and decoded, with ar_batch and mtp_batch specifically optimizing for high-throughput concurrent processing.

MTPLX is an open-source inference engine designed to optimize large language model serving through advanced batching strategies. Understanding the different scheduler modes for concurrent batched decoding is essential for maximizing throughput and minimizing latency in production deployments. The runtime behavior is governed by the SchedulerMode enum defined in mtplx/batching/state.py, which exposes granular control over how multiple requests share GPU resources during the decoding phase.

Overview of Scheduler Modes

The SchedulerMode enum in [mtplx/batching/state.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py) defines six distinct concurrency strategies. These modes determine whether requests are processed sequentially, interleaved, or grouped into batches for parallel token generation.

Serial Mode

Serial mode processes one request at a time with no batching overhead. According to the source code at line 21, SchedulerMode.SERIAL disables all batching logic, making it ideal for debugging scenarios or latency-critical workloads where deterministic, single-stream performance is required. This mode eliminates contention between requests but cannot leverage GPU parallelism for throughput gains.

Cooperative Mode

Cooperative admission allows the scheduler to interleave work across heterogeneous workloads without creating fixed-size batches. As implemented at line 22 in state.py, SchedulerMode.COOPERATIVE provides flexible scheduling that admits requests based on resource availability rather than strict batching constraints. This mode suits deployments with mixed request sizes where rigid batching would cause head-of-line blocking.

AR Batch Mode

AR Batch (Autoregressive Batching) groups multiple requests into a single decoding step where each request advances exactly one token. Defined at line 23 as SchedulerMode.AR_BATCH, this mode excels when serving many short generations simultaneously. The underlying implementation batches one token per request across the request pool, allowing the GPU kernel to process multiple sequences in parallel while maintaining simple autoregressive semantics.

MTP Batch Mode

MTP Batch (Multi-Token Prediction) represents the primary high-throughput configuration for concurrent batched decoding. Located at line 24 in mtplx/batching/state.py, SchedulerMode.MTP_BATCH enables a single request to generate multiple speculative tokens in a batch while still admitting other requests concurrently. This approach yields the highest throughput for long generations by amortizing kernel launch overhead across multiple tokens and requests simultaneously, leveraging underlying kernels like Flash-Attention and MQA.

MTP Cohort Experimental

MTP Cohort Experimental is an unstable variant that groups requests into “cohorts” to share speculative work. Defined at line 25 as SchedulerMode.MTP_COHORT_EXPERIMENTAL, this mode requires explicit opt-in via experimental flags and is not recommended for production use. It attempts to further increase throughput by allowing cross-request speculation, though the implementation remains under active development.

Hyper Mode

Hyper mode dedicates the entire batching machinery to a single external request (width = 1). Implemented at line 31 as SchedulerMode.HYPER, this mode behaves like serial admission but routes through the full batch pipeline, enabling precise profiling and benchmarking of solitary requests without interference from concurrent traffic.

Configuring Scheduler Modes

MTPLX exposes scheduler configuration through command-line flags parsed in mtplx/server/openai.py and mtplx/commands/public.py. You can specify the desired mode when launching the server to optimize for your specific throughput requirements.


# Launch with AR Batch mode for many short requests

import subprocess, shlex

cmd_ar = shlex.split(
    "python -m mtplx.server.openai serve --scheduler-mode ar_batch"
)
subprocess.Popen(cmd_ar)

# Launch with MTP Batch mode for high-throughput long generation

cmd_mtp = shlex.split(
    "python -m mtplx.server.openai serve --scheduler-mode mtp_batch"
)
subprocess.Popen(cmd_mtp)

To verify the active configuration programmatically, query the server status endpoint:

from mtplx.client import OpenAIClient

client = OpenAIClient(base_url="http://localhost:8000/v1")
resp = client.get("/status")
print("Active scheduler mode:", resp["scheduler_mode"])

Core Implementation Files

The scheduler architecture spans several key files that handle mode parsing, state management, and execution:

  • mtplx/batching/state.py – Contains the SchedulerMode enum definition and related scheduling primitives (lines 21-31).

  • mtplx/server/openai.py – Implements CLI flag parsing for --scheduler-mode and injects the selected mode into the runtime state.

  • mtplx/commands/public.py – Defines the public CLI interface (mtplx serve) and propagates scheduler arguments to the server initialization.

  • mtplx/server/mtp_batch.py – Houses the core implementation of the MTP batch scheduler, managing speculative token generation for mtp_batch mode.

  • mtplx/server/hyper.py – Contains specialized logic for the hyper mode, handling single-request dedicated batching for profiling scenarios.

Summary

  • MTPLX offers six scheduler modes defined in mtplx/batching/state.py: serial, cooperative, ar_batch, mtp_batch, mtp_cohort_experimental, and hyper.
  • AR Batch optimizes throughput for many short requests by batching one token per request across the request pool.
  • MTP Batch provides the highest throughput for long generations through speculative multi-token prediction while maintaining concurrent request admission.
  • Serial and Hyper modes isolate single requests for debugging or profiling, while Cooperative offers flexible interleaving without fixed batch sizes.
  • Configuration occurs via the --scheduler-mode CLI flag parsed in mtplx/server/openai.py and mtplx/commands/public.py.

Frequently Asked Questions

What is the difference between ar_batch and mtp_batch in MTPLX?

AR Batch processes multiple requests simultaneously but generates exactly one token per request per decoding step, making it efficient for short sequences. MTP Batch allows individual requests to generate multiple speculative tokens in parallel while still admitting other requests, maximizing throughput for longer generations by better utilizing GPU compute units.

Which scheduler mode should I use for production workloads?

For most production scenarios requiring concurrent batched decoding, mtp_batch provides the optimal balance of throughput and latency, particularly for long-form content generation. If your workload consists primarily of short queries (fewer than 50 tokens), ar_batch may offer more predictable latency characteristics.

How do I enable experimental cohort batching?

The mtp_cohort_experimental mode is gated behind experimental flags and not enabled by default. You must explicitly opt-in through environment variables or configuration flags parsed by mtplx/commands/public.py, as the implementation in mtplx/batching/state.py (line 25) remains unstable and subject to breaking changes.

Where is the scheduler mode state managed in the codebase?

The canonical definition resides in mtplx/batching/state.py where the SchedulerMode enum is declared. The active mode propagates through the system starting from CLI parsing in mtplx/commands/public.py, through server initialization in mtplx/server/openai.py, and finally to the specific executor implementations such as mtplx/server/mtp_batch.py or mtplx/server/hyper.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →