MTPLX Scheduler Modes Explained: Serial, MTP Batch, Hyper, AR Batch, and Solo MTP
MTPLX supports five distinct scheduler modes—serial, mtp_batch, hyper, ar_batch, and solo_mtp—that control how generation requests are batched and executed, configurable via the --scheduler-mode CLI flag or the scheduler_mode configuration key.
MTPLX is an open-source inference engine optimized for high-performance language model serving. Understanding the available scheduler modes in MTPLX is critical for tuning throughput and latency based on your workload characteristics. The scheduler determines how incoming requests are queued, batched, and executed across GPU resources, with implementation details found in the youssofal/MTPLX repository.
Overview of MTPLX Scheduler Modes
The scheduler architecture controls the batching strategy applied to inference requests. According to the source code in mtplx/batching/state.py, the system defines a SchedulerMode enum that validates runtime configurations parsed through mtplx/commands/public.py. Each mode implements distinct execution semantics ranging from deterministic single-request processing to aggressive multi-request batching strategies.
The Five Scheduler Modes in MTPLX
Serial Mode
Serial mode executes each request one-by-one without internal batching. This is the default scheduler mode in MTPLX, providing deterministic ordering and predictable per-request latency. Use this mode for debugging, simple use-cases, or when strict request isolation is required.
MTP Batch Mode
MTP batch mode groups tokens into fixed-width micro-batches (default width of 8) to improve throughput while preserving multi-token-per-step (MTP) semantics. This mode balances efficiency with the specialized requirements of MTP decoding, making it ideal for most large-language-model serving scenarios where throughput optimization is essential.
Hyper Mode
Hyper mode implements a hyper-chassis configuration where only a single external request is admitted at a time. The scheduler dedicates the entire batch width to that request, eliminating gather windows and inter-request contention. This mode maximizes throughput for single-client scenarios requiring dedicated GPU capacity without interference.
AR Batch Mode
AR batch mode aggregates multiple autoregressive (AR) decoding requests into a single batch. By parallelizing short AR sequences, this mode increases GPU utilization when serving many concurrent autoregressive generation tasks simultaneously.
Solo MTP Mode
Solo MTP mode runs a single MTP request without any batching overhead. This isolation is specifically designed for debugging or profiling the MTP pipeline without interference from other concurrent workloads.
How to Configure Scheduler Modes in MTPLX
You can specify the scheduler through command-line arguments, Python configuration objects, or by inspecting runtime state payloads.
Command-Line Configuration
Pass the --scheduler-mode flag when starting the MTPLX server:
# Run with the default serial scheduler
mtplx serve
# Enable MTP batching with micro-batch width 8
mtplx serve --scheduler-mode mtp_batch
# Use hyper mode for single-client dedication
mtplx serve --scheduler-mode hyper
# Enable AR batching for autoregressive workloads
mtplx serve --scheduler-mode ar_batch
Python Configuration Access
Access the current scheduler mode through the configuration object:
from mtplx.config import Config
cfg = Config.load()
print("Scheduler mode:", cfg.scheduler_mode) # e.g., "mtp_batch"
Runtime State Inspection
Inspect the scheduler state in request payloads:
payload = openai._mtplx_scheduler_state(state)
print(payload["scheduler_mode"]) # → "mtp_batch"
print(payload["scheduler_policy"]) # e.g., "fixed_mtp_batch_width_8"
Source Code Implementation Details
The scheduler mode validation logic resides in mtplx/commands/public.py, which parses the --scheduler-mode argument and validates it against the SchedulerMode enum defined in mtplx/batching/state.py. If an invalid value is supplied, MTPLX raises a validation error listing the allowed choices. The server initialization logic in mtplx/server/openai.py reads these configurations to instantiate the appropriate scheduler, while mtplx/config.py persists the scheduler_mode configuration key across sessions.
Summary
- MTPLX provides five scheduler modes: serial, mtp_batch, hyper, ar_batch, and solo_mtp.
- Serial is the default mode offering deterministic, single-request execution without batching.
- MTP batch optimizes throughput for multi-token-per-step workloads using fixed-width micro-batches.
- Hyper mode dedicates full batch capacity to a single client, eliminating contention and gather windows.
- AR batch parallelizes autoregressive requests for better GPU utilization during concurrent decoding.
- Solo MTP isolates individual MTP requests for debugging and pipeline profiling.
- Configuration occurs via the
--scheduler-modeCLI flag or thescheduler_modeconfig key, with validation enforced by the enum inmtplx/batching/state.py.
Frequently Asked Questions
What is the default scheduler mode in MTPLX?
The default scheduler mode is serial, which processes requests one-by-one without batching. This provides deterministic latency and is suitable for debugging or latency-sensitive applications where request ordering must be preserved.
How do I validate a scheduler mode configuration?
MTPLX validates the scheduler_mode value against the SchedulerMode enum defined in mtplx/batching/state.py. If you provide an invalid mode through the CLI or configuration file, the validation logic in mtplx/commands/public.py raises an error displaying the complete list of supported modes.
Can I switch scheduler modes at runtime?
Scheduler modes are set during server initialization via the --scheduler-mode flag or configuration file. While you cannot switch modes dynamically without restarting the server, you can inspect the current active mode at runtime through the scheduler state payload using openai._mtplx_scheduler_state().
Which scheduler mode offers the highest throughput?
Hyper mode typically offers the highest single-client throughput by dedicating the entire batch width to one request, eliminating gather windows entirely. For multi-client aggregate throughput, mtp_batch generally provides the best performance for multi-token-per-step workloads, while ar_batch optimizes concurrent autoregressive decoding scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →