How to Configure the Scheduler Mode in MTPLX

You can configure the scheduler mode in MTPLX via the scheduler_mode key in ~/.mtplx/config.toml, the --scheduler-mode CLI flag, or the scheduler_mode parameter in ServerArgs objects, with valid options including serial, cooperative, ar_batch, mtp_batch, mtp_cohort_experimental, and hyper.

MTPLX is an open-source inference engine that optimizes request execution through a pluggable runtime scheduler. Configuring the scheduler mode determines how the engine batches and executes inference requests, directly impacting throughput and latency for workloads ranging from single-request scenarios to high-concurrency batches. This guide covers the configuration mechanisms and implementation details based on the youssofal/MTPLX source code.

Available Scheduler Modes in MTPLX

The available scheduling strategies are defined as a StrEnum in mtplx/batching/state.py (lines 18-33). Each mode corresponds to a concrete ModelWorkScheduler implementation:

  • serial: Solo MTP oracle processing one request at a time (default).
  • cooperative: Cooperative batching for mixed foreground and background workloads.
  • ar_batch: Batched autoregressive decode lane optimized for concurrent, decode-heavy requests.
  • mtp_batch: Batched MTP decode lane designed for prefill-heavy inference loads.
  • mtp_cohort_experimental: Experimental cohort-aware MTP scheduling for research scenarios.
  • hyper: Width-1 "hyper" chassis for single-request speculative decoding.

According to the implementation in mtplx/server/openai.py (lines 792-794), the runtime enforces mode-specific invariants, such as requiring mtp_batch for balanced MTP configurations.

Configuration Methods

You can configure the scheduler mode through three interfaces, all reading from the same UserConfig schema.

Config File

Create or edit ~/.mtplx/config.toml and set the scheduler_mode key:


# ~/.mtplx/config.toml

scheduler_mode = "mtp_batch"

The configuration loader in mtplx/config.py (lines 88-90) normalizes this value using _str_or_none and stores it in UserConfig.scheduler_mode.

CLI Flag

Override the config file setting using the --scheduler-mode flag available in commands like mtplx serve and mtplx start:

mtplx serve --scheduler-mode ar_batch

The argument is registered in mtplx/cli.py (lines 65-78) with "serial" as the default value and choices limited to the SchedulerMode enum values.

Programmatic API

When building server arguments programmatically, pass scheduler_mode as a string parameter to ServerArgs or OpenAIArgs:

from mtplx.commands.public import ServerArgs

args = ServerArgs(
    model="models/example",
    scheduler_mode="hyper"
)

The UserConfig object processes this parameter identically to CLI and file-based configuration.

Technical Implementation Details

Configuration Propagation

When MTPLX builds the backend server command in mtplx/commands/public.py (lines 11523-11524), it conditionally appends --scheduler-mode <mode> to the generated command unless the mode matches the default serial. This ensures the scheduler selection propagates to the runtime environment.

Runtime Enforcement

Inside the server implementation, specifically in mtplx/server/openai.py and mtplx/server/hyper.py, the scheduler mode string determines the ModelWorkScheduler configuration. The code validates that specific features, such as balanced MTP batching, require the corresponding scheduler_mode setting.

Selecting the Right Scheduler Mode

Choose your scheduler mode based on traffic patterns and latency requirements:

  • serial: Use for lowest-latency single-request workloads or when debugging inference behavior.
  • cooperative: Best for mixed workloads requiring cooperative batching without AR-specific optimizations.
  • ar_batch: Ideal for high-throughput, decode-heavy traffic with concurrent requests.
  • mtp_batch: Optimized for prefill-heavy loads where batching MTP decode steps improves throughput.
  • mtp_cohort_experimental: Deploy for research into cohort-aware MTP scheduling strategies.
  • hyper: Suitable for single-request speculative decoding scenarios requiring width-1 chassis.

Summary

  • MTPLX supports six scheduler modes defined in mtplx/batching/state.py: serial, cooperative, ar_batch, mtp_batch, mtp_cohort_experimental, and hyper.
  • Configure via ~/.mtplx/config.toml using the scheduler_mode key, via CLI with --scheduler-mode, or programmatically via ServerArgs.
  • The default mode is serial, enforced through mtplx/cli.py and propagated via mtplx/commands/public.py.
  • Runtime validation in mtplx/server/openai.py ensures mode-specific constraints are satisfied.

Frequently Asked Questions

What is the default scheduler mode in MTPLX?

The default scheduler mode is serial, which processes one request at a time using the solo MTP oracle. This default is defined in the CLI argument registration in mtplx/cli.py and applies when no configuration override is provided.

Can I change the scheduler mode without restarting the server?

No, the scheduler mode is determined at server initialization when the ServerArgs are processed and the backend command is constructed in mtplx/commands/public.py. Changing modes requires restarting the MTPLX server with the new configuration.

Why does my configuration fail when using balanced MTP?

The runtime enforces that balanced MTP configurations require scheduler_mode="mtp_batch". If you receive validation errors in mtplx/server/openai.py (lines 792-794), ensure your config file or CLI flag explicitly sets the mode to mtp_batch rather than the default serial.

What is the difference between ar_batch and mtp_batch?

According to the SchedulerMode enum in mtplx/batching/state.py, ar_batch configures a batched autoregressive decode lane for decode-heavy workloads, while mtp_batch configures a batched MTP decode lane optimized for prefill-heavy inference loads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →