MTPLX Concurrency Modes: Complete Guide to Scheduler Behaviors and Configuration

MTPLX supports six distinct scheduler modes—serial, cooperative, ar_batch, mtp_batch, mtp_cohort_experimental, and hyper—that control whether requests execute sequentially, interleave work, or batch decode operations while maintaining strict per-request isolation.

MTPLX concurrency modes separate model execution (which always runs on a single owner thread) from scheduling logic that governs how requests share computational resources. The chosen scheduler determines whether the backend processes requests one-at-a-time or batches them for improved throughput, with each mode offering specific trade-offs between latency, throughput, and hardware utilization. Understanding these behaviors is essential for optimizing inference performance based on your specific backend capabilities and workload patterns.

Available Scheduler Modes

MTPLX determines request scheduling at server startup via the --scheduler-mode flag. According to docs/concurrency.md, the following values are supported:

Serial Mode

serial runs a single request at a time on the backend’s normal generation route. This is the default mode when no scheduler is explicitly configured. While simple and predictable, this mode underutilizes hardware capable of parallel batch processing because each request monopolizes the compute width until completion.

Cooperative Mode

cooperative enables interleaving of independently owned request work when the backend supports cooperative execution. This mode allows the scheduler to pause and resume requests at safe points, improving throughput compared to serial execution without requiring full batching support from the underlying model implementation.

Autoregressive Batching

ar_batch batches autoregressive decode steps on compatible backends, allowing multiple requests to share a single forward pass while maintaining separate prompt histories and sampler states. This mode is ideal for throughput-sensitive applications where multiple clients submit requests simultaneously and the backend supports grouped matrix operations during the decode phase.

Native MTP Batching

mtp_batch leverages native-MTP (Multi-Token Prediction) decode using a model-specific installed MTP lane. Unlike autoregressive batching, this mode requires explicit model support for MTP lanes but can achieve higher throughput by predicting multiple future tokens in parallel across the batched requests.

Experimental Cohort Scheduling

mtp_cohort_experimental provides opt-in access to experimental native-MTP cohort scheduling algorithms. As noted in the source documentation, this mode is subject to change and should only be used for testing new scheduling heuristics that may eventually stabilize into production modes.

Hyper Mode

hyper is a special-case scheduler that serves one request at a time (with extra requests queuing FIFO like serial) but reserves the full batch width for speculative branches of that single request. No second client ever shares the width; instead, the width is dedicated to self-generated speculative tokens for the active request.

This mode imposes strict configuration constraints. When hyper is selected, the server disallows flags such as --max-active-requests, --decode-batch-max, positive --batch-wait-ms, and the agent or throughput presets. The health endpoint reports this mode under scheduler.hyper with fields showing the reserved width, queue depth, and receipts.

The Ownership Contract

Regardless of the selected concurrency mode, each admitted request retains exclusive ownership of its execution context. As implemented in mtplx/server/openai.py and documented in docs/concurrency.md, this contract guarantees that requests never read or mutate another request’s state even when sharing batched buffers.

Every request maintains its own:

  • Prompt and generated tokens
  • Logical KV cache and recurrent state
  • Target and draft sampler settings
  • Seeded random-number stream
  • Token budget, stop state, cancellation event, and output stream

Shared batched buffers are physically reused for efficiency, but row-specific masks, offsets, commits, and rewind mechanisms ensure strict isolation between concurrent requests.

Configuring Concurrency Modes

You can specify the scheduler mode via command-line flags or persistent configuration files.

CLI Configuration

Pass the mode directly when starting the server:

mtplx serve \
  --model <model-or-path> \
  --scheduler-mode <serial|cooperative|ar_batch|mtp_batch|mtp_cohort_experimental|hyper>

The mtplx/cli.py entry point parses this flag and forwards it to the server initialization logic in mtplx/server/openai.py.

Persistent Configuration

To avoid specifying the flag on every startup, persist the setting in the user configuration:

mtplx config set scheduler_mode mtp_batch
mtplx config set max_active_requests 8
mtplx config set decode_batch_max 8
mtplx config show --json

Configuration values are stored and validated by mtplx/config.py. Changing the scheduler requires stopping the server, updating the configuration, and restarting the service to apply the new scheduling behavior.

Verifying Concurrent Behavior

The health endpoint (/health) exposes telemetry that proves whether concurrent execution is active. For non-hyper modes, you should observe last_real_width ≥ 2 and a non-empty batch_histogram indicating that multiple requests shared a forward pass.

Query the current scheduler state:

curl -s http://127.0.0.1:8000/health | jq '.scheduler | {mode, active_lane, config, telemetry}'

For hyper mode, the response will show scheduler.hyper with the speculative width reserved for the single active request.

Practical Examples

Testing All Modes Sequentially

This bash script starts the server in each mode and verifies the active configuration:

for mode in serial cooperative ar_batch mtp_batch hyper; do
  echo "=== Starting MTPLX with $mode scheduler ==="
  mtplx serve --model <model> --scheduler-mode $mode &
  SERVER_PID=$!
  sleep 5
  curl -s http://127.0.0.1:8000/health | jq '.scheduler.mode'
  kill $SERVER_PID
done

Concurrent Client Testing

Use Python threading to stress-test batching modes with simultaneous requests:

import threading
import requests
import json

def send_req():
    payload = {
        "model": "gpt-4o-mini",
        "messages": [{"role": "user", "content": "Hello"}]
    }
    r = requests.post(
        "http://127.0.0.1:8000/v1/chat/completions",
        json=payload
    )
    print(r.json()["choices"][0]["message"]["content"])

threads = [threading.Thread(target=send_req) for _ in range(2)]
for t in threads:
    t.start()
for t in threads:
    t.join()

# Verify concurrency was exercised

health = requests.get("http://127.0.0.1:8000/health").json()
print(json.dumps(health["scheduler"], indent=2))

Summary

  • MTPLX concurrency modes determine how the scheduler allocates batch width across requests, ranging from single-threaded serial execution to speculative hyper mode.
  • Six schedulers are available: serial, cooperative, ar_batch, mtp_batch, mtp_cohort_experimental, and hyper, each optimized for different backend capabilities and latency requirements.
  • Strict ownership guarantees that batched requests never share logical state, maintaining separate KV caches, sampler settings, and RNG streams even during shared forward passes.
  • Configuration occurs via --scheduler-mode CLI flags or persistent settings in mtplx config, with changes requiring a server restart.
  • Validation is possible through the /health endpoint, which reports real-time batch width and scheduler telemetry.

Frequently Asked Questions

What is the difference between ar_batch and mtp_batch modes?

ar_batch performs autoregressive batching compatible with standard transformer backends, grouping decode steps from multiple requests into shared matrix operations. mtp_batch requires native Multi-Token Prediction lane support in the model architecture, allowing simultaneous prediction of multiple future tokens per request. While ar_batch works with most compatible backends, mtp_batch demands specific model installations but offers potentially higher throughput for MTP-capable models.

Why does hyper mode reject --max-active-requests flags?

hyper mode is designed specifically for speculative decoding where the batch width is reserved for the single active request's self-generated branches rather than for concurrent clients. Since the mode strictly limits external admission to one request, flags controlling multi-request concurrency like --max-active-requests or --decode-batch-max are incompatible and rejected by the argument validator in mtplx/config.py.

How can I confirm that requests are actually executing concurrently?

Query the health endpoint at http://127.0.0.1:8000/health and inspect scheduler.telemetry. Effective concurrency in non-hyper modes produces a last_real_width value greater than or equal to 2 and a populated batch_histogram showing historical batch sizes. If the system is falling back to serial behavior, these fields will show width 1 and empty histograms despite the configured scheduler mode.

Do I need to restart the server when changing scheduler modes?

Yes, changing the scheduler mode requires a full server restart. The mode is locked at initialization time in mtplx/server/openai.py based on the --scheduler-mode flag or persisted configuration value. Runtime switching is not supported because the scheduler selection determines fundamental buffer allocation strategies and thread pool structures that cannot be safely reconfigured while serving active requests.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →