# How Concurrency Scheduler Modes in MTPLX Work: From Serial to Pipelined Execution

> Explore MTPLX concurrency scheduler modes from serial to pipelined execution. Learn how single-thread strategies optimize performance without parallel threads.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: internals
- Published: 2026-09-13

---

**MTPLX implements concurrency not through parallel threads but through admission policies and batching strategies on a single-owner thread, with scheduler modes ranging from strict serial execution to fixed-width pipelined micro-batches.**

The MTPLX inference server (youssofal/MTPLX) runs all model work on a **single-owner thread** to maintain compatibility with the thread-affine MLX runtime on Apple Silicon. Rather than using parallel threads, the system achieves concurrency through a set of **scheduler modes** that control admission gates and batch formation. These modes determine how many requests may enter the foreground work band simultaneously and how individual decode steps are grouped into micro-batches.

## The Single-Owner Thread Architecture

All heavy MLX computation in MTPLX executes on one dedicated thread because the MLX runtime requires thread affinity. In [`mtplx/model_scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_scheduler.py), the `ModelWorkScheduler` owns three distinct work bands:

- **Foreground**: User-visible token generation
- **Idle-post-commit**: Snapshot work that occurs after request completion
- **Idle-persistence**: Durability operations for long-term storage

A tiny scheduler loop decides which band runs next, with **foreground priority** always taking precedence over idle work. Concurrency emerges not from parallel execution but from the **admission policies** that govern how many requests occupy the foreground band and how the scheduler batches their operations together.

## MTPLX Scheduler Modes Explained

The runtime exposes six distinct scheduler modes defined in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py) via the `SchedulerMode` enum. Each mode plugs a different admission-and-batching strategy into the single-owner loop.

### serial

The `serial` mode represents the baseline admission policy. Exactly **one request** occupies the foreground band at any time, with additional requests queuing in FIFO order. No batching width is applied—every request executes strictly sequentially on the model-owner thread.

Implementation resides in [`mtplx/model_scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_scheduler.py) within `ModelWorkScheduler`, which enforces foreground priority without grouping operations.

### cooperative and ar_batch

The `cooperative` mode (aliased as `ar_batch` for clarity) implements the classic "AR-batch" path inspired by vLLM-style servers. While requests enter the foreground one-by-one, the scheduler may **group multiple decode steps** into a micro-batch using the `batch_histogram` telemetry field.

This creates **micro-batching of autoregressive steps**; though only one request initiates per turn, several tokens may generate in a single batch when requests become ready simultaneously. The logic lives in `record_batch_step` and `_take_next`, which determines when idle work becomes runnable and drains the batch.

### mtp_batch

The `mtp_batch` mode enables "Multi-Turn-Pipeline" execution. The scheduler creates **fixed-width micro-batches** (default width of 8) that pipeline across turns, processing several requests in lock-step. This mode trades latency for throughput by admitting multiple requests simultaneously.

Implementation appears in [`mtplx/server/mtp_batch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/mtp_batch.py), which adds `"scheduler_mode": "mtp_batch"` to the health payload and manages the fixed-width admission gate. The mode also supports throughput presets via `SchedulerPreset` configurations.

### mtp_cohort_experimental

An experimental variant of the pipelined approach, `mtp_cohort_experimental` clusters requests into **cohorts** before forming batches. This mode enables research into adaptive cohort formation, processing each cluster as a distinct batch unit. It shares implementation files with `mtp_batch` but uses the mode string `"mtp_cohort_experimental"` to trigger cohort-aware logic.

### hyper

The `hyper` mode operates as a **singleton-passthrough** with internal speculation. Only one request may be in-flight externally, but the scheduler generates **speculative rows** for that request through self-generated micro-batches. This creates a width-1 batch externally that can achieve internal width greater than one for self-generated rows.

Implementation resides in [`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py) via `HyperAdmissionGate` and the `HyperWidthExecutor` seam. The `HyperWidthExecutor.plan` method determines speculative width, while `HyperAdmissionGate.bind` falls back to the serial path when width equals one.

## How Batching Controls Concurrency

Concurrency behavior depends on how each mode implements **admission gates** and **batch telemetry**:

- **Admission gates** control entry into the foreground band. `HyperAdmissionGate` in `hyper` mode admits only singletons, while `mtp_batch` admits up to the fixed width.
- **Batch histograms** track micro-batch sizes via `ModelWorkScheduler.record_batch_step`. This telemetry allows the health endpoint to report actual throughput characteristics per mode.
- **Persistence pump**: In `mtp_batch` mode, the `_persistence_pump_budget` prevents starvation by allowing durability work to run even when idle bands contain pending operations.

## Configuring and Monitoring Scheduler Modes

Select your concurrency strategy via the `--scheduler-mode` CLI argument parsed in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py). The server wires the appropriate admission gate at startup based on your selection.

Monitor runtime behavior through the health endpoint:

```python

# Query scheduler telemetry during operation

import requests

resp = requests.get("http://localhost:8080/health")
stats = resp.json()["scheduler"]

print("Current mode:", stats["scheduler_policy"])
print("Batch histogram:", stats["batch_histogram"])

```

Start a server with a specific mode:

```python

# Example: launch in MTP-batch mode

from mtplx.commands.public import build_parser

args = build_parser().parse_args(["serve", "--scheduler-mode", "mtp_batch"])

# args.scheduler_mode == "mtp_batch"

```

For advanced use cases, interact directly with the admission gates:

```python

# Direct usage of HyperAdmissionGate (advanced)

from mtplx.server.hyper import HyperAdmissionGate, HyperTicket

gate = HyperAdmissionGate()
ticket = gate.admit(request_id="req1", prompt_tokens=256)

def serial_closure():
    return {"tokens": [1, 2, 3]}

bound = gate.bind(ticket, serial_closure, generation_mode="mtp")
result = bound()  # Executes serial path or width>1 plan

```

## Summary

- MTPLX uses a **single-owner thread** for all MLX work due to runtime thread affinity, implementing concurrency through admission policies rather than parallel threads.
- **Six scheduler modes** provide distinct concurrency behaviors: `serial` (strictly sequential), `cooperative`/`ar_batch` (dynamic micro-batching), `mtp_batch` (fixed-width pipelining), `mtp_cohort_experimental` (cohort clustering), and `hyper` (speculative singleton).
- **Admission gates** like `HyperAdmissionGate` and `ModelWorkScheduler` control how many requests enter the foreground band, while `batch_histogram` telemetry tracks actual micro-batch sizes.
- Configuration occurs via `--scheduler-mode` CLI arguments parsed in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py), with runtime monitoring available through the health endpoint in [`mtplx/server/mtp_batch.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/mtp_batch.py).

## Frequently Asked Questions

### What is the default scheduler mode in MTPLX?

The default mode depends on your server configuration, but `serial` provides the baseline strictly sequential execution. Most production deployments select either `cooperative` for dynamic batching or `mtp_batch` for throughput-oriented fixed-width pipelining.

### Why does MTPLX use a single-owner thread instead of parallel threads?

The MLX runtime on Apple Silicon maintains thread affinity that makes parallel thread execution unsafe or inefficient. By funneling all model work through a single-owner thread in [`mtplx/model_scheduler.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/model_scheduler.py), MTPLX ensures safe interaction with the MLX runtime while achieving concurrency through sophisticated admission policies and batch formation.

### How does the `hyper` mode differ from `serial` mode?

While both modes admit only one external request at a time, `hyper` mode can generate **speculative rows** internally, creating self-generated micro-batches with width greater than one. This happens through the `HyperWidthExecutor` seam, whereas `serial` mode processes exactly one token generation step per scheduling turn with no internal parallelism.

### Can I change the scheduler mode without restarting the server?

No, the scheduler mode is wired at startup when the CLI parser in [`mtplx/commands/public.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/commands/public.py) instantiates the appropriate admission gate. Changing modes requires restarting the server with a different `--scheduler-mode` argument.