# What Is the Default Concurrency Mode in MTPLX?

> Discover the default concurrency mode in MTPLX. Learn how MTPLX processes requests serially by default and how to change it for better performance.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: deep-dive
- Published: 2026-09-02

---

**MTPLX defaults to serial concurrency mode**, processing one request at a time unless explicitly configured otherwise via the `--scheduler-mode` flag or API configuration.

When deploying inference servers with the **youssofal/MTPLX** repository, understanding the scheduler behavior is critical for performance tuning. The codebase explicitly defines **serial** as the fallback concurrency strategy across both documentation and command-line interfaces, ensuring predictable single-threaded execution out of the box.

## Source Code Evidence for the Serial Default

### Documentation Declaration

According to [`docs/concurrency.md`](https://github.com/youssofal/MTPLX/blob/main/docs/concurrency.md) lines 15-18, the documentation lists `serial` as "Run one request at a time" and explicitly marks it as **"this is the default"**【/cache/repos/github.com/youssofal/MTPLX/main/docs/concurrency.md#L15-L18】. This declaration serves as the primary reference for users configuring the scheduler.

### CLI Argument Parser Implementation

The command-line interface reinforces this default in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) lines 5023-5028, where the `--scheduler-mode` argument defaults to `"serial"` when the flag is omitted【/cache/repos/github.com/youssofal/MTPLX/main/mtplx/cli.py#L5 023‑L5 028】. This ensures consistent behavior whether users interact via CLI or Python API.

## Alternative Concurrency Modes

While **serial** is the default, MTPLX supports several alternative schedulers for high-throughput scenarios:

- **`ar_batch`** - Automatic Regressive batching
- **`mtp_batch`** - Multi-Token Prediction batching
- **`hyper`** - Hyper-parallel processing

These modes must be explicitly specified to override the serial fallback.

## Practical Configuration Examples

### Starting the Server with Defaults

To launch the server using the implicit serial mode:

```bash
mtplx serve --model Qwen3.6-35B

```

### Explicitly Selecting a Different Mode

Override the default by specifying an alternative scheduler:

```bash
mtplx serve --model Qwen3.6-35B --scheduler-mode mtp_batch

```

### Inspecting the Default Programmatically

Access the default value through the `SchedulerMode` enum defined in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py):

```python
from mtplx.batching.state import SchedulerMode

print(SchedulerMode.serial.value)   # → "serial"

```

## Summary

- **Serial concurrency** is the hardcoded default across MTPLX documentation and CLI infrastructure.
- The default is declared in [`docs/concurrency.md`](https://github.com/youssofal/MTPLX/blob/main/docs/concurrency.md) and enforced in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) lines 5023-5028.
- Users must explicitly pass `--scheduler-mode` to enable batch processing or parallel execution.
- The `SchedulerMode` enum in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py) provides programmatic access to configuration values.

## Frequently Asked Questions

### What is the default concurrency mode in MTPLX?

MTPLX uses **serial** mode by default, meaning it processes one request at a time. This is explicitly documented in [`docs/concurrency.md`](https://github.com/youssofal/MTPLX/blob/main/docs/concurrency.md) and hardcoded into the CLI argument parser in [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) lines 5023-5028.

### How do I switch from serial to batch processing?

Pass the `--scheduler-mode` flag when starting the server, such as `--scheduler-mode mtp_batch` or `--scheduler-mode ar_batch`. Alternatively, configure the scheduler through the Python API by setting the mode parameter to the desired `SchedulerMode` enum value.

### What are the performance implications of serial mode?

Serial mode ensures deterministic, sequential processing but limits throughput to single-request handling. For high-traffic deployments, switching to `mtp_batch` or `hyper` modes enables parallel processing and improved GPU utilization.

### Where is the default value stored in the codebase?

The default string `"serial"` is defined in two critical locations: the user-facing documentation at [`docs/concurrency.md`](https://github.com/youssofal/MTPLX/blob/main/docs/concurrency.md) lines 15-18, and the CLI implementation at [`mtplx/cli.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cli.py) lines 5023-5028. The underlying enum definition resides in [`mtplx/batching/state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/batching/state.py).