What Is the Default Concurrency Mode in MTPLX?
MTPLX defaults to serial concurrency mode, processing one request at a time unless explicitly configured otherwise via the --scheduler-mode flag or API configuration.
When deploying inference servers with the youssofal/MTPLX repository, understanding the scheduler behavior is critical for performance tuning. The codebase explicitly defines serial as the fallback concurrency strategy across both documentation and command-line interfaces, ensuring predictable single-threaded execution out of the box.
Source Code Evidence for the Serial Default
Documentation Declaration
According to docs/concurrency.md lines 15-18, the documentation lists serial as "Run one request at a time" and explicitly marks it as "this is the default"【/cache/repos/github.com/youssofal/MTPLX/main/docs/concurrency.md#L15-L18】. This declaration serves as the primary reference for users configuring the scheduler.
CLI Argument Parser Implementation
The command-line interface reinforces this default in mtplx/cli.py lines 5023-5028, where the --scheduler-mode argument defaults to "serial" when the flag is omitted【/cache/repos/github.com/youssofal/MTPLX/main/mtplx/cli.py#L5 023‑L5 028】. This ensures consistent behavior whether users interact via CLI or Python API.
Alternative Concurrency Modes
While serial is the default, MTPLX supports several alternative schedulers for high-throughput scenarios:
ar_batch- Automatic Regressive batchingmtp_batch- Multi-Token Prediction batchinghyper- Hyper-parallel processing
These modes must be explicitly specified to override the serial fallback.
Practical Configuration Examples
Starting the Server with Defaults
To launch the server using the implicit serial mode:
mtplx serve --model Qwen3.6-35B
Explicitly Selecting a Different Mode
Override the default by specifying an alternative scheduler:
mtplx serve --model Qwen3.6-35B --scheduler-mode mtp_batch
Inspecting the Default Programmatically
Access the default value through the SchedulerMode enum defined in mtplx/batching/state.py:
from mtplx.batching.state import SchedulerMode
print(SchedulerMode.serial.value) # → "serial"
Summary
- Serial concurrency is the hardcoded default across MTPLX documentation and CLI infrastructure.
- The default is declared in
docs/concurrency.mdand enforced inmtplx/cli.pylines 5023-5028. - Users must explicitly pass
--scheduler-modeto enable batch processing or parallel execution. - The
SchedulerModeenum inmtplx/batching/state.pyprovides programmatic access to configuration values.
Frequently Asked Questions
What is the default concurrency mode in MTPLX?
MTPLX uses serial mode by default, meaning it processes one request at a time. This is explicitly documented in docs/concurrency.md and hardcoded into the CLI argument parser in mtplx/cli.py lines 5023-5028.
How do I switch from serial to batch processing?
Pass the --scheduler-mode flag when starting the server, such as --scheduler-mode mtp_batch or --scheduler-mode ar_batch. Alternatively, configure the scheduler through the Python API by setting the mode parameter to the desired SchedulerMode enum value.
What are the performance implications of serial mode?
Serial mode ensures deterministic, sequential processing but limits throughput to single-request handling. For high-traffic deployments, switching to mtp_batch or hyper modes enables parallel processing and improved GPU utilization.
Where is the default value stored in the codebase?
The default string "serial" is defined in two critical locations: the user-facing documentation at docs/concurrency.md lines 15-18, and the CLI implementation at mtplx/cli.py lines 5023-5028. The underlying enum definition resides in mtplx/batching/state.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →