What Are the Concurrency Modes in MTPLX? AR, MTP, and Hyper Explained
MTPLX implements three distinct concurrency pipelines: the Auto-Regressive (AR) lane for single-prompt FIFO generation, the Multi-Token-Parallel (MTP) batch mode for grouped high-throughput inference with prefill limits, and the Hyper mode for strict definitional FIFO scheduling without tunable width parameters.
The MTPLX inference server by youssofal/MTPLX provides specialized concurrency modes to optimize request scheduling for different latency and throughput requirements. Understanding these modes—AR, MTP, and Hyper—is essential for configuring the server to handle everything from interactive chat to large batch workloads efficiently.
The Three Concurrency Modes in MTPLX
MTPLX routes requests through three distinct lanes based on workload characteristics. The concurrency-adaptive dispatcher selects the appropriate pipeline automatically, though you can override this behavior via client configuration.
| Concurrency Mode | Description | Key Constraint | Source Location |
|---|---|---|---|
| AR Lane | Default single-prompt generation | FIFO execution, adaptive concurrency | [mtplx/server/openai.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py#L3120) |
| MTP Batch | Multi-token parallel batch processing | Prefill concurrency capped at 2 | [mtplx/mtp_batch_numerics.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_batch_numerics.py#L9) |
| Hyper | Definitional external scheduling | Non-tunable width, strict FIFO | [mtplx/server/hyper.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py#L49) |
AR (Auto-Regressive) Lane
The AR lane serves as the default mode for handling single-prompt generation requests. As implemented in [mtplx/server/openai.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py#L3120), this mode maintains a strict FIFO queue where the model runs step-by-step, "pumping on host round-trips under concurrency" to yield tokens as they become available. When real concurrency is detected, the AR lane shares the same concurrency-adaptive lane as batched workloads, ensuring that resource limits are enforced while preserving per-request ordering.
MTP (Multi-Token-Parallel) Batch Mode
For high-throughput scenarios, MTP batch mode aggregates multiple prompts into a single batch and generates tokens in parallel. The implementation in [mtplx/mtp_batch_numerics.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_batch_numerics.py#L9) exposes the public MTP concurrency routes and explicitly caps prefill concurrency to 2 regardless of decode worker availability, keeping memory utilization predictable. This mode allows higher parallelism during the decode phase, making it the primary path for processing bulk completions.
Hyper Mode
The Hyper mode implements an external concurrency model that treats scheduling as definitional rather than tunable. According to the source in [mtplx/server/hyper.py](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py#L49), this mode devotes the entire system width to the hyper scheduler and enforces a strict FIFO order across all incoming requests. Unlike the AR and MTP lanes, Hyper mode does not expose width parameters, guaranteeing that no request is starved even under high connection volume.
Configuring Concurrency Modes in MTPLX
You can select the appropriate concurrency mode through the Python client, direct HTTP requests, or OpenAI-compatible API calls.
Python Client Configuration
Use the concurrency_mode parameter when initializing MTPLXClient to specify "mtp", "ar", or "hyper":
from mtplx import MTPLXClient
# Initialize client for MTP batch mode with explicit prefill limits
client = MTPLXClient(
base_url="http://localhost:8000",
model="qwen2.5-7b-mtp",
concurrency_mode="mtp", # Options: "ar" | "mtp" | "hyper"
max_batch_size=8,
)
resp = client.batch_generate(
prompts=[
"Explain quantum entanglement in simple terms.",
"Summarize the plot of Inception.",
]
)
print(resp)
Command-Line Invocation
For Hyper mode endpoints, use direct HTTP POST requests:
curl -X POST http://localhost:8000/v1/hyper/completions \
-H "Content-Type: application/json" \
-d '{"model":"qwen2.5-7b","prompt":"Write a haiku about sunrise."}'
OpenAI-Compatible API
When using the OpenAI-compatible endpoint, the server automatically selects the optimal lane based on request volume:
import openai
openai.api_base = "http://localhost:8000/v1"
openai.api_key = "placeholder"
response = openai.ChatCompletion.create(
model="qwen2.5-7b",
messages=[{"role": "user", "content": "What is the capital of France?"}]
)
print(response.choices[0].message["content"])
Summary
- AR Lane: Default FIFO mode for single prompts; shares adaptive concurrency with batched workloads as implemented in
mtplx/server/openai.py. - MTP Batch: High-throughput mode with explicit prefill concurrency capped at 2; defined in
mtplx/mtp_batch_numerics.py. - Hyper Mode: Definitional concurrency with strict FIFO and no tunable width; source comments in
mtplx/server/hyper.pyspecify this behavior. - Configuration: Use the
concurrency_modeparameter inMTPLXClientor rely on automatic dispatcher routing via the OpenAI-compatible API.
Frequently Asked Questions
What is the difference between AR and MTP concurrency modes?
The AR lane processes requests sequentially in FIFO order with step-by-step token generation, making it ideal for interactive latency-sensitive applications. The MTP batch mode groups requests to generate multiple tokens in parallel, capping prefill concurrency at 2 to manage memory while maximizing decode throughput for bulk operations.
Why is prefill concurrency limited to 2 in MTP mode?
The prefill concurrency cap of 2 in MTP mode, as defined in mtplx/mtp_batch_numerics.py, prevents excessive memory utilization during the initial prompt processing phase while still allowing higher parallelism during token generation. This trade-off ensures predictable resource usage under heavy batch loads.
Can I tune the concurrency width in Hyper mode?
No. According to the implementation in mtplx/server/hyper.py, Hyper mode treats concurrency as definitional rather than tunable, meaning the system dedicates its full width to maintaining strict FIFO ordering without exposing configurable concurrency parameters to the user.
Which mode should I use for high-throughput batch processing?
Use MTP batch mode by setting concurrency_mode="mtp" in your client configuration. This mode is explicitly designed for high-throughput scenarios, providing optimized parallelism during decode phases while enforcing the prefill limit to maintain system stability.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →