# What Are the Concurrency Modes in MTPLX? AR, MTP, and Hyper Explained

> Explore MTPLX concurrency modes: AR for single-prompt FIFO, MTP for high-throughput batching, and Hyper for strict FIFO. Understand the pipelines for efficient inference.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: internals
- Published: 2026-09-02

---

**MTPLX implements three distinct concurrency pipelines: the Auto-Regressive (AR) lane for single-prompt FIFO generation, the Multi-Token-Parallel (MTP) batch mode for grouped high-throughput inference with prefill limits, and the Hyper mode for strict definitional FIFO scheduling without tunable width parameters.**

The MTPLX inference server by `youssofal/MTPLX` provides specialized concurrency modes to optimize request scheduling for different latency and throughput requirements. Understanding these modes—**AR**, **MTP**, and **Hyper**—is essential for configuring the server to handle everything from interactive chat to large batch workloads efficiently.

## The Three Concurrency Modes in MTPLX

MTPLX routes requests through three distinct lanes based on workload characteristics. The concurrency-adaptive dispatcher selects the appropriate pipeline automatically, though you can override this behavior via client configuration.

| Concurrency Mode | Description | Key Constraint | Source Location |
|------------------|-------------|----------------|-----------------|
| **AR Lane** | Default single-prompt generation | FIFO execution, adaptive concurrency | [[`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py#L3120) |
| **MTP Batch** | Multi-token parallel batch processing | Prefill concurrency capped at 2 | [[`mtplx/mtp_batch_numerics.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_batch_numerics.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_batch_numerics.py#L9) |
| **Hyper** | Definitional external scheduling | Non-tunable width, strict FIFO | [[`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py#L49) |

### AR (Auto-Regressive) Lane

The **AR lane** serves as the default mode for handling single-prompt generation requests. As implemented in [[`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py#L3120), this mode maintains a strict FIFO queue where the model runs step-by-step, "pumping on host round-trips under concurrency" to yield tokens as they become available. When real concurrency is detected, the AR lane shares the same concurrency-adaptive lane as batched workloads, ensuring that resource limits are enforced while preserving per-request ordering.

### MTP (Multi-Token-Parallel) Batch Mode

For high-throughput scenarios, **MTP batch mode** aggregates multiple prompts into a single batch and generates tokens in parallel. The implementation in [[`mtplx/mtp_batch_numerics.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_batch_numerics.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_batch_numerics.py#L9) exposes the public MTP concurrency routes and explicitly caps **prefill concurrency to 2** regardless of decode worker availability, keeping memory utilization predictable. This mode allows higher parallelism during the decode phase, making it the primary path for processing bulk completions.

### Hyper Mode

The **Hyper mode** implements an external concurrency model that treats scheduling as definitional rather than tunable. According to the source in [[`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py)](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py#L49), this mode devotes the entire system width to the hyper scheduler and enforces a strict FIFO order across all incoming requests. Unlike the AR and MTP lanes, Hyper mode does not expose width parameters, guaranteeing that no request is starved even under high connection volume.

## Configuring Concurrency Modes in MTPLX

You can select the appropriate concurrency mode through the Python client, direct HTTP requests, or OpenAI-compatible API calls.

### Python Client Configuration

Use the `concurrency_mode` parameter when initializing `MTPLXClient` to specify `"mtp"`, `"ar"`, or `"hyper"`:

```python
from mtplx import MTPLXClient

# Initialize client for MTP batch mode with explicit prefill limits

client = MTPLXClient(
    base_url="http://localhost:8000",
    model="qwen2.5-7b-mtp",
    concurrency_mode="mtp",  # Options: "ar" | "mtp" | "hyper"

    max_batch_size=8,
)

resp = client.batch_generate(
    prompts=[
        "Explain quantum entanglement in simple terms.",
        "Summarize the plot of Inception.",
    ]
)
print(resp)

```

### Command-Line Invocation

For Hyper mode endpoints, use direct HTTP POST requests:

```bash
curl -X POST http://localhost:8000/v1/hyper/completions \
     -H "Content-Type: application/json" \
     -d '{"model":"qwen2.5-7b","prompt":"Write a haiku about sunrise."}'

```

### OpenAI-Compatible API

When using the OpenAI-compatible endpoint, the server automatically selects the optimal lane based on request volume:

```python
import openai

openai.api_base = "http://localhost:8000/v1"
openai.api_key = "placeholder"

response = openai.ChatCompletion.create(
    model="qwen2.5-7b",
    messages=[{"role": "user", "content": "What is the capital of France?"}]
)
print(response.choices[0].message["content"])

```

## Summary

- **AR Lane**: Default FIFO mode for single prompts; shares adaptive concurrency with batched workloads as implemented in [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py).
- **MTP Batch**: High-throughput mode with explicit prefill concurrency capped at 2; defined in [`mtplx/mtp_batch_numerics.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_batch_numerics.py).
- **Hyper Mode**: Definitional concurrency with strict FIFO and no tunable width; source comments in [`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py) specify this behavior.
- **Configuration**: Use the `concurrency_mode` parameter in `MTPLXClient` or rely on automatic dispatcher routing via the OpenAI-compatible API.

## Frequently Asked Questions

### What is the difference between AR and MTP concurrency modes?

The **AR lane** processes requests sequentially in FIFO order with step-by-step token generation, making it ideal for interactive latency-sensitive applications. The **MTP batch mode** groups requests to generate multiple tokens in parallel, capping prefill concurrency at 2 to manage memory while maximizing decode throughput for bulk operations.

### Why is prefill concurrency limited to 2 in MTP mode?

The prefill concurrency cap of 2 in MTP mode, as defined in [`mtplx/mtp_batch_numerics.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/mtp_batch_numerics.py), prevents excessive memory utilization during the initial prompt processing phase while still allowing higher parallelism during token generation. This trade-off ensures predictable resource usage under heavy batch loads.

### Can I tune the concurrency width in Hyper mode?

No. According to the implementation in [`mtplx/server/hyper.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/hyper.py), Hyper mode treats concurrency as definitional rather than tunable, meaning the system dedicates its full width to maintaining strict FIFO ordering without exposing configurable concurrency parameters to the user.

### Which mode should I use for high-throughput batch processing?

Use **MTP batch mode** by setting `concurrency_mode="mtp"` in your client configuration. This mode is explicitly designed for high-throughput scenarios, providing optimized parallelism during decode phases while enforcing the prefill limit to maintain system stability.