# How Pool Size (`--num_pipelines`) Affects Concurrent Session Capacity in Hugging Face Speech-to-Speech

> Discover how increasing --num_pipelines boosts concurrent session capacity in Hugging Face Speech-to-Speech. Learn its limitations for realtime sessions and Apple Silicon.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: performance
- Published: 2026-07-09

---

**Increasing `--num_pipelines` creates additional pipeline instances that allow the Hugging Face Speech-to-Speech server to handle multiple simultaneous realtime sessions, though it is restricted to realtime mode and limited on Apple Silicon due to MLX resource contention.**

The `huggingface/speech-to-speech` library processes audio through dedicated pipelines, where each pipeline manages a single user session. When deploying a realtime server, the `--num_pipelines` argument controls how many concurrent sessions the system can accommodate before rejecting additional connections.

---

## Where `--num_pipelines` Is Defined

The argument is declared in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) as part of the module configuration class:

```python
num_pipelines: int = field(metadata={"help": "num_pipelines; further connections are rejected. Only valid for --mode realtime. Default is 1."})

```

This **field** establishes the default pool size of `1`, meaning the server handles one realtime session at a time unless explicitly configured otherwise.

---

## Effect on Concurrent Session Capacity

Each pipeline instance corresponds to one active session. The library calculates the pool size using:

```python
pool_size = max(1, module_kwargs.num_pipelines)

```

This value determines how many pipeline threads are spawned to process incoming audio streams. When the pool is exhausted, the system rejects new connections until an existing session terminates.

---

## Critical Restrictions on Pool Size

The library enforces strict validation to prevent resource conflicts and unsupported configurations.

### Realtime Mode Requirement

The `--num_pipelines` argument is only valid when `--mode realtime` is specified. In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the code explicitly validates this combination:

```python
if args.module_kwargs.num_pipelines > 1 and args.module_kwargs.mode != "realtime":
    raise ValueError("--num_pipelines > 1 is only supported with --mode realtime")

```

Attempting to use multiple pipelines in asynchronous or batch modes triggers a **ValueError** immediately at startup.

### Apple Silicon Limitations

On macOS (`darwin`) platforms, MLX resource contention prevents stable multi-pipeline operation when live transcription is enabled. The library detects this condition and automatically downgrades the pool:

```python
if args.module_kwargs.num_pipelines > 1 and platform == "darwin" and args.module_kwargs.enable_live_transcription:
    logger.warning("MLX contention: --num_pipelines=%d > 1 on Apple Silicon → disabling live transcription", args.module_kwargs.num_pipelines)
    args.module_kwargs.num_pipelines = 1

```

In this scenario, the system **warns the user** and forces the pool size back to `1`, effectively disabling live transcription to maintain stability.

---

## Practical Configuration Examples

### Launch a 4-Session Realtime Server

To support four concurrent users on a Linux or CUDA-enabled system:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --num_pipelines 4 \
    --port 8765

```

### Invalid Configuration (Non-Realtime Mode)

The following command fails because `--num_pipelines` exceeds `1` without realtime mode:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode async \
    --num_pipelines 2

```

This produces:

```text
ValueError: --num_pipelines > 1 is only supported with --mode realtime

```

### macOS Auto-Downgrade Behavior

On Apple Silicon with live transcription enabled, attempting multiple pipelines triggers a warning and fallback:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --num_pipelines 2 \
    --enable_live_transcription true

```

Output:

```text
Warning: MLX contention: --num_pipelines=2 > 1 on Apple Silicon → disabling live transcription

```

---

## Summary

- **`--num_pipelines`** sets the number of simultaneous realtime sessions the server can handle, defaulting to `1`.
- **Pool creation** occurs in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) via `pool_size = max(1, module_kwargs.num_pipelines)`.
- **Realtime restriction**: Values greater than `1` require `--mode realtime`; otherwise, the system raises a `ValueError`.
- **Apple Silicon guard**: On macOS with live transcription enabled, the library automatically reduces the pool to `1` to prevent MLX resource contention.
- **Capacity limit**: Once the pool is saturated, the server rejects additional connections until a pipeline becomes available.

---

## Frequently Asked Questions

### What happens when `--num_pipelines` is set higher than available hardware resources?

The library does not perform automatic hardware detection. If you set `--num_pipelines` to a value exceeding your CPU/GPU capacity, sessions will compete for compute resources, potentially causing latency or timeouts. You must manually tune this value based on your hardware specifications and model size.

### Can I change the pool size dynamically without restarting the server?

No. The pool size is determined at startup in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) and cannot be adjusted dynamically. To modify concurrent capacity, you must restart the service with a new `--num_pipelines` value.

### Why does the library reject `--num_pipelines 2` when using `--mode async`?

The validation logic in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) explicitly restricts multi-pipeline configurations to realtime mode because asynchronous and batch processing pipelines are designed for single-session, file-based workflows rather than concurrent stream handling. Multi-pipeline support is only implemented for the realtime WebSocket server architecture.

### Does increasing `--num_pipelines` improve latency for a single user?

No. Latency for an individual session is determined by the model inference speed and audio chunk processing time. Increasing `--num_pipelines` only affects **throughput** (number of simultaneous users), not the latency experienced by any single active session.