# How the Hugging Face Speech-to-Speech Pipeline Handles Concurrent WebSocket Connections with num_pipelines

> Learn how the Hugging Face speech-to-speech pipeline efficiently manages concurrent WebSocket connections using num_pipelines and thread pools for optimal performance.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-11

---

**The Hugging Face speech-to-speech pipeline manages concurrent WebSocket connections by instantiating a thread pool sized according to the `num_pipelines` argument, where each pipeline instance processes exactly one connection at a time and the server rejects additional clients when the pool is saturated.**

The `huggingface/speech-to-speech` repository implements a scalable architecture to handle multiple simultaneous audio streams using the `num_pipelines` parameter. This design creates a managed pool of independent pipeline instances that determines precisely how many concurrent WebSocket clients the real-time server can accept before applying backpressure.

## Thread Pool Architecture and Pipeline Instantiation

The concurrency model centers on a **thread pool** created at startup in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). The pool size is calculated as the maximum of 1 and the user-provided value to ensure at least one pipeline always exists:

```python
pool_size = max(1, module_kwargs.num_pipelines)  # s2s_pipeline.py, line 642

```

The [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py) module implements the actual pool management, spawning a separate **worker thread** for each pipeline instance. Each instance operates independently within its own thread, allowing true parallel processing of audio streams across multiple CPU cores without blocking the main server loop.

## Connection Routing and Backpressure Handling

When a client connects via WebSocket, the server assigns the connection to the next available pipeline in the pool. If all pipelines are currently occupied processing active streams, the server implements **backpressure** by rejecting new connections immediately. According to the source code in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) (line 70), the system responds with a message indicating that *"further connections are rejected"* until an existing client disconnects and frees a pipeline instance.

## Configuration Constraints and Platform Limitations

The `num_pipelines` feature carries specific runtime restrictions. First, multiple pipelines are **only supported in realtime mode**. The CLI validation in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) (lines 1010–1014) raises a `ValueError` if `num_pipelines > 1` while the mode is set to anything other than `"realtime"`:

```python
if args.module_kwargs.num_pipelines > 1 and args.module_kwargs.mode != "realtime":
    raise ValueError(
        f"--num_pipelines > 1 is only supported with --mode realtime "
        f"(got mode={args.module_kwargs.mode!r}, num_pipelines={args.module_kwargs.num_pipelines})"
    )

```

Second, on **Apple Silicon** (platform `darwin`), enabling multiple pipelines forces the system to disable **live transcription** due to MLX framework contention. The runtime check in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) (lines 1022–1026) logs a warning when this condition is detected:

```python
if args.module_kwargs.num_pipelines > 1 and platform == "darwin" and args.module_kwargs.enable_live_transcription:
    logger.warning(
        "MLX contention: --num_pipelines=%d > 1 on Apple Silicon → disabling live transcription ",
        args.module_kwargs.num_pipelines,
    )

```

## Practical Implementation Example

To launch a server capable of handling three concurrent WebSocket clients:

```bash
python -m speech_to_speech.demo.server \
    --mode realtime \
    --num_pipelines 3 \
    --port 8000

```

Connect using a Python WebSocket client:

```python
import websockets
import asyncio

async def transcribe():
    uri = "ws://localhost:8000/realtime"
    async with websockets.connect(uri) as ws:
        await ws.send(b"...")  # Send audio chunks

        response = await ws.recv()  # Receive synthesized audio

        print(response)

asyncio.run(transcribe())

```

Attempting to connect a fourth client while three are active results in immediate connection rejection, preventing resource exhaustion.

## Summary

- The **thread pool** size is determined by `max(1, num_pipelines)` in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)
- Each pipeline instance handles **exactly one WebSocket connection** at a time
- Excess connections are **rejected** when the pool is exhausted, as defined in [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py)
- Multiple pipelines require **realtime mode** and disable live transcription on **Apple Silicon** due to MLX contention

## Frequently Asked Questions

### What happens when the number of concurrent connections exceeds num_pipelines?

When all pipeline instances in the thread pool are occupied, the server rejects new WebSocket connections immediately. According to the implementation in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py), the server responds with a rejection message indicating that no further connections are accepted until an existing client disconnects.

### Can num_pipelines be used in batch or offline processing modes?

No. The validation logic in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 1010–1014) explicitly raises a `ValueError` if `num_pipelines` is greater than 1 while the mode is not set to `"realtime"`. This restriction exists because the concurrent connection handling architecture is specifically designed for real-time streaming scenarios.

### Why does Apple Silicon disable live transcription when using multiple pipelines?

On Apple Silicon devices (platform `darwin`), the MLX framework experiences resource contention when multiple pipeline instances run simultaneously with live transcription enabled. The code in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) (lines 1022–1026) detects this configuration and automatically disables live transcription to prevent performance degradation, logging a warning to notify the user of the adjustment.

### Where is the thread pool for managing pipelines actually implemented?

The thread pool logic resides in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py), which manages the lifecycle of worker threads. The pipeline initialization code in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (line 642) calculates the required pool size and delegates thread management to this utility module, ensuring each pipeline instance runs in its own isolated execution context.