# Performance and Memory Implications of Using `--num_pipelines` with Multiple Concurrent WebSocket Connections

> Explore performance and memory implications of Hugging Face speech-to-speech `--num_pipelines`. Understand how increasing pipelines impacts latency, memory, and GPU/CPU usage with concurrent WebSocket connections.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: performance
- Published: 2026-07-10

---

**Increasing `--num_pipelines` enables the Hugging Face speech-to-speech server to accept more simultaneous WebSocket clients, but each additional pipeline instantiates a complete duplicate of the VAD, STT, LLM, and TTS models, resulting in linear memory growth and potential GPU/CPU contention that increases latency.**

The `--num_pipelines` flag in the `huggingface/speech-to-speech` repository determines how many isolated realtime pipelines the server creates to handle incoming WebSocket connections. Understanding the performance and memory implications of using `--num_pipelines` with multiple concurrent WebSocket connections is essential for production deployments, as the architecture deliberately trades resource efficiency for strict client isolation.

## How `--num_pipelines` Limits Concurrent Connections

The server enforces a hard concurrency cap equal to the value of `--num_pipelines`. In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 989–1014), the start-up logic validates that the number of active WebSocket sessions never exceeds this configured limit.

When a client attempts to connect while all pipelines are occupied, the server rejects the connection immediately. This design ensures that each accepted WebSocket connection receives a dedicated pipeline instance, preventing resource starvation among active clients.

## Linear Memory Scaling with Model Duplication

Each pipeline instantiated by `--num_pipelines` loads its own complete model stack, including **Voice Activity Detection (VAD)**, **Speech-to-Text (STT)**, **Large Language Model (LLM)**, and **Text-to-Speech (TTS)** components. Memory consumption grows roughly linearly with the pipeline count because no model weights are shared between instances.

The **LLM** and **TTS** models dominate the memory footprint. For example, if a single LLM instance occupies 8 GB of GPU memory, launching two pipelines requires approximately 16 GB, three pipelines require 24 GB, and so on. This multiplicative effect makes memory the primary bottleneck when scaling `--num_pipelines` beyond 2–3 instances on consumer hardware.

## GPU, CPU Contention and Latency Impact

With multiple pipelines sharing the same physical hardware, **GPU** and **CPU** resources contend for compute cycles. This contention can increase per-request latency, particularly when the device approaches saturation.

On **Apple Silicon** devices, the library detects this resource pressure and automatically disables live transcription when `--num_pipelines` exceeds 1 (as implemented in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) lines 1022–1024). This automatic fallback prevents performance degradation but removes the real-time transcription feature for multi-pipeline configurations on macOS.

## Threading Overhead and Pipeline Isolation

Each pipeline runs in its own dedicated thread managed by [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py). While Python's Global Interpreter Lock (GIL) and context-switching overhead add modest per-pipeline costs, these are negligible compared to the heavy inference compute of the LLM and TTS models.

The threading architecture deliberately isolates client state to prevent data mixing between concurrent WebSocket connections. This isolation guarantees that audio streams and conversation contexts from one client never leak into another pipeline's processing queue.

## Deployment Patterns for High Concurrency

Rather than increasing `--num_pipelines` to serve dozens of clients, the recommended scalability pattern involves running multiple separate server processes—each with `--num_pipelines=1`—behind a load balancer. This approach:

- Prevents a single process from exhausting available GPU memory
- Distributes CPU contention across multiple OS processes
- Maintains the isolation guarantees without the memory multiplication of high pipeline counts

Single-process deployments with high `--num_pipelines` values should be reserved for scenarios where the aggregate model sizes fit comfortably within available VRAM and latency requirements tolerate resource contention.

## Code Examples

Configure a realtime server to handle up to three concurrent clients:

```bash
speech-to-speech \
    --mode realtime \
    --num_pipelines 3 \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3

```

On macOS, explicitly disable live transcription when using multiple pipelines to avoid automatic performance degradation:

```bash
speech-to-speech \
    --mode realtime \
    --num_pipelines 2 \
    --enable_live_transcription false \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3

```

Connect a Python client to a multi-pipeline server:

```python
import websockets
import asyncio

async def talk():
    uri = "ws://localhost:8765/v1/realtime"
    async with websockets.connect(uri) as ws:
        await ws.send('{"type":"session.update","session":{"type":"realtime"}}')
        async for msg in ws:
            print(msg)

asyncio.run(talk())

```

## Summary

- **`--num_pipelines` sets a hard limit** on simultaneous WebSocket connections; additional clients are rejected when the limit is reached according to [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).
- **Memory scales linearly** with pipeline count, with LLM and TTS models consuming the majority of GPU VRAM.
- **Resource contention** increases latency when multiple pipelines share GPU or CPU resources, prompting automatic feature degradation on Apple Silicon.
- **Thread isolation** prevents state leakage between clients but adds minimal overhead compared to model inference costs.
- **Horizontal scaling** via multiple single-pipeline processes behind a load balancer is preferred over vertical scaling with high `--num_pipelines` values.

## Frequently Asked Questions

### What happens when WebSocket connections exceed the `--num_pipelines` limit?

The server rejects new connections once the number of active WebSocket sessions equals the configured `--num_pipelines` value. This enforcement occurs in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 989–1014), ensuring that each connected client receives a dedicated pipeline instance without resource overcommitment.

### How does memory usage scale when increasing `--num_pipelines`?

Memory consumption grows roughly linearly because each pipeline loads independent instances of the VAD, STT, LLM, and TTS models. If your LLM requires 8 GB of GPU memory, running three pipelines consumes approximately 24 GB for the LLM components alone, plus additional memory for the audio processing stacks.

### Why is live transcription disabled on macOS when using multiple pipelines?

On Apple Silicon devices, the library automatically disables live transcription in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 1022–1024) when `--num_pipelines` exceeds 1. This prevents GPU/CPU contention from degrading real-time audio processing performance, though it removes the real-time text feedback feature for multi-client scenarios on macOS.

### What is the recommended architecture for serving many concurrent users?

Deploy multiple server processes with `--num_pipelines=1` behind a load balancer rather than scaling a single process to high pipeline counts. This horizontal scaling pattern prevents memory exhaustion and distributes compute contention across separate OS processes while maintaining the isolation guarantees of the pipeline architecture.