Performance and Memory Implications of Using `--num_pipelines` with Multiple Concurrent WebSocket Connections
Increasing --num_pipelines enables the Hugging Face speech-to-speech server to accept more simultaneous WebSocket clients, but each additional pipeline instantiates a complete duplicate of the VAD, STT, LLM, and TTS models, resulting in linear memory growth and potential GPU/CPU contention that increases latency.
The --num_pipelines flag in the huggingface/speech-to-speech repository determines how many isolated realtime pipelines the server creates to handle incoming WebSocket connections. Understanding the performance and memory implications of using --num_pipelines with multiple concurrent WebSocket connections is essential for production deployments, as the architecture deliberately trades resource efficiency for strict client isolation.
How --num_pipelines Limits Concurrent Connections
The server enforces a hard concurrency cap equal to the value of --num_pipelines. In src/speech_to_speech/s2s_pipeline.py (lines 989–1014), the start-up logic validates that the number of active WebSocket sessions never exceeds this configured limit.
When a client attempts to connect while all pipelines are occupied, the server rejects the connection immediately. This design ensures that each accepted WebSocket connection receives a dedicated pipeline instance, preventing resource starvation among active clients.
Linear Memory Scaling with Model Duplication
Each pipeline instantiated by --num_pipelines loads its own complete model stack, including Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) components. Memory consumption grows roughly linearly with the pipeline count because no model weights are shared between instances.
The LLM and TTS models dominate the memory footprint. For example, if a single LLM instance occupies 8 GB of GPU memory, launching two pipelines requires approximately 16 GB, three pipelines require 24 GB, and so on. This multiplicative effect makes memory the primary bottleneck when scaling --num_pipelines beyond 2–3 instances on consumer hardware.
GPU, CPU Contention and Latency Impact
With multiple pipelines sharing the same physical hardware, GPU and CPU resources contend for compute cycles. This contention can increase per-request latency, particularly when the device approaches saturation.
On Apple Silicon devices, the library detects this resource pressure and automatically disables live transcription when --num_pipelines exceeds 1 (as implemented in src/speech_to_speech/s2s_pipeline.py lines 1022–1024). This automatic fallback prevents performance degradation but removes the real-time transcription feature for multi-pipeline configurations on macOS.
Threading Overhead and Pipeline Isolation
Each pipeline runs in its own dedicated thread managed by src/speech_to_speech/utils/thread_manager.py. While Python's Global Interpreter Lock (GIL) and context-switching overhead add modest per-pipeline costs, these are negligible compared to the heavy inference compute of the LLM and TTS models.
The threading architecture deliberately isolates client state to prevent data mixing between concurrent WebSocket connections. This isolation guarantees that audio streams and conversation contexts from one client never leak into another pipeline's processing queue.
Deployment Patterns for High Concurrency
Rather than increasing --num_pipelines to serve dozens of clients, the recommended scalability pattern involves running multiple separate server processes—each with --num_pipelines=1—behind a load balancer. This approach:
- Prevents a single process from exhausting available GPU memory
- Distributes CPU contention across multiple OS processes
- Maintains the isolation guarantees without the memory multiplication of high pipeline counts
Single-process deployments with high --num_pipelines values should be reserved for scenarios where the aggregate model sizes fit comfortably within available VRAM and latency requirements tolerate resource contention.
Code Examples
Configure a realtime server to handle up to three concurrent clients:
speech-to-speech \
--mode realtime \
--num_pipelines 3 \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3
On macOS, explicitly disable live transcription when using multiple pipelines to avoid automatic performance degradation:
speech-to-speech \
--mode realtime \
--num_pipelines 2 \
--enable_live_transcription false \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3
Connect a Python client to a multi-pipeline server:
import websockets
import asyncio
async def talk():
uri = "ws://localhost:8765/v1/realtime"
async with websockets.connect(uri) as ws:
await ws.send('{"type":"session.update","session":{"type":"realtime"}}')
async for msg in ws:
print(msg)
asyncio.run(talk())
Summary
--num_pipelinessets a hard limit on simultaneous WebSocket connections; additional clients are rejected when the limit is reached according tosrc/speech_to_speech/s2s_pipeline.py.- Memory scales linearly with pipeline count, with LLM and TTS models consuming the majority of GPU VRAM.
- Resource contention increases latency when multiple pipelines share GPU or CPU resources, prompting automatic feature degradation on Apple Silicon.
- Thread isolation prevents state leakage between clients but adds minimal overhead compared to model inference costs.
- Horizontal scaling via multiple single-pipeline processes behind a load balancer is preferred over vertical scaling with high
--num_pipelinesvalues.
Frequently Asked Questions
What happens when WebSocket connections exceed the --num_pipelines limit?
The server rejects new connections once the number of active WebSocket sessions equals the configured --num_pipelines value. This enforcement occurs in src/speech_to_speech/s2s_pipeline.py (lines 989–1014), ensuring that each connected client receives a dedicated pipeline instance without resource overcommitment.
How does memory usage scale when increasing --num_pipelines?
Memory consumption grows roughly linearly because each pipeline loads independent instances of the VAD, STT, LLM, and TTS models. If your LLM requires 8 GB of GPU memory, running three pipelines consumes approximately 24 GB for the LLM components alone, plus additional memory for the audio processing stacks.
Why is live transcription disabled on macOS when using multiple pipelines?
On Apple Silicon devices, the library automatically disables live transcription in src/speech_to_speech/s2s_pipeline.py (lines 1022–1024) when --num_pipelines exceeds 1. This prevents GPU/CPU contention from degrading real-time audio processing performance, though it removes the real-time text feedback feature for multi-client scenarios on macOS.
What is the recommended architecture for serving many concurrent users?
Deploy multiple server processes with --num_pipelines=1 behind a load balancer rather than scaling a single process to high pipeline counts. This horizontal scaling pattern prevents memory exhaustion and distributes compute contention across separate OS processes while maintaining the isolation guarantees of the pipeline architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →