# OpenAI Realtime Protocol Handler: Managing Concurrent WebSocket Sessions in Speech-to-Speech

> Learn how the Hugging Face Speech-to-Speech library manages concurrent WebSocket sessions using asyncio Semaphore and RuntimeConfig for efficient real-time communication. Optimize your performance today.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: architecture
- Published: 2026-07-09

---

**The Hugging Face Speech-to-Speech library controls concurrent WebSocket connections through an `asyncio.Semaphore` initialized from the `RuntimeConfig.max_concurrent_sessions` setting, configurable via the `S2S_MAX_CONCURRENT_SESSIONS` environment variable.**

The `huggingface/speech-to-speech` repository implements the OpenAI Realtime protocol for real-time audio processing. Managing concurrent WebSocket sessions is essential for resource allocation and stability, enforced through a centralized configuration system that limits simultaneous connections using a semaphore-based session manager.

## How Concurrent Session Limits Work

The architecture separates configuration from session management logic. The `RuntimeConfig` dataclass defines the concurrency ceiling, while the `RealtimeService` implements the actual semaphore-based gatekeeping.

### The RuntimeConfig Dataclass

In [`src/speech_to_speech/api/openai_realtime/runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/runtime_config.py), the `RuntimeConfig` dataclass encapsulates all service parameters:

```python
@dataclass
class RuntimeConfig:
    max_concurrent_sessions: int = 10
    enable_logging: bool = False
    allowed_origins: List[str] = field(default_factory=lambda: ["*"])
    timeout_seconds: int = 300

```

The **`max_concurrent_sessions`** attribute defaults to **10** simultaneous connections. The class provides a **`from_env()`** method that reads the **`S2S_MAX_CONCURRENT_SESSIONS`** environment variable to override this default.

### The Session Manager Semaphore

The `RealtimeService` class in [`src/speech_to_speech/api/openai_realtime/pipeline_unit.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/pipeline_unit.py) initializes an **`asyncio.Semaphore`** using the value from `RuntimeConfig.max_concurrent_sessions`. This semaphore acts as the session manager, controlling access to WebSocket handler resources.

Each incoming connection request triggers an **`acquire()`** call on the semaphore. If the semaphore counter reaches zero, the server rejects new connections until existing sessions release their slots.

## Configuring Concurrent Sessions

You can configure the concurrency limit either through environment variables for containerized deployments or programmatically for embedded use cases.

### Environment Variable Configuration

For Docker or Kubernetes deployments, set the **`S2S_MAX_CONCURRENT_SESSIONS`** variable before starting the server:

```bash
export S2S_MAX_CONCURRENT_SESSIONS=20
export S2S_TIMEOUT_SECONDS=600
uvicorn speech_to_speech.demo.server:app --host 0.0.0.0 --port 8000

```

The **`from_env()`** method in `RuntimeConfig` converts these variables into the appropriate types:

```python
@classmethod
def from_env(cls) -> "RuntimeConfig":
    max_sessions = int(os.getenv("S2S_MAX_CONCURRENT_SESSIONS", "10"))
    enable_logging = _bool_from_env(os.getenv("S2S_ENABLE_LOGGING"), False)
    origins_raw = os.getenv("S2S_ALLOWED_ORIGINS", "*")
    allowed_origins = [origin.strip() for origin in origins_raw.split(",") if origin]
    timeout = int(os.getenv("S2S_TIMEOUT_SECONDS", "300"))
    return cls(
        max_concurrent_sessions=max_sessions,
        enable_logging=enable_logging,
        allowed_origins=allowed_origins,
        timeout_seconds=timeout,
    )

```

### Programmatic Configuration

For custom server setups, instantiate `RuntimeConfig` directly and pass it to the application factory:

```python
from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig
from speech_to_speech.demo.server import create_app

custom_config = RuntimeConfig(
    max_concurrent_sessions=50,
    enable_logging=True,
    allowed_origins=["https://myapp.example.com"],
    timeout_seconds=300,
)

app = create_app(runtime_config=custom_config)

```

## WebSocket Session Lifecycle Management

The `WebSocketStreamer` in [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py) wraps FastAPI WebSocket objects and only initializes after acquiring a semaphore slot from the session manager.

### Connection Acceptance

When a client connects, the handler attempts to acquire the semaphore before establishing the WebSocket stream:

```python
async def on_connect(websocket: WebSocket, config: RuntimeConfig):
    try:
        await config.session_manager.acquire()
        streamer = WebSocketStreamer(websocket)
        await streamer.start()
    except RuntimeError:
        await websocket.close(code=1008, reason="Maximum sessions reached")
        return

```

### Connection Rejection and Timeouts

If the semaphore is exhausted, the server returns a **403 Forbidden** response or closes the WebSocket with code **1008** and the reason "Maximum sessions reached".

Additionally, the **`timeout_seconds`** setting (default 300 seconds) automatically closes idle sessions, releasing their semaphore slots back to the pool. The `RealtimeService` monitors activity and triggers cleanup when the idle threshold expires.

## Key Implementation Files

- **[`src/speech_to_speech/api/openai_realtime/runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/runtime_config.py)** – Defines the `RuntimeConfig` dataclass, the `_bool_from_env` helper, environment variable parsing, and the `default_config` singleton.
- **[`src/speech_to_speech/api/openai_realtime/pipeline_unit.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/pipeline_unit.py)** – Implements the `RealtimeService` class and manages the `asyncio.Semaphore` for session concurrency.
- **[`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py)** – Wraps FastAPI WebSocket connections and handles binary audio frame streaming.
- **[`demo/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/demo/server.py)** – FastAPI entry point that injects `RuntimeConfig` and mounts the WebSocket router.

## Summary

- The Speech-to-Speech library uses an **`asyncio.Semaphore`** to enforce the `max_concurrent_sessions` limit from `RuntimeConfig`.
- Configure the limit via the **`S2S_MAX_CONCURRENT_SESSIONS`** environment variable or programmatically through the `RuntimeConfig` constructor.
- The **`from_env()`** method provides type-safe parsing of environment variables using the **`_bool_from_env`** helper for booleans.
- Exceeded capacity results in **WebSocket close code 1008** or a 403 response, while idle sessions automatically timeout after the configured **`timeout_seconds`**.
- Key files include **[`runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/runtime_config.py)** for configuration and **[`pipeline_unit.py`](https://github.com/huggingface/speech-to-speech/blob/main/pipeline_unit.py)** for the semaphore-based session manager.

## Frequently Asked Questions

### How do I increase the maximum number of concurrent WebSocket sessions?

Set the **`S2S_MAX_CONCURRENT_SESSIONS`** environment variable to your desired integer value before starting the Uvicorn server. Alternatively, create a custom `RuntimeConfig` instance with a higher `max_concurrent_sessions` value and pass it to `create_app()` in [`demo/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/demo/server.py).

### What happens when the server reaches its concurrent session limit?

When the semaphore counter reaches zero, new connection attempts receive a **403 Forbidden** response or the WebSocket closes with code **1008** and the reason "Maximum sessions reached". The client must retry after an existing session disconnects or times out.

### How is the idle timeout configured for OpenAI Realtime sessions?

The **`S2S_TIMEOUT_SECONDS`** environment variable controls the idle timeout, defaulting to **300 seconds** (5 minutes). The `RealtimeService` monitors session activity and automatically closes connections that exceed this idle threshold, releasing the semaphore slot for new clients.

### Can I customize the RuntimeConfig programmatically instead of using environment variables?

Yes. Import `RuntimeConfig` from [`src/speech_to_speech/api/openai_realtime/runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/runtime_config.py), instantiate it with your desired parameters, and pass it to the application factory. This approach is useful for testing or when embedding the Speech-to-Speech server within larger Python applications.