OpenAI Realtime Protocol Handler: Managing Concurrent WebSocket Sessions in Speech-to-Speech

The Hugging Face Speech-to-Speech library controls concurrent WebSocket connections through an asyncio.Semaphore initialized from the RuntimeConfig.max_concurrent_sessions setting, configurable via the S2S_MAX_CONCURRENT_SESSIONS environment variable.

The huggingface/speech-to-speech repository implements the OpenAI Realtime protocol for real-time audio processing. Managing concurrent WebSocket sessions is essential for resource allocation and stability, enforced through a centralized configuration system that limits simultaneous connections using a semaphore-based session manager.

How Concurrent Session Limits Work

The architecture separates configuration from session management logic. The RuntimeConfig dataclass defines the concurrency ceiling, while the RealtimeService implements the actual semaphore-based gatekeeping.

The RuntimeConfig Dataclass

In src/speech_to_speech/api/openai_realtime/runtime_config.py, the RuntimeConfig dataclass encapsulates all service parameters:

@dataclass
class RuntimeConfig:
    max_concurrent_sessions: int = 10
    enable_logging: bool = False
    allowed_origins: List[str] = field(default_factory=lambda: ["*"])
    timeout_seconds: int = 300

The max_concurrent_sessions attribute defaults to 10 simultaneous connections. The class provides a from_env() method that reads the S2S_MAX_CONCURRENT_SESSIONS environment variable to override this default.

The Session Manager Semaphore

The RealtimeService class in src/speech_to_speech/api/openai_realtime/pipeline_unit.py initializes an asyncio.Semaphore using the value from RuntimeConfig.max_concurrent_sessions. This semaphore acts as the session manager, controlling access to WebSocket handler resources.

Each incoming connection request triggers an acquire() call on the semaphore. If the semaphore counter reaches zero, the server rejects new connections until existing sessions release their slots.

Configuring Concurrent Sessions

You can configure the concurrency limit either through environment variables for containerized deployments or programmatically for embedded use cases.

Environment Variable Configuration

For Docker or Kubernetes deployments, set the S2S_MAX_CONCURRENT_SESSIONS variable before starting the server:

export S2S_MAX_CONCURRENT_SESSIONS=20
export S2S_TIMEOUT_SECONDS=600
uvicorn speech_to_speech.demo.server:app --host 0.0.0.0 --port 8000

The from_env() method in RuntimeConfig converts these variables into the appropriate types:

@classmethod
def from_env(cls) -> "RuntimeConfig":
    max_sessions = int(os.getenv("S2S_MAX_CONCURRENT_SESSIONS", "10"))
    enable_logging = _bool_from_env(os.getenv("S2S_ENABLE_LOGGING"), False)
    origins_raw = os.getenv("S2S_ALLOWED_ORIGINS", "*")
    allowed_origins = [origin.strip() for origin in origins_raw.split(",") if origin]
    timeout = int(os.getenv("S2S_TIMEOUT_SECONDS", "300"))
    return cls(
        max_concurrent_sessions=max_sessions,
        enable_logging=enable_logging,
        allowed_origins=allowed_origins,
        timeout_seconds=timeout,
    )

Programmatic Configuration

For custom server setups, instantiate RuntimeConfig directly and pass it to the application factory:

from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig
from speech_to_speech.demo.server import create_app

custom_config = RuntimeConfig(
    max_concurrent_sessions=50,
    enable_logging=True,
    allowed_origins=["https://myapp.example.com"],
    timeout_seconds=300,
)

app = create_app(runtime_config=custom_config)

WebSocket Session Lifecycle Management

The WebSocketStreamer in src/speech_to_speech/connections/websocket_streamer.py wraps FastAPI WebSocket objects and only initializes after acquiring a semaphore slot from the session manager.

Connection Acceptance

When a client connects, the handler attempts to acquire the semaphore before establishing the WebSocket stream:

async def on_connect(websocket: WebSocket, config: RuntimeConfig):
    try:
        await config.session_manager.acquire()
        streamer = WebSocketStreamer(websocket)
        await streamer.start()
    except RuntimeError:
        await websocket.close(code=1008, reason="Maximum sessions reached")
        return

Connection Rejection and Timeouts

If the semaphore is exhausted, the server returns a 403 Forbidden response or closes the WebSocket with code 1008 and the reason "Maximum sessions reached".

Additionally, the timeout_seconds setting (default 300 seconds) automatically closes idle sessions, releasing their semaphore slots back to the pool. The RealtimeService monitors activity and triggers cleanup when the idle threshold expires.

Key Implementation Files

Summary

  • The Speech-to-Speech library uses an asyncio.Semaphore to enforce the max_concurrent_sessions limit from RuntimeConfig.
  • Configure the limit via the S2S_MAX_CONCURRENT_SESSIONS environment variable or programmatically through the RuntimeConfig constructor.
  • The from_env() method provides type-safe parsing of environment variables using the _bool_from_env helper for booleans.
  • Exceeded capacity results in WebSocket close code 1008 or a 403 response, while idle sessions automatically timeout after the configured timeout_seconds.
  • Key files include runtime_config.py for configuration and pipeline_unit.py for the semaphore-based session manager.

Frequently Asked Questions

How do I increase the maximum number of concurrent WebSocket sessions?

Set the S2S_MAX_CONCURRENT_SESSIONS environment variable to your desired integer value before starting the Uvicorn server. Alternatively, create a custom RuntimeConfig instance with a higher max_concurrent_sessions value and pass it to create_app() in demo/server.py.

What happens when the server reaches its concurrent session limit?

When the semaphore counter reaches zero, new connection attempts receive a 403 Forbidden response or the WebSocket closes with code 1008 and the reason "Maximum sessions reached". The client must retry after an existing session disconnects or times out.

How is the idle timeout configured for OpenAI Realtime sessions?

The S2S_TIMEOUT_SECONDS environment variable controls the idle timeout, defaulting to 300 seconds (5 minutes). The RealtimeService monitors session activity and automatically closes connections that exceed this idle threshold, releasing the semaphore slot for new clients.

Can I customize the RuntimeConfig programmatically instead of using environment variables?

Yes. Import RuntimeConfig from src/speech_to_speech/api/openai_realtime/runtime_config.py, instantiate it with your desired parameters, and pass it to the application factory. This approach is useful for testing or when embedding the Speech-to-Speech server within larger Python applications.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →