OpenAI Realtime Protocol Handler: Managing Concurrent WebSocket Sessions in Speech-to-Speech
The Hugging Face Speech-to-Speech library controls concurrent WebSocket connections through an asyncio.Semaphore initialized from the RuntimeConfig.max_concurrent_sessions setting, configurable via the S2S_MAX_CONCURRENT_SESSIONS environment variable.
The huggingface/speech-to-speech repository implements the OpenAI Realtime protocol for real-time audio processing. Managing concurrent WebSocket sessions is essential for resource allocation and stability, enforced through a centralized configuration system that limits simultaneous connections using a semaphore-based session manager.
How Concurrent Session Limits Work
The architecture separates configuration from session management logic. The RuntimeConfig dataclass defines the concurrency ceiling, while the RealtimeService implements the actual semaphore-based gatekeeping.
The RuntimeConfig Dataclass
In src/speech_to_speech/api/openai_realtime/runtime_config.py, the RuntimeConfig dataclass encapsulates all service parameters:
@dataclass
class RuntimeConfig:
max_concurrent_sessions: int = 10
enable_logging: bool = False
allowed_origins: List[str] = field(default_factory=lambda: ["*"])
timeout_seconds: int = 300
The max_concurrent_sessions attribute defaults to 10 simultaneous connections. The class provides a from_env() method that reads the S2S_MAX_CONCURRENT_SESSIONS environment variable to override this default.
The Session Manager Semaphore
The RealtimeService class in src/speech_to_speech/api/openai_realtime/pipeline_unit.py initializes an asyncio.Semaphore using the value from RuntimeConfig.max_concurrent_sessions. This semaphore acts as the session manager, controlling access to WebSocket handler resources.
Each incoming connection request triggers an acquire() call on the semaphore. If the semaphore counter reaches zero, the server rejects new connections until existing sessions release their slots.
Configuring Concurrent Sessions
You can configure the concurrency limit either through environment variables for containerized deployments or programmatically for embedded use cases.
Environment Variable Configuration
For Docker or Kubernetes deployments, set the S2S_MAX_CONCURRENT_SESSIONS variable before starting the server:
export S2S_MAX_CONCURRENT_SESSIONS=20
export S2S_TIMEOUT_SECONDS=600
uvicorn speech_to_speech.demo.server:app --host 0.0.0.0 --port 8000
The from_env() method in RuntimeConfig converts these variables into the appropriate types:
@classmethod
def from_env(cls) -> "RuntimeConfig":
max_sessions = int(os.getenv("S2S_MAX_CONCURRENT_SESSIONS", "10"))
enable_logging = _bool_from_env(os.getenv("S2S_ENABLE_LOGGING"), False)
origins_raw = os.getenv("S2S_ALLOWED_ORIGINS", "*")
allowed_origins = [origin.strip() for origin in origins_raw.split(",") if origin]
timeout = int(os.getenv("S2S_TIMEOUT_SECONDS", "300"))
return cls(
max_concurrent_sessions=max_sessions,
enable_logging=enable_logging,
allowed_origins=allowed_origins,
timeout_seconds=timeout,
)
Programmatic Configuration
For custom server setups, instantiate RuntimeConfig directly and pass it to the application factory:
from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig
from speech_to_speech.demo.server import create_app
custom_config = RuntimeConfig(
max_concurrent_sessions=50,
enable_logging=True,
allowed_origins=["https://myapp.example.com"],
timeout_seconds=300,
)
app = create_app(runtime_config=custom_config)
WebSocket Session Lifecycle Management
The WebSocketStreamer in src/speech_to_speech/connections/websocket_streamer.py wraps FastAPI WebSocket objects and only initializes after acquiring a semaphore slot from the session manager.
Connection Acceptance
When a client connects, the handler attempts to acquire the semaphore before establishing the WebSocket stream:
async def on_connect(websocket: WebSocket, config: RuntimeConfig):
try:
await config.session_manager.acquire()
streamer = WebSocketStreamer(websocket)
await streamer.start()
except RuntimeError:
await websocket.close(code=1008, reason="Maximum sessions reached")
return
Connection Rejection and Timeouts
If the semaphore is exhausted, the server returns a 403 Forbidden response or closes the WebSocket with code 1008 and the reason "Maximum sessions reached".
Additionally, the timeout_seconds setting (default 300 seconds) automatically closes idle sessions, releasing their semaphore slots back to the pool. The RealtimeService monitors activity and triggers cleanup when the idle threshold expires.
Key Implementation Files
src/speech_to_speech/api/openai_realtime/runtime_config.py– Defines theRuntimeConfigdataclass, the_bool_from_envhelper, environment variable parsing, and thedefault_configsingleton.src/speech_to_speech/api/openai_realtime/pipeline_unit.py– Implements theRealtimeServiceclass and manages theasyncio.Semaphorefor session concurrency.src/speech_to_speech/connections/websocket_streamer.py– Wraps FastAPI WebSocket connections and handles binary audio frame streaming.demo/server.py– FastAPI entry point that injectsRuntimeConfigand mounts the WebSocket router.
Summary
- The Speech-to-Speech library uses an
asyncio.Semaphoreto enforce themax_concurrent_sessionslimit fromRuntimeConfig. - Configure the limit via the
S2S_MAX_CONCURRENT_SESSIONSenvironment variable or programmatically through theRuntimeConfigconstructor. - The
from_env()method provides type-safe parsing of environment variables using the_bool_from_envhelper for booleans. - Exceeded capacity results in WebSocket close code 1008 or a 403 response, while idle sessions automatically timeout after the configured
timeout_seconds. - Key files include
runtime_config.pyfor configuration andpipeline_unit.pyfor the semaphore-based session manager.
Frequently Asked Questions
How do I increase the maximum number of concurrent WebSocket sessions?
Set the S2S_MAX_CONCURRENT_SESSIONS environment variable to your desired integer value before starting the Uvicorn server. Alternatively, create a custom RuntimeConfig instance with a higher max_concurrent_sessions value and pass it to create_app() in demo/server.py.
What happens when the server reaches its concurrent session limit?
When the semaphore counter reaches zero, new connection attempts receive a 403 Forbidden response or the WebSocket closes with code 1008 and the reason "Maximum sessions reached". The client must retry after an existing session disconnects or times out.
How is the idle timeout configured for OpenAI Realtime sessions?
The S2S_TIMEOUT_SECONDS environment variable controls the idle timeout, defaulting to 300 seconds (5 minutes). The RealtimeService monitors session activity and automatically closes connections that exceed this idle threshold, releasing the semaphore slot for new clients.
Can I customize the RuntimeConfig programmatically instead of using environment variables?
Yes. Import RuntimeConfig from src/speech_to_speech/api/openai_realtime/runtime_config.py, instantiate it with your desired parameters, and pass it to the application factory. This approach is useful for testing or when embedding the Speech-to-Speech server within larger Python applications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →