What Is RuntimeConfig in Speech-to-Speech? Configuring Session Parameters

RuntimeConfig is a mutable, shared configuration object that stores Realtime session state and conversation history for each WebSocket connection in the Speech-to-Speech service.

Defined in the huggingface/speech-to-speech repository, RuntimeConfig serves as the central hub for managing session parameters in OpenAI Realtime-compatible pipelines. It encapsulates audio configuration settings, turn detection parameters, and conversation buffers that persist throughout a connection's lifecycle.

RuntimeConfig Structure and Default Values

The RuntimeConfig class is implemented in [src/speech_to_speech/api/openai_realtime/runtime_config.py](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/runtime_config.py) as a Pydantic BaseModel. It contains two primary fields that define the runtime behavior of a speech-to-speech session.

Core Fields

The model declaration (lines 27–45) defines:

class RuntimeConfig(BaseModel):
    chat: Chat = Field(default_factory=lambda: Chat(10))
    session: RealtimeSessionCreateRequest = Field(
        default_factory=lambda: RealtimeSessionCreateRequest(type="realtime"),
        validate_default=True,
    )
  • chat: A Chat instance that holds conversation history with a default buffer size of 10 turns.
  • session: A RealtimeSessionCreateRequest object that mirrors the full OpenAI Realtime session payload, including audio input/output configurations and tool definitions.

Audio Structure Validation

A Pydantic field validator (_ensure_audio_structure, lines 46–57) guarantees that nested audio configuration objects are never None. This protects downstream handlers from null reference errors when accessing audio parameters.

@field_validator("session", mode="after")
@classmethod
def _ensure_audio_structure(cls, v: RealtimeSessionCreateRequest) -> RealtimeSessionCreateRequest:
    if v.audio is None:
        v.audio = RealtimeAudioConfig()
    if v.audio.input is None:
        v.audio.input = RealtimeAudioConfigInput()
    if v.audio.output is None:
        v.audio.output = RealtimeAudioConfigOutput()
    return v

This validator ensures that session.audio, session.audio.input, and session.audio.output always contain valid objects, even if the client provides an incomplete configuration.

Configuring Session Parameters

RuntimeConfig provides mechanisms for reading critical session parameters and applying incremental updates without replacing the entire configuration object.

Interrupt Response Handling

The interrupt_response_enabled property (lines 58–77) extracts the turn_detection.interrupt_response flag from the session's audio input configuration. It defaults to True to match the OpenAI API default and handles both Pydantic models and raw dictionaries for robustness.

@property
def interrupt_response_enabled(self) -> bool:
    td = self.session.audio.input.turn_detection
    # Handles missing fields, dicts, or model objects

    ...

Applying Incremental Updates

When a client sends a session.update event, RuntimeConfig.apply_session_update (lines 78–82) delegates to an internal helper that performs a selective merge. Only the fields explicitly set in the update (tracked via update.model_fields_set) overwrite the current session values.

def apply_session_update(self, update: RealtimeSessionCreateRequest) -> None:
    _apply_update(self.session, update)

This approach preserves existing configuration values while allowing fine-grained adjustments to specific parameters like voice settings or turn detection thresholds.

Integration with the Speech-to-Speech Pipeline

RuntimeConfig bridges the gap between the WebSocket API layer and the audio processing pipeline, ensuring that each connection maintains isolated, mutable state.

Per-Connection State Management

In [src/speech_to_speech/api/openai_realtime/service.py](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/service.py), the ConnState class embeds RuntimeConfig as a field:

runtime_config: RuntimeConfig = Field(default_factory=RuntimeConfig)

This design ensures that every WebSocket connection receives its own RuntimeConfig instance, preventing cross-contamination of session parameters between concurrent users.

Handler Access Patterns

Pipeline handlers—such as VAD, LLM, and TTS components—read from runtime_config.session to obtain processing parameters. For example, the VAD handler in [src/speech_to_speech/VAD/vad_handler.py](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) accesses runtime_config.session.audio to apply turn-detection settings. Because the Python GIL guarantees atomicity for simple attribute reads and writes, no additional locking is required for these primitive values, making the configuration lightweight and fast to access across threads.

Practical Usage Examples

Creating a RuntimeConfig Manually

To instantiate RuntimeConfig with a custom conversation buffer size:

from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig
from speech_to_speech.LLM.chat import Chat

# Custom chat buffer size (e.g., keep last 20 turns)

cfg = RuntimeConfig(chat=Chat(20))

print(cfg.session.type)           # → "realtime"

print(cfg.session.audio.input)  # always a RealtimeAudioConfigInput instance

Updating Session Parameters at Runtime

To modify turn-detection settings without replacing the entire configuration:

from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig
from openai.types.realtime import RealtimeSessionCreateRequest
from openai.types.realtime.realtime_audio_config_input import RealtimeAudioConfigInput

cfg = RuntimeConfig()  # default instance

# Build a partial update that only changes turn‑detection settings

partial_update = RealtimeSessionCreateRequest(
    audio=RealtimeAudioConfigInput(
        input=RealtimeAudioConfigInput(
            turn_detection={"type": "server_vad", "interrupt_response": False}
        )
    )
)

# Apply the update – only the explicitly set fields are merged

cfg.apply_session_update(partial_update)

print(cfg.interrupt_response_enabled)  # → False

Accessing Config from a Connection Handler

Inside a Realtime service handler, access the configuration through the connection state:


# Inside a RealtimeService handler

state = self._state(conn_id)          # ConnState instance

rt_cfg = state.runtime_config

# Example: decide whether to abort an ongoing response

if rt_cfg.interrupt_response_enabled and event.interrupt_response:
    # cancel the current response

    ...

Summary

  • RuntimeConfig is a Pydantic BaseModel defined in runtime_config.py that manages per-connection session state and conversation history in the Speech-to-Speech service.
  • It contains a chat field (buffering conversation turns) and a session field (mirroring OpenAI Realtime API structure), with validators ensuring nested audio objects are never None.
  • The interrupt_response_enabled property provides convenient access to turn-detection settings, defaulting to True.
  • apply_session_update merges partial updates selectively, preserving existing configuration while applying only explicit changes.
  • Each WebSocket connection stores its own RuntimeConfig instance via ConnState, enabling isolated session management without requiring additional thread locks.

Frequently Asked Questions

What fields does RuntimeConfig manage in speech-to-speech?

RuntimeConfig manages two primary fields: chat, which stores the conversation history buffer (default size 10), and session, which holds the RealtimeSessionCreateRequest object containing audio input/output configurations, turn detection parameters, and tool definitions for the active Realtime session.

How does RuntimeConfig handle partial session updates?

When apply_session_update receives a RealtimeSessionCreateRequest update, it delegates to an internal helper that walks the nested model tree and copies only the fields present in update.model_fields_set. This ensures that unspecified fields retain their existing values rather than being overwritten with defaults.

Is RuntimeConfig thread-safe for concurrent access?

Yes, RuntimeConfig relies on Python's Global Interpreter Lock (GIL), which guarantees that simple attribute reads and writes on primitive values are atomic. Therefore, no additional locking mechanisms are required when pipeline handlers access configuration values, making the design lightweight for multi-threaded audio processing.

Where is RuntimeConfig instantiated in the codebase?

RuntimeConfig is instantiated per WebSocket connection in [src/speech_to_speech/api/openai_realtime/service.py](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/service.py) within the ConnState class, which uses default_factory=RuntimeConfig to ensure each connection receives an isolated configuration instance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →