Understanding RuntimeConfig and session.update in Hugging Face Speech-to-Speech

RuntimeConfig is a mutable, thread-safe configuration container that stores conversation state and session parameters, while session.update merges partial changes in-place to dynamically adjust VAD, LLM, and TTS behavior without restarting the pipeline.

In the huggingface/speech-to-speech repository, real-time speech processing relies on a central configuration object to coordinate between voice activity detection, language models, and text-to-speech components. RuntimeConfig serves as this mutable state container defined in src/speech_to_speech/api/openai_realtime/runtime_config.py, enabling the session.update mechanism to modify pipeline behavior mid-conversation while maintaining thread safety and validation.

What is RuntimeConfig?

RuntimeConfig acts as the single source of truth for session-level configuration inside each realtime connection. It is implemented as a Pydantic BaseModel with validate_assignment=True (lines 27-45), ensuring that any attribute modification triggers immediate validation.

Conversation State and Session Storage

The class maintains two primary fields:

  • chat: A Chat object that tracks conversation history and manages token windows for the language model context.
  • session: A complete RealtimeSessionCreateRequest model that mirrors the OpenAI Realtime API specification, containing audio settings, tool definitions, turn-detection parameters, and model configuration.

Because these fields live in a shared, mutable container, components across the pipeline can read the current configuration without passing explicit parameters through every function call.

Thread Safety and Validation

The Pydantic-based implementation ensures that updates are type-checked immediately upon assignment. This prevents invalid configuration states from propagating to downstream handlers like the VAD or TTS components, which depend on specific schema structures to function correctly.

How session.update Modifies Pipeline Behavior

When an OpenAI Realtime client sends a session.update event, the system must apply these changes without disrupting the active audio stream or conversation context. This is handled through a specific update mechanism that preserves unspecified fields while applying new values atomically.

Receiving Updates via SessionHandler

The SessionHandler in src/speech_to_speech/api/openai_realtime/handlers/session.py (lines 23-31) receives incoming events and routes session.update payloads to RuntimeConfig.apply_session_update. This method initiates the merge process that updates the stored configuration.

The Merge Algorithm

The actual merge logic resides in the private helper _apply_update in runtime_config.py (lines 10-24). This function implements a recursive merge strategy with specific semantics:

  1. Explicit fields only: Only fields present in the incoming model's model_fields_set are considered—silent or default fields are ignored.
  2. Recursive merging: If a field contains a nested BaseModel, the function recurses into that object rather than replacing it entirely, preserving sibling fields that weren't included in the update.
  3. Scalar replacement: For primitive values, the new value (including explicit None) replaces the old value directly.
  4. In-place modification: The existing session object is mutated rather than replaced, ensuring that references held by other components remain valid.

Real-Time Impact on Pipeline Components

All pipeline stages consult RuntimeConfig during processing, meaning session.update changes take effect immediately:

  • Voice Activity Detection (VAD): The handler in src/speech_to_speech/VAD/vad_handler.py checks runtime_config.interrupt_response_enabled (derived from session.audio.input.turn_detection) to determine whether new user audio should abort an ongoing assistant response.
  • Language Model: Components reading from runtime_config.session respect updated parameters like tool_choice, temperature, and dynamically added tool definitions when generating the next response.
  • Text-to-Speech: TTS handlers may reference session.audio.output.voice to select the speaking voice for synthesized audio, allowing clients to change voices mid-conversation.

Working with RuntimeConfig: Code Example

The following example demonstrates creating a configuration instance, inspecting defaults, and applying partial updates that modify pipeline behavior:


# 1️⃣ Create a fresh RuntimeConfig (default audio structure is ensured)

from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig
cfg = RuntimeConfig()

# 2️⃣ Inspect the default voice (uses “alloy” as placeholder)

print(cfg.session.audio.output.voice)   # → None (will be set by the client)

# 3️⃣ Apply a partial update coming from a session.update event

from openai.types.realtime import RealtimeSessionCreateRequest
update = RealtimeSessionCreateRequest(
    audio={"output": {"voice": "coral"}},   # change only the output voice

    tool_choice="auto"
)
cfg.apply_session_update(update)

# 4️⃣ The change is visible to all components

print(cfg.session.audio.output.voice)   # → "coral"

print(cfg.session.tool_choice)          # → "auto"

# 5️⃣ Later, a client sends another update that clears turn_detection

clear_update = RealtimeSessionCreateRequest(
    audio={"input": {"turn_detection": None}}
)
cfg.apply_session_update(clear_update)

# 6️⃣ VAD now sees `turn_detection` as None, affecting interrupt logic

print(cfg.session.audio.input.turn_detection)  # → None

Summary

  • RuntimeConfig stores the Chat object and RealtimeSessionCreateRequest in a thread-safe, validated container at src/speech_to_speech/api/openai_realtime/runtime_config.py.
  • session.update events are processed by SessionHandler and applied via RuntimeConfig.apply_session_update, which delegates to _apply_update.
  • The merge algorithm preserves untouched fields recursively while allowing explicit None values to clear existing configuration.
  • Changes immediately affect VAD interruption logic, LLM parameters (tools, temperature), and TTS voice selection without requiring pipeline restarts.

Frequently Asked Questions

Where is RuntimeConfig defined in the codebase?

RuntimeConfig is defined in src/speech_to_speech/api/openai_realtime/runtime_config.py at lines 27-45. This file also contains the private _apply_update helper function (lines 10-24) that implements the recursive merge logic used by apply_session_update.

How does session.update handle partial configuration changes?

The apply_session_update method merges only fields explicitly set in the incoming request (identified by model_fields_set), leaving all other existing configuration values intact. For nested objects like audio.input, it recurses into the structure to update specific children without overwriting sibling fields that weren't included in the update payload.

Can RuntimeConfig changes affect an ongoing conversation?

Yes. Because RuntimeConfig is a mutable, shared reference that pipeline components read during processing, any session.update that arrives mid-conversation instantly influences subsequent steps. For example, changing turn_detection immediately alters VAD behavior for the next audio chunk, and modifying tool_choice affects the next LLM inference call.

What happens when a client sends an explicit null value in session.update?

The _apply_update helper treats explicit None values as intentional clearing operations. When the incoming model has None set for a field in model_fields_set, the function replaces the existing value with None, effectively removing that configuration parameter for downstream components.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →