# RuntimeConfig Deep-Merge Explained: How session.update Events Configure Speech-to-Speech Sessions

> Discover how RuntimeConfig in huggingface/speech-to-speech uses deep-merge with session.update events to efficiently configure speech-to-speech sessions without overwriting existing settings.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-10

---

**The `RuntimeConfig` class in `huggingface/speech-to-speech` serves as a mutable, shared state container that applies `session.update` events through a recursive deep-merge algorithm, preserving untouched fields while updating only explicitly provided values.**

The `RuntimeConfig` class forms the central configuration hub for the OpenAI Realtime API integration within the `huggingface/speech-to-speech` repository. It maintains the active conversational state and session parameters in memory, enabling dynamic reconfiguration without service restarts. When clients emit `session.update` events, the system executes a sophisticated deep-merge process that targets specific configuration paths while leaving existing nested structures intact.

## What is RuntimeConfig?

Located in [`src/speech_to_speech/api/openai_realtime/runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/runtime_config.py), the `RuntimeConfig` class acts as a thread-safe, mutable object that persists for the lifetime of the service. It bridges the gap between the OpenAI Realtime API and the local speech processing pipeline, providing handlers with instantaneous access to current parameters.

### Core Responsibilities

The configuration object stores two primary data structures:

1.  The **`Chat`** instance that drives the conversational flow and manages dialogue history.
2.  The **`RealtimeSessionCreateRequest`** (`session`) object, which mirrors the exact session payload sent to the OpenAI Realtime API.

Pipeline handlers—including Voice Activity Detection (VAD), Large Language Model (LLM), and Text-to-Speech (TTS) components—read from this shared object to obtain the latest instructions, audio formats, and turn detection settings.

### Mutable State Architecture

Because Python’s Global Interpreter Lock (GIL) makes simple attribute reads and writes atomic, `RuntimeConfig` operates without explicit locking mechanisms for primitive values. This design choice minimizes overhead while ensuring that concurrent pipeline stages can safely access configuration snapshots. The object initializes with default structures and remains mutable throughout the session, allowing real-time adjustments via WebSocket events.

## How session.update Deep-Merges Configuration

When a client transmits a `session.update` event, the payload contains only the fields requiring modification—not the complete configuration tree. The deep-merge mechanism ensures that omitted fields retain their existing values, even within deeply nested objects.

### The Recursive Merge Helper (`_apply_update`)

The private function `_apply_update` (lines 10–25 in [`runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/runtime_config.py)) implements the recursive merging logic:

- It iterates exclusively over the `model_fields_set` of the incoming `BaseModel` update, identifying only the fields explicitly included in the payload.
- For each field present in the update, it copies the value onto the corresponding attribute of the current configuration.
- When both the existing value and the new value are `BaseModel` instances, the function recurses to merge the nested objects rather than replacing them entirely.

This recursion prevents partial updates from wiping out untouched sub-fields. For example, updating `audio.input.turn_detection` does not destroy the existing `audio.output.volume` setting.

### The Public Interface (`apply_session_update`)

The `apply_session_update` method (lines 78–82) serves as the public entry point for configuration updates:

```python
def apply_session_update(self, session: RealtimeSessionCreateRequest) -> None:
    """Apply a partial update to the session configuration."""
    self._session = _apply_update(self._session, session)

```

This method receives a `RealtimeSessionCreateRequest` object representing the client’s update and delegates the merging process to `_apply_update`. The result is a new configuration state where only the specified fields have been overwritten.

### Handling Nested Objects

The deep-merge algorithm treats nested `BaseModel` objects as separate merge targets. If an update provides a partial definition for a nested structure—such as setting only `sample_rate` within `audio.input`—the merge preserves all other existing properties of that nested object (like `turn_detection` or `buffer_duration`). This granular approach allows clients to adjust specific parameters without resending the entire session configuration.

## Practical Implementation Examples

### Creating a Runtime Config and Updating It

```python
from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig
from openai.types.realtime import RealtimeSessionCreateRequest
from openai.types.realtime.realtime_audio_config import RealtimeAudioConfig
from openai.types.realtime.realtime_audio_config_input import RealtimeAudioConfigInput

# Initialise the mutable config

rt_cfg = RuntimeConfig()

# Current session state (initially empty)

print(rt_cfg.session.audio.input.turn_detection)   # → None (default)

# Simulate a session.update payload that only changes turn_detection.interrupt_response

update = RealtimeSessionCreateRequest(
    audio=RealtimeAudioConfig(
        input=RealtimeAudioConfigInput(
            turn_detection={"interrupt_response": False}
        )
    )
)

# Apply the update – only the provided field is merged, all others stay unchanged

rt_cfg.apply_session_update(update)

# Verify the deep‑merged result

print(rt_cfg.session.audio.input.turn_detection.interrupt_response)  # → False

```

### Partial Nested Update Preserving Untouched Fields

```python

# Existing config already has an output volume set

rt_cfg.session.audio.output = RealtimeAudioConfigOutput(volume=0.8)

# New update only touches the input sample_rate

partial_update = RealtimeSessionCreateRequest(
    audio=RealtimeAudioConfig(
        input=RealtimeAudioConfigInput(sample_rate=16000)
    )
)

rt_cfg.apply_session_update(partial_update)

# The output volume is still present

assert rt_cfg.session.audio.output.volume == 0.8

# The new input sample_rate is applied

assert rt_cfg.session.audio.input.sample_rate == 16000

```

These examples demonstrate how the deep-merge preserves previously established configuration branches while integrating new values received via `session.update`.

## Validation and Safety Guarantees

To prevent invalid states, the `RuntimeConfig` implementation includes the `_ensure_audio_structure` validator (lines 46–56 in [`runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/runtime_config.py)). This validator ensures that audio sub-objects within the session configuration are never set to `None`, maintaining the structural integrity required by downstream audio processing components. Even when partial updates omit audio fields, the validation guarantees that the existing audio configuration remains valid and accessible.

## Summary

- **RuntimeConfig** acts as a mutable, in-memory state container for OpenAI Realtime sessions, storing the `Chat` instance and `RealtimeSessionCreateRequest` payload.
- The **`_apply_update`** helper function recursively merges incoming `session.update` events by iterating over `model_fields_set` and preserving untouched nested fields.
- The public **`apply_session_update`** method provides the interface for applying partial configuration updates without requiring full session restarts.
- **Deep-merge logic** ensures that pipeline handlers always access consistent, valid configuration states, even during concurrent updates to nested structures.
- The **`_ensure_audio_structure`** validator maintains audio object integrity, preventing `None` values in critical audio configuration paths.

## Frequently Asked Questions

### What is the purpose of RuntimeConfig in the speech-to-speech pipeline?

**RuntimeConfig** serves as a centralized, mutable configuration store that maintains the current state of OpenAI Realtime sessions throughout the service lifecycle. It holds the `Chat` instance managing conversation flow and the `RealtimeSessionCreateRequest` object that mirrors the API payload, allowing pipeline components like VAD, LLM, and TTS handlers to access the latest settings without restarting the service.

### How does the deep-merge process differ from a standard dictionary update?

Unlike a standard dictionary update that would replace entire nested objects, the deep-merge implemented in `_apply_update` recursively traverses only the fields explicitly set in the incoming `BaseModel` (identified by `model_fields_set`). When both existing and new values are `BaseModel` instances, the function merges them recursively, preserving sub-fields that were not included in the update payload rather than overwriting the entire parent object.

### Is RuntimeConfig thread-safe for concurrent access by multiple handlers?

Yes, **RuntimeConfig** is designed for concurrent access without explicit locking mechanisms. Simple attribute reads and writes to primitive values are atomic operations protected by Python's Global Interpreter Lock (GIL). However, complex deep-merge operations via `apply_session_update` should be executed sequentially to prevent race conditions during nested object modifications.

### What ensures that audio configuration remains valid after partial updates?

The **`_ensure_audio_structure`** validator, defined at lines 46–56 in [`src/speech_to_speech/api/openai_realtime/runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/runtime_config.py), automatically validates that audio sub-objects within the session configuration are never set to `None`. This safety mechanism ensures that even when clients send partial updates that omit audio fields, the configuration maintains valid audio structures required by the speech processing pipeline.