# OpenAI Realtime-Compatible Server WebSocket Event Types: Complete Reference

> Explore all 20 OpenAI Realtime WebSocket event types implemented by Hugging Face speech-to-speech. Get a complete reference for streaming audio, managing conversations, and LLM responses.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: api-reference
- Published: 2026-08-08

---

**The Hugging Face `speech-to-speech` repository implements the full OpenAI Realtime API specification, exposing 20 distinct WebSocket event types for bidirectional streaming audio, conversation management, and LLM response handling.**

The `huggingface/speech-to-speech` project provides an open-source OpenAI Realtime-compatible server that enables low-latency, bidirectional speech-to-speech interactions through WebSocket connections. This server mirrors the official OpenAI Realtime protocol, making it compatible with existing OpenAI client SDKs while integrating custom pipeline components for voice activity detection (VAD), text-to-speech (TTS), and large language model (LLM) inference. Understanding the exact WebSocket event types supported by this server is essential for building compliant clients or debugging conversation flows.

## Client-to-Server Event Types

The server accepts specific JSON event payloads from clients to control session parameters, manage conversation history, and stream audio input. These events are validated and dispatched by the `WebSocketRouter` class defined in [`tests/openai_realtime/test_websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_websocket_router.py).

### Session Management

**`session.update`** modifies runtime configuration such as voice selection, temperature, and turn detection parameters. The router handles this via `_handle_session_update`, applying changes immediately and echoing confirmation back to the client.

### Conversation Item Control

Clients manipulate conversation history using four discrete event types:

- **`conversation.item.create`** – Initializes a new message object (user or assistant role) via `_handle_conversation_item_create`.
- **`conversation.item.truncate`** – Shortens an existing item's content, processed by `_handle_conversation_item_truncate` to handle partial transcription cleanup.
- **`conversation.item.delete`** – Removes a specific item from context using `_handle_conversation_item_delete`.
- **`conversation.item.input_audio`** – Streams base64-encoded PCM audio chunks associated with a user turn, handled by the audio subsystem referenced in [`tests/openai_realtime/test_audio_client.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_audio_client.py) (see `_send_audio_frames`).

### Response Generation

To trigger or halt LLM inference:

- **`response.create`** – Requests the server to generate a response, optionally including function-call definitions. The router delegates to `_handle_response_create`.
- **`response.cancel`** – Aborts an in-flight generation, implemented in `_handle_response_cancel`.

### Connection Keep-Alive

**`heartbeat`** maintains persistent connections through the `_handle_heartbeat` handler, preventing timeout disconnections during silent periods.

## Server-to-Client Event Types

The server emits typed events to signal state changes, stream generated content, and report errors. These events are constructed using classes defined in [`pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/pipeline/events.py) and emitted through the `EventQueue` system.

### Session Lifecycle

Upon connection establishment:

- **`session.created`** – Sent immediately after handshake completion, confirming session initialization (see `_emit_session_created` in [`tests/openai_realtime/test_realtime_service.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_realtime_service.py)).
- **`session.updated`** – Echoes configuration changes applied via `session.update` (see `_emit_session_updated`).

### Conversation State Changes

As the conversation context mutates, the server broadcasts:

- **`conversation.item.created`** – Confirms addition of new items (see `_emit_item_created` in [`tests/openai_realtime/test_conversation_events.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_conversation_events.py)).
- **`conversation.item.truncated`** – Signals that an item's content was shortened (see `_emit_item_truncated`).
- **`conversation.item.deleted`** – Acknowledges item removal (see `_emit_item_deleted`).

### Response Streaming

During LLM generation and audio synthesis:

- **`response.created`** – Marks the start of response processing (see `_emit_response_created` in [`tests/openai_realtime/test_response_events.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_response_events.py)).
- **`response.output_text.delta`** – Transmits incremental text fragments as they are generated by the LLM (see `_emit_text_delta`).
- **`response.output_audio.delta`** – Streams synthesized audio chunks back to the client (see `_emit_audio_delta`).
- **`response.done`** – Indicates completion of the entire response, including any function-call results (see `_emit_response_done`).

### Audio Buffer Management

For handling real-time audio input:

- **`input_audio_buffer.commit`** – Acknowledges receipt and buffering of a user audio chunk (see `_emit_buffer_commit` in [`tests/openai_realtime/test_audio_client.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_audio_client.py)).
- **`audio.delta`** – Continuous server-side audio playback feed used for TTS streaming (see `_emit_audio_delta` in the audio output notifier).

### Error Handling

**`error`** events report protocol violations, invalid message formats, or internal pipeline failures. These are structured as `ErrorEvent` objects and emitted through the standard event queue.

## Implementation Architecture

The event protocol is implemented across several key modules. The `WebSocketRouter` in [`tests/openai_realtime/test_websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_websocket_router.py) validates incoming JSON against supported event types, dispatching to private handler methods (e.g., `_handle_session_update`). Outbound events are managed by an `EventQueue` defined in [`pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/pipeline/events.py), which serializes typed event objects into WebSocket frames. Graceful shutdown is coordinated through a `stop_event` mechanism (referenced in [`tests/openai_realtime/test_runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_runtime_config.py)), ensuring final error messages or close frames are transmitted before disconnection.

## Practical Integration Examples

The repository includes a reference JavaScript client at [`demo/ws/s2s-ws-client.js`](https://github.com/huggingface/speech-to-speech/blob/main/demo/ws/s2s-ws-client.js) demonstrating proper event usage.

Initializing a session and configuring parameters:

```javascript
const ws = new WebSocket('ws://localhost:8000/ws');
ws.addEventListener('open', () => {
  ws.send(JSON.stringify({
    type: 'session.update',
    session: { voice: 'alloy', temperature: 0.7 }
  }));
});
ws.addEventListener('message', (e) => {
  const event = JSON.parse(e.data);
  if (event.type === 'session.updated') {
    console.log('Session parameters applied:', event.session);
  }
});

```

Streaming microphone audio to the server:

```javascript
navigator.mediaDevices.getUserMedia({ audio: true }).then((stream) => {
  const recorder = new MediaRecorder(stream);
  recorder.ondataavailable = (e) => {
    const reader = new FileReader();
    reader.onloadend = () => {
      ws.send(JSON.stringify({
        type: 'conversation.item.input_audio',
        item_id: 'user-turn-1',
        audio: reader.result.split(',')[1] // base64 data
      }));
    };
    reader.readAsDataURL(e.data);
  };
  recorder.start(100); // 100ms chunks
});

```

Handling streaming responses:

```javascript
ws.addEventListener('message', (e) => {
  const ev = JSON.parse(e.data);
  switch (ev.type) {
    case 'response.created':
      console.log('Generation started');
      break;
    case 'response.output_text.delta':
      appendTextToUI(ev.delta);
      break;
    case 'response.output_audio.delta':
      playAudioChunk(ev.delta); // base64 PCM
      break;
    case 'response.done':
      console.log('Response complete');
      break;
    case 'error':
      console.error('Server error:', ev.error);
      break;
  }
});

```

## Summary

- The **OpenAI Realtime-compatible server** in `huggingface/speech-to-speech` supports **20 distinct WebSocket event types** covering session control, conversation management, audio streaming, and error reporting.
- **Client-to-server events** (8 types) are routed through `WebSocketRouter` in [`tests/openai_realtime/test_websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_websocket_router.py), with specific handlers like `_handle_session_update` and `_handle_conversation_item_create`.
- **Server-to-client events** (12 types) include lifecycle signals (`session.created`), incremental streaming (`response.output_text.delta`, `response.output_audio.delta`), and buffer management (`input_audio_buffer.commit`).
- All events are typed and queued through the `EventQueue` system defined in [`pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/pipeline/events.py), ensuring type-safe serialization over WebSocket connections.
- The reference implementation in [`demo/ws/s2s-ws-client.js`](https://github.com/huggingface/speech-to-speech/blob/main/demo/ws/s2s-ws-client.js) demonstrates compatible client-side usage of the full event protocol.

## Frequently Asked Questions

### What audio format does the `conversation.item.input_audio` event expect?

The server expects **base64-encoded PCM16 audio** at 24kHz sample rate, transmitted as a string in the `audio` field of the event payload. The [`demo/ws/s2s-ws-client.js`](https://github.com/huggingface/speech-to-speech/blob/main/demo/ws/s2s-ws-client.js) file demonstrates proper encoding using `FileReader` to convert `Blob` objects from `MediaRecorder` into the required format.

### How does the server handle unsupported or malformed event types?

Invalid events trigger an **`error`** event response formatted according to the `ErrorEvent` class in [`pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/pipeline/events.py). The `WebSocketRouter` validates the `type` field against the supported enumeration; unrecognized types immediately return an error frame without modifying conversation state.

### Can I cancel a response after calling `response.create`?

Yes. Clients may send a **`response.cancel`** event at any time during active generation. The server processes this through `_handle_response_cancel` in the router, immediately terminating the LLM inference and TTS synthesis pipelines, and emits a final `response.done` event to confirm cancellation.

### What is the difference between `response.output_audio.delta` and `audio.delta`?

**`response.output_audio.delta`** carries synthesized speech chunks associated with a specific LLM response turn, emitted during active generation. In contrast, **`audio.delta`** represents continuous server-side audio playback buffering, used primarily for maintaining low-latency audio streams independent of specific response boundaries. Both are handled in [`tests/openai_realtime/test_response_events.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_response_events.py) and [`test_audio_client.py`](https://github.com/huggingface/speech-to-speech/blob/main/test_audio_client.py) respectively.