# How to Configure Live Transcription and Transcription Events in the Realtime API

> Learn to configure live transcription and events in the Realtime API. Enable live transcription, set intervals, and emit partial and completed events for seamless audio processing with huggingface/speech-to-speech.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-08

---

**Enable live transcription by setting `enable_live_transcription=True` and configuring `live_transcription_update_interval` in `ModuleArguments`, which triggers the VAD handler to emit partial results that the `ConversationHandler` converts into OpenAI-compatible `delta` and `completed` events.**

The Hugging Face `speech-to-speech` repository implements an OpenAI-compatible Realtime API that streams audio through a speech-to-text pipeline. Configuring live transcription allows clients to receive progressive text updates as speech is processed, rather than waiting for final transcripts. This guide walks through the exact configuration flags and internal code paths that enable real-time transcription events in the `huggingface/speech-to-speech` project.

## Configuration Flags for Live Transcription

The system exposes two primary arguments in the `ModuleArguments` dataclass to control live transcription behavior.

### ModuleArguments Dataclass

Located in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) (lines 48-57), the configuration surface provides:

| Argument | Default | Effect |
|----------|---------|--------|
| `--enable_live_transcription` | `True` | Activates progressive STT output streaming. |
| `--live_transcription_update_interval` | `0.5` | Sets the interval in seconds between partial transcription emissions. |

When `enable_live_transcription` is set to `True`, the pipeline initializes the VAD handler to forward intermediate speech recognition results. The `live_transcription_update_interval` controls the throttling frequency for these updates.

## Pipeline Wiring and VAD Handler Setup

After parsing arguments, the pipeline builder in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) propagates these values to the VAD handler's runtime configuration.

### Configuring VADHandler

Around lines 29-33 of [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the pipeline logic maps the module arguments to the VAD handler keyword arguments:

```python

# s2s_pipeline.py – around line 29

if module_kwargs.enable_live_transcription:
    vad_kw.enable_realtime_transcription = True
    vad_kw.realtime_processing_pause = module_kwargs.live_transcription_update_interval

```

The `enable_realtime_transcription` flag instructs the `VADHandler` to process and forward every partial result. The `realtime_processing_pause` parameter, stored in the handler's internal state (lines 70-88 of [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py)), determines the cooldown period between successive partial transcription emissions.

## Processing Partial and Final Transcriptions

Inside the VAD processing loop, the `TranscriptionNotifier` component bridges raw STT output to the event system.

### TranscriptionNotifier Role

The `TranscriptionNotifier.process` method in [`src/speech_to_speech/STT/transcription_notifier.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/transcription_notifier.py) (lines 43-55) inspects incoming messages and transforms them into two internal types:

- **`PartialTranscription`** – Represents intermediate speech segments that become delta events.
- **`Transcription`** – Represents the finalized transcript for a completed turn.

When the notifier receives a `PartialTranscription` with non-empty text, it pushes the content onto the `text_output_queue`, which feeds into the realtime event pipeline.

## Realtime Protocol Integration

The `ConversationHandler` translates internal pipeline events into OpenAI-compatible realtime protocol messages.

### Event Types and Handlers

The handler implements two specific methods in [`src/speech_to_speech/api/openai_realtime/handlers/conversation.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/handlers/conversation.py):

**`on_partial_transcription`** (lines 100-110): Emits `conversation.item.input_audio_transcription.delta` events containing incremental text updates.

**`on_transcription_completed`** (lines 112-128): Emits `conversation.item.input_audio_transcription.completed` events when the speech segment ends.

Both methods generate properly typed event objects with fresh `event_id` values and the correct `item_id` mapping to the current input turn. The `RealtimeService` registers these handlers in its constructor (lines 216-218 of [`src/speech_to_speech/api/openai_realtime/service.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/service.py)) and forwards events to connected websocket clients.

## Implementation Examples

### Command Line Configuration

Run the pipeline with custom transcription intervals using the CLI:

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --enable_live_transcription true \
    --live_transcription_update_interval 0.3 \
    --vad_threshold 0.6 \
    --parakeet_tdt_stt_backend parakeet-tdt \
    --tts qwen3

```

### Programmatic Configuration

Configure the pipeline directly in Python:

```python
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
from speech_to_speech.s2s_pipeline import parse_arguments, build_pipeline, ThreadManager
from threading import Event
from queue import Queue

# Initialize with live transcription enabled

module_args = ModuleArguments(
    mode="realtime",
    enable_live_transcription=True,
    live_transcription_update_interval=0.25,
)

# Parse and inject arguments

parsed = parse_arguments()
parsed.module_kwargs = module_args

# Prepare communication queues

queues_and_events = {
    "stop_event": Event(),
    "should_listen": Event(),
    "recv_audio_chunks_queue": Queue(),
    "send_audio_chunks_queue": Queue(),
    "spoken_prompt_queue": Queue(),
    "stt_output_queue": Queue(),
    "text_prompt_queue": Queue(),
    "lm_response_queue": Queue(),
    "lm_processed_queue": Queue(),
}

# Build and start pipeline

manager = build_pipeline(**parsed.__dict__, queues_and_events=queues_and_events)
manager.start()

```

### Client-Side Event Handling

Connect to the realtime server and display transcription updates:

```bash
python scripts/listen_and_play_realtime.py \
    --ws_host localhost \
    --ws_port 8000

```

The client script registers websocket callbacks for `conversation.item.input_audio_transcription.delta` to display partial results and `conversation.item.input_audio_transcription.completed` to finalize the transcript.

## Summary

- **Enable live transcription** by setting `enable_live_transcription=True` in `ModuleArguments` (default is `True`).
- **Control update frequency** with `live_transcription_update_interval` (default 0.5 seconds).
- **Pipeline flow**: `ModuleArguments` → [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) → `VADHandler` → `TranscriptionNotifier` → `ConversationHandler` → `RealtimeService`.
- **Event types**: Clients receive `conversation.item.input_audio_transcription.delta` for partial updates and `conversation.item.input_audio_transcription.completed` for final transcripts.
- **Key files**: Configuration resides in [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py), processing logic in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py) and [`transcription_notifier.py`](https://github.com/huggingface/speech-to-speech/blob/main/transcription_notifier.py), and protocol translation in [`conversation.py`](https://github.com/huggingface/speech-to-speech/blob/main/conversation.py).

## Frequently Asked Questions

### How do I disable live transcription if I only want final transcripts?

Set `--enable_live_transcription false` when launching the pipeline. This prevents the `VADHandler` from emitting partial results, so clients receive only the `conversation.item.input_audio_transcription.completed` event at the end of each speech segment.

### What is the minimum practical value for `live_transcription_update_interval`?

While the system accepts any float value, intervals below 0.1 seconds may cause excessive websocket traffic and client-side processing overhead. The default of 0.5 seconds provides a balance between responsiveness and efficiency, though 0.25-0.3 seconds works well for low-latency applications.

### Where are the transcription events formatted into OpenAI protocol messages?

The `ConversationHandler` in [`src/speech_to_speech/api/openai_realtime/handlers/conversation.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/handlers/conversation.py) creates typed event objects. Specifically, `on_partial_transcription` generates `ConversationItemInputAudioTranscriptionDeltaEvent` objects and `on_transcription_completed` generates the corresponding `completed` events, both adhering to the OpenAI realtime API specification.

### Can I modify the transcription text before it reaches the client?

Yes, by intercepting the flow in `TranscriptionNotifier.process` (lines 43-55 of [`src/speech_to_speech/STT/transcription_notifier.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/transcription_notifier.py)). You can transform the text content of `PartialTranscription` or `Transcription` objects before they are queued for the `ConversationHandler`, enabling custom preprocessing such as profanity filtering or formatting.