How to Configure Live Transcription and Transcription Events in the Realtime API

Enable live transcription by setting enable_live_transcription=True and configuring live_transcription_update_interval in ModuleArguments, which triggers the VAD handler to emit partial results that the ConversationHandler converts into OpenAI-compatible delta and completed events.

The Hugging Face speech-to-speech repository implements an OpenAI-compatible Realtime API that streams audio through a speech-to-text pipeline. Configuring live transcription allows clients to receive progressive text updates as speech is processed, rather than waiting for final transcripts. This guide walks through the exact configuration flags and internal code paths that enable real-time transcription events in the huggingface/speech-to-speech project.

Configuration Flags for Live Transcription

The system exposes two primary arguments in the ModuleArguments dataclass to control live transcription behavior.

ModuleArguments Dataclass

Located in src/speech_to_speech/arguments_classes/module_arguments.py (lines 48-57), the configuration surface provides:

Argument Default Effect
--enable_live_transcription True Activates progressive STT output streaming.
--live_transcription_update_interval 0.5 Sets the interval in seconds between partial transcription emissions.

When enable_live_transcription is set to True, the pipeline initializes the VAD handler to forward intermediate speech recognition results. The live_transcription_update_interval controls the throttling frequency for these updates.

Pipeline Wiring and VAD Handler Setup

After parsing arguments, the pipeline builder in s2s_pipeline.py propagates these values to the VAD handler's runtime configuration.

Configuring VADHandler

Around lines 29-33 of src/speech_to_speech/s2s_pipeline.py, the pipeline logic maps the module arguments to the VAD handler keyword arguments:


# s2s_pipeline.py – around line 29

if module_kwargs.enable_live_transcription:
    vad_kw.enable_realtime_transcription = True
    vad_kw.realtime_processing_pause = module_kwargs.live_transcription_update_interval

The enable_realtime_transcription flag instructs the VADHandler to process and forward every partial result. The realtime_processing_pause parameter, stored in the handler's internal state (lines 70-88 of src/speech_to_speech/VAD/vad_handler.py), determines the cooldown period between successive partial transcription emissions.

Processing Partial and Final Transcriptions

Inside the VAD processing loop, the TranscriptionNotifier component bridges raw STT output to the event system.

TranscriptionNotifier Role

The TranscriptionNotifier.process method in src/speech_to_speech/STT/transcription_notifier.py (lines 43-55) inspects incoming messages and transforms them into two internal types:

  • PartialTranscription – Represents intermediate speech segments that become delta events.
  • Transcription – Represents the finalized transcript for a completed turn.

When the notifier receives a PartialTranscription with non-empty text, it pushes the content onto the text_output_queue, which feeds into the realtime event pipeline.

Realtime Protocol Integration

The ConversationHandler translates internal pipeline events into OpenAI-compatible realtime protocol messages.

Event Types and Handlers

The handler implements two specific methods in src/speech_to_speech/api/openai_realtime/handlers/conversation.py:

on_partial_transcription (lines 100-110): Emits conversation.item.input_audio_transcription.delta events containing incremental text updates.

on_transcription_completed (lines 112-128): Emits conversation.item.input_audio_transcription.completed events when the speech segment ends.

Both methods generate properly typed event objects with fresh event_id values and the correct item_id mapping to the current input turn. The RealtimeService registers these handlers in its constructor (lines 216-218 of src/speech_to_speech/api/openai_realtime/service.py) and forwards events to connected websocket clients.

Implementation Examples

Command Line Configuration

Run the pipeline with custom transcription intervals using the CLI:

python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --enable_live_transcription true \
    --live_transcription_update_interval 0.3 \
    --vad_threshold 0.6 \
    --parakeet_tdt_stt_backend parakeet-tdt \
    --tts qwen3

Programmatic Configuration

Configure the pipeline directly in Python:

from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
from speech_to_speech.s2s_pipeline import parse_arguments, build_pipeline, ThreadManager
from threading import Event
from queue import Queue

# Initialize with live transcription enabled

module_args = ModuleArguments(
    mode="realtime",
    enable_live_transcription=True,
    live_transcription_update_interval=0.25,
)

# Parse and inject arguments

parsed = parse_arguments()
parsed.module_kwargs = module_args

# Prepare communication queues

queues_and_events = {
    "stop_event": Event(),
    "should_listen": Event(),
    "recv_audio_chunks_queue": Queue(),
    "send_audio_chunks_queue": Queue(),
    "spoken_prompt_queue": Queue(),
    "stt_output_queue": Queue(),
    "text_prompt_queue": Queue(),
    "lm_response_queue": Queue(),
    "lm_processed_queue": Queue(),
}

# Build and start pipeline

manager = build_pipeline(**parsed.__dict__, queues_and_events=queues_and_events)
manager.start()

Client-Side Event Handling

Connect to the realtime server and display transcription updates:

python scripts/listen_and_play_realtime.py \
    --ws_host localhost \
    --ws_port 8000

The client script registers websocket callbacks for conversation.item.input_audio_transcription.delta to display partial results and conversation.item.input_audio_transcription.completed to finalize the transcript.

Summary

  • Enable live transcription by setting enable_live_transcription=True in ModuleArguments (default is True).
  • Control update frequency with live_transcription_update_interval (default 0.5 seconds).
  • Pipeline flow: ModuleArguments → s2s_pipeline.py → VADHandler → TranscriptionNotifier → ConversationHandler → RealtimeService.
  • Event types: Clients receive conversation.item.input_audio_transcription.delta for partial updates and conversation.item.input_audio_transcription.completed for final transcripts.
  • Key files: Configuration resides in module_arguments.py, processing logic in vad_handler.py and transcription_notifier.py, and protocol translation in conversation.py.

Frequently Asked Questions

How do I disable live transcription if I only want final transcripts?

Set --enable_live_transcription false when launching the pipeline. This prevents the VADHandler from emitting partial results, so clients receive only the conversation.item.input_audio_transcription.completed event at the end of each speech segment.

What is the minimum practical value for live_transcription_update_interval?

While the system accepts any float value, intervals below 0.1 seconds may cause excessive websocket traffic and client-side processing overhead. The default of 0.5 seconds provides a balance between responsiveness and efficiency, though 0.25-0.3 seconds works well for low-latency applications.

Where are the transcription events formatted into OpenAI protocol messages?

The ConversationHandler in src/speech_to_speech/api/openai_realtime/handlers/conversation.py creates typed event objects. Specifically, on_partial_transcription generates ConversationItemInputAudioTranscriptionDeltaEvent objects and on_transcription_completed generates the corresponding completed events, both adhering to the OpenAI realtime API specification.

Can I modify the transcription text before it reaches the client?

Yes, by intercepting the flow in TranscriptionNotifier.process (lines 43-55 of src/speech_to_speech/STT/transcription_notifier.py). You can transform the text content of PartialTranscription or Transcription objects before they are queued for the ConversationHandler, enabling custom preprocessing such as profanity filtering or formatting.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →