How to Enable and Configure Live Transcription Events in Hugging Face Speech-to-Speech
Enable live transcription events by setting enable_live_transcription=True in ModuleArguments or using the --enable-live-transcription CLI flag, and configure the streaming interval with live_transcription_update_interval (default 0.5 seconds).
Live transcription events stream partial recognition results to your application while the user is still speaking, enabling real-time feedback. In the Hugging Face speech-to-speech repository, this feature is implemented through a coordinated pipeline involving voice activity detection (VAD) handlers and event notifiers. This guide covers how to enable and configure live transcription events using both the command-line interface and the Python API.
What Are Live Transcription Events?
Live transcription events (also called partial transcriptions) are incremental speech recognition updates emitted before the final transcript is complete. Unlike final transcription events that fire after speech ends, these partial updates allow you to build "live-typing" interfaces. According to the source code, the system emits partial_transcription events through the TranscriptionNotifier class while the VADHandler processes audio chunks in realtime mode.
Configuration Options in ModuleArguments
The entry point for live transcription configuration is the ModuleArguments class in src/speech_to_speech/arguments_classes/module_arguments.py (line 49). This dataclass exposes two critical parameters:
enable_live_transcription: Boolean flag to toggle the feature (defaults toTrue)live_transcription_update_interval: Float specifying the seconds between partial updates (defaults to0.5)
How to Enable Live Transcription Events
Via Command Line Interface
To enable live transcription when running the demo server, append the flag and optional interval parameter:
python -m speech_to_speech.demo.server \
--enable-live-transcription \
--live-transcription-update-interval 0.3
Via Python API
Programmatically, pass the configuration through module_kwargs when constructing the pipeline:
from speech_to_speech import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
args = ModuleArguments(
enable_live_transcription=True,
live_transcription_update_interval=0.3,
)
pipeline = SpeechToSpeechPipeline(module_kwargs=args)
Configuring the Update Interval
The live_transcription_update_interval controls how frequently the system emits partial results. Lower values (e.g., 0.1-0.2 seconds) provide more responsive feedback but increase computational overhead. The default of 0.5 seconds balances responsiveness with resource usage. This value is propagated through S2SPipeline in src/speech_to_speech/s2s_pipeline.py (lines 557-572) to set the realtime_processing_pause on the VAD handler.
How Live Transcription Works Under the Hood
When enabled, the pipeline activates three coordinated components:
VAD Handler Realtime Mode
In src/speech_to_speech/STT/parakeet_tdt_handler.py (lines 127-132), the VAD handler checks enable_live_transcription and switches to realtime mode by setting enable_realtime_transcription = True and configuring realtime_processing_pause to match your specified interval. This allows the handler to process audio chunks while speech is still ongoing rather than waiting for speech to end.
TranscriptionNotifier Event Propagation
The TranscriptionNotifier class in src/speech_to_speech/STT/transcription_notifier.py receives partial transcriptions from the STT handler and forwards them as partial_transcription events to the rest of the pipeline. These events are defined in src/speech_to_speech/pipeline/events.py and contain the intermediate text recognized from the current audio buffer.
Handling Partial Transcription Events
To consume live transcriptions in your application, register an event handler for the partial_transcription event:
from speech_to_speech import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
module_args = ModuleArguments(
enable_live_transcription=True,
live_transcription_update_interval=0.2,
)
pipeline = SpeechToSpeechPipeline(module_kwargs=module_args)
def on_partial(event):
print(f"Live: {event['partial']}")
pipeline.register_event_handler("partial_transcription", on_partial)
pipeline.run()
Disabling Live Transcription for Final-Only Output
Set enable_live_transcription=False when you only need final transcripts or when running on unsupported platforms:
module_args = ModuleArguments(
enable_live_transcription=False,
)
pipeline = SpeechToSpeechPipeline(module_kwargs=module_args)
When disabled, the VAD handler skips realtime processing (as noted in s2s_pipeline.py lines 1053-1059 for macOS compatibility), and only transcription_completed events are emitted after speech ends.
Summary
- Configure live transcription through
ModuleArgumentsinsrc/speech_to_speech/arguments_classes/module_arguments.pyusingenable_live_transcriptionandlive_transcription_update_interval - The VAD handler in
parakeet_tdt_handler.pyactivates realtime mode when the flag is enabled TranscriptionNotifieremitspartial_transcriptionevents defined insrc/speech_to_speech/pipeline/events.pyat the configured interval- Use
pipeline.register_event_handler("partial_transcription", callback)to consume streaming results - Disable the feature for batch processing or macOS multi-pipeline deployments to avoid compatibility issues
Frequently Asked Questions
What is the default update interval for live transcription?
The default value for live_transcription_update_interval is 0.5 seconds (500 milliseconds). You can reduce this to 0.1 seconds for near-instantaneous feedback or increase it to reduce processing overhead.
Can I use live transcription on macOS?
Live transcription is automatically disabled when running multiple pipelines on macOS. The S2SPipeline in src/speech_to_speech/s2s_pipeline.py (lines 1053-1059) contains a platform guard that forces enable_live_transcription=False on macOS when multiprocessing is required, as the realtime VAD mode is not supported in that configuration.
How do I consume live transcription events in my client application?
Register a callback using pipeline.register_event_handler("partial_transcription", your_function). The event payload contains a partial key with the current transcript text. For WebSocket clients, the demo server in demo/server.py streams these events automatically when the feature is enabled.
Does enabling live transcription affect latency?
Yes, enabling live transcription adds minimal latency proportional to your live_transcription_update_interval. Setting the interval too low (e.g., 0.05 seconds) may increase CPU usage without perceptible benefit, while the default 0.5 seconds provides a good balance for most applications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →