How `live_transcription_update_interval` Controls Streaming Transcription Latency in Speech-to-Speech
live_transcription_update_interval (default 0.5 seconds) determines how frequently the pipeline flushes progressive transcription results to the user, directly governing the perceived responsiveness of streaming output.
In the huggingface/speech-to-speech repository, real-time transcription relies on this configurable timing parameter to balance low-latency feedback against computational overhead. By adjusting the live_transcription_update_interval, developers can tune whether the system prioritizes immediate partial results or batched processing for better resource efficiency.
Configuration Source and Default Value
The parameter is declared as a dataclass field in src/speech_to_speech/arguments_classes/module_arguments.py:
live_transcription_update_interval: float = field(
default=0.5,
metadata={"help": "Update interval for live transcription in seconds (default: 0.5s = 500ms)"}
)
This defines the default 500 ms window that controls the emission cadence of partial transcriptions while a speaker is still talking.
Pipeline Propagation Pathways
When the SpeechToSpeechPipeline initializes, the value flows into two critical subsystems that coordinate streaming latency.
Voice Activity Detection (VAD) Integration
In src/speech_to_speech/s2s_pipeline.py, the interval is assigned to the VAD handler's realtime processing pause:
vad_kw.realtime_processing_pause = module_kwargs.live_transcription_update_interval
According to the huggingface/speech-to-speech source code, the VAD runs in realtime mode and deliberately pauses for realtime_processing_pause seconds after each audio chunk. During these pauses, the system may emit progressive (partial) transcription results, meaning the interval directly dictates the minimum granularity of streaming updates.
Parakeet TDT STT Handler Integration
The same value propagates to the Parakeet TDT handler in src/speech_to_speech/STT/parakeet_tdt_handler.py as the emission_interval:
emission_interval=self.live_transcription_update_interval,
This configuration ensures the STT model emits intermediate tokens at the same cadence defined by the global setting, synchronizing the acoustic model's output rate with the VAD's processing rhythm.
Latency vs. Resource Trade-offs
The interval creates a direct trade-off between responsiveness and system load:
- Smaller intervals (e.g., 0.1–0.2 s) trigger more frequent pauses and emissions, reducing perceived streaming transcription latency but increasing CPU usage and UI update frequency.
- Larger intervals (e.g., 1.0+ s) batch progressive results longer, lowering overhead but introducing noticeable delays before partial text appears.
State Management During Streaming
When the final (full) transcription arrives, the handler resets the live transcription state via _reset_live_transcription_state() and clears any pending UI line through _clear_live_transcription_line(). If live transcription is disabled entirely, the VAD does not reopen turns on partial results, eliminating the latency trade-off as verified by the test test_vad_reopens_speculative_turn_when_live_transcription_disabled.
Practical Configuration Example
To launch the pipeline with a more responsive 200 ms update interval:
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
args = ModuleArguments(
enable_live_transcription=True,
live_transcription_update_interval=0.2, # 200 ms for lower latency
)
pipeline = SpeechToSpeechPipeline(module_kwargs=args)
pipeline.run()
Internally, the Parakeet TDT handler consumes this value as follows:
class ParakeetTDTHandler:
def __init__(self, enable_live_transcription: bool = False,
live_transcription_update_interval: float = 0.5):
self.enable_live_transcription = enable_live_transcription
self.live_transcription_update_interval = live_transcription_update_interval
if self.enable_live_transcription:
# Progressive chunks emit every live_transcription_update_interval
self._vad = VADHandler(emission_interval=self.live_transcription_update_interval)
Summary
live_transcription_update_intervaldefaults to 0.5 seconds and controls the emission cadence of progressive transcriptions insrc/speech_to_speech/arguments_classes/module_arguments.py.- The value propagates to both the VAD handler (
realtime_processing_pause) and the Parakeet TDT handler (emission_interval) viasrc/speech_to_speech/s2s_pipeline.py. - Smaller values reduce streaming transcription latency at the cost of higher CPU usage and more frequent UI updates.
- When disabled, the system bypasses the live transcription logic entirely, removing the latency considerations and only emitting final transcriptions.
Frequently Asked Questions
What is the default value of live_transcription_update_interval?
The default value is 0.5 seconds (500 ms), defined in src/speech_to_speech/arguments_classes/module_arguments.py. This provides a balance between responsiveness and resource consumption for most real-time applications.
How does decreasing the interval affect CPU usage?
Reducing the interval below the default 0.5 seconds causes the VAD and STT handlers to process and emit partial results more frequently. This lowers perceived latency but increases CPU utilization because the pipeline wakes up more often to flush intermediate transcriptions.
Where is the interval value applied in the VAD logic?
In src/speech_to_speech/s2s_pipeline.py, the value is copied to vad_kw.realtime_processing_pause. The VAD handler uses this to pause between processing audio chunks in realtime mode, effectively setting the throttle for how often progressive results can surface.
Can I disable live transcription entirely to eliminate latency concerns?
Yes. Setting enable_live_transcription=False disables the progressive emission logic. According to the huggingface/speech-to-speech test suite, when disabled, the VAD does not reopen speculative turns based on partial results, bypassing the latency trade-off entirely and only emitting final transcriptions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →