# How `live_transcription_update_interval` Controls Streaming Transcription Latency in Speech-to-Speech

> Learn how live_transcription_update_interval controls streaming transcription latency in speech-to-speech. Adjust this setting for faster, more responsive audio processing.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: performance
- Published: 2026-07-11

---

**`live_transcription_update_interval`** (default **0.5 seconds**) determines how frequently the pipeline flushes progressive transcription results to the user, directly governing the perceived responsiveness of streaming output.

In the `huggingface/speech-to-speech` repository, real-time transcription relies on this configurable timing parameter to balance low-latency feedback against computational overhead. By adjusting the `live_transcription_update_interval`, developers can tune whether the system prioritizes immediate partial results or batched processing for better resource efficiency.

## Configuration Source and Default Value

The parameter is declared as a dataclass field in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py):

```python
live_transcription_update_interval: float = field(
    default=0.5,
    metadata={"help": "Update interval for live transcription in seconds (default: 0.5s = 500ms)"}
)

```

This defines the default **500 ms** window that controls the emission cadence of partial transcriptions while a speaker is still talking.

## Pipeline Propagation Pathways

When the `SpeechToSpeechPipeline` initializes, the value flows into two critical subsystems that coordinate streaming latency.

### Voice Activity Detection (VAD) Integration

In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the interval is assigned to the VAD handler's realtime processing pause:

```python
vad_kw.realtime_processing_pause = module_kwargs.live_transcription_update_interval

```

According to the `huggingface/speech-to-speech` source code, the VAD runs in realtime mode and deliberately pauses for `realtime_processing_pause` seconds after each audio chunk. During these pauses, the system may emit progressive (partial) transcription results, meaning the interval directly dictates the minimum granularity of streaming updates.

### Parakeet TDT STT Handler Integration

The same value propagates to the Parakeet TDT handler in [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) as the `emission_interval`:

```python
emission_interval=self.live_transcription_update_interval,

```

This configuration ensures the STT model emits intermediate tokens at the same cadence defined by the global setting, synchronizing the acoustic model's output rate with the VAD's processing rhythm.

## Latency vs. Resource Trade-offs

The interval creates a direct trade-off between responsiveness and system load:

- **Smaller intervals** (e.g., 0.1–0.2 s) trigger more frequent pauses and emissions, reducing perceived streaming transcription latency but increasing CPU usage and UI update frequency.
- **Larger intervals** (e.g., 1.0+ s) batch progressive results longer, lowering overhead but introducing noticeable delays before partial text appears.

## State Management During Streaming

When the final (full) transcription arrives, the handler resets the live transcription state via `_reset_live_transcription_state()` and clears any pending UI line through `_clear_live_transcription_line()`. If live transcription is disabled entirely, the VAD does not reopen turns on partial results, eliminating the latency trade-off as verified by the test `test_vad_reopens_speculative_turn_when_live_transcription_disabled`.

## Practical Configuration Example

To launch the pipeline with a more responsive 200 ms update interval:

```python
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments

args = ModuleArguments(
    enable_live_transcription=True,
    live_transcription_update_interval=0.2,  # 200 ms for lower latency

)

pipeline = SpeechToSpeechPipeline(module_kwargs=args)
pipeline.run()

```

Internally, the Parakeet TDT handler consumes this value as follows:

```python
class ParakeetTDTHandler:
    def __init__(self, enable_live_transcription: bool = False,
                 live_transcription_update_interval: float = 0.5):
        self.enable_live_transcription = enable_live_transcription
        self.live_transcription_update_interval = live_transcription_update_interval

        if self.enable_live_transcription:
            # Progressive chunks emit every live_transcription_update_interval

            self._vad = VADHandler(emission_interval=self.live_transcription_update_interval)

```

## Summary

- **`live_transcription_update_interval`** defaults to **0.5 seconds** and controls the emission cadence of progressive transcriptions in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py).
- The value propagates to both the VAD handler (`realtime_processing_pause`) and the Parakeet TDT handler (`emission_interval`) via [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).
- Smaller values reduce streaming transcription latency at the cost of higher CPU usage and more frequent UI updates.
- When disabled, the system bypasses the live transcription logic entirely, removing the latency considerations and only emitting final transcriptions.

## Frequently Asked Questions

### What is the default value of `live_transcription_update_interval`?

The default value is **0.5 seconds** (500 ms), defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py). This provides a balance between responsiveness and resource consumption for most real-time applications.

### How does decreasing the interval affect CPU usage?

Reducing the interval below the default 0.5 seconds causes the VAD and STT handlers to process and emit partial results more frequently. This lowers perceived latency but increases CPU utilization because the pipeline wakes up more often to flush intermediate transcriptions.

### Where is the interval value applied in the VAD logic?

In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the value is copied to `vad_kw.realtime_processing_pause`. The VAD handler uses this to pause between processing audio chunks in realtime mode, effectively setting the throttle for how often progressive results can surface.

### Can I disable live transcription entirely to eliminate latency concerns?

Yes. Setting `enable_live_transcription=False` disables the progressive emission logic. According to the `huggingface/speech-to-speech` test suite, when disabled, the VAD does not reopen speculative turns based on partial results, bypassing the latency trade-off entirely and only emitting final transcriptions.