# How Live Transcription Works During Active Speech Detection in Hugging Face Speech-to-Speech

> Discover how live transcription leverages active speech detection in Hugging Face Speech-to-Speech. Learn how VAD buffers audio and processes segments for accurate real-time transcription.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-30

---

**Live transcription during active speech detection in the huggingface/speech-to-speech repository works by buffering audio chunks only when the Voice Activity Detection (VAD) component detects speech above a probability threshold, then emitting complete segments to the STT pipeline once the configured silence duration (`speech_continuation_ms`) is exceeded.**

The huggingface/speech-to-speech library implements a real-time audio processing pipeline that coordinates Voice Activity Detection (VAD) with Automatic Speech Recognition (ASR) to enable efficient live transcription. Understanding the interaction between the VAD subsystem and the speech-to-text pipeline reveals how the system achieves near real-time responsiveness while maintaining accuracy.

## Audio Ingestion and Stream Initialization

The pipeline begins with the **LocalAudioStreamer** in [`src/speech_to_speech/connections/local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/local_audio_streamer.py), which captures raw PCM audio from the microphone. Instead of passing raw bytes downstream, the streamer wraps each chunk in a **VADAudio** message object. This structure carries a mode flag indicating whether the audio represents ongoing speech (`progressive` mode) or an explicit end-of-turn signal (`final` mode), allowing the pipeline to handle audio as discrete semantic events rather than continuous streams.

## Voice Activity Detection Implementation

The core detection logic resides in **VADHandler** ([`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py)), a subclass of `BaseHandler` that receives the stream of `VADAudio` objects. Internally, the handler creates a **VADIterator** (defined in [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py)) that analyzes each audio chunk using a lightweight PyTorch model to determine speech probability.

### Configuration and Thresholds

Detection sensitivity is governed by **VADHandlerArguments** in [`src/speech_to_speech/arguments_classes/vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/vad_arguments.py):

```python
class VADHandlerArguments:
    threshold: float = 0.5          # probability above which speech is considered active

    speech_continuation_ms: int = 200  # silence window that still belongs to the same utterance

```

### Soft-End Buffering Strategy

The **VADIterator** yields a boolean `speech` flag for each audio chunk. When the probability drops below `threshold`, the iterator implements a *soft-end* grace period: if the elapsed silence remains shorter than `speech_continuation_ms`, the handler continues buffering audio as though speech were still active. This prevents premature segmentation during natural conversational pauses. Only when silence persists beyond the configured window does the handler finalize the segment.

## Turn Lifecycle and Transcription Triggering

Upon determining that a speech segment is complete—either through silence timeout or receiving a `final` mode signal—the **VADHandler** emits a **VADOutItem** (defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py)). This data structure contains:

- The concatenated audio bytes for the complete utterance
- A unique **turn_id** (UUID) identifying the conversation turn
- An incrementing revision number to track segment iterations

The `VADOutItem` travels downstream to the STT handler (such as `WhisperSTTHandler` or `FasterWhisperSTTHandler`), which converts the buffered audio into text. The pipeline distinguishes between **PartialTranscription** events—emitted during active recognition to provide real-time feedback—and final **Transcription** objects sent when inference completes.

## End-to-End Pipeline Configuration

The following example demonstrates initializing the complete pipeline with VAD, Whisper-based STT, a language model, and TTS capabilities:

```python
from speech_to_speech.s2s_pipeline import build_s2s_pipeline
from speech_to_speech.arguments_classes.vad_arguments import VADHandlerArguments
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTArguments
from speech_to_speech.arguments_classes.chat_completions_language_model_arguments import ChatCompletionsLanguageModelArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments

# Initialize with default VAD settings

pipeline = build_s2s_pipeline(
    vad_handler_kwargs=VADHandlerArguments(),
    stt_handler_kwargs=WhisperSTTArguments(),
    lm_handler_kwargs=ChatCompletionsLanguageModelArguments(),
    tts_handler_kwargs=Qwen3TTSArguments(),
)

# Process live audio stream

for event in pipeline.run():
    if isinstance(event, PartialTranscription):
        print(f"⏳ Partial: {event.text}")
    elif isinstance(event, Transcription):
        print(f"✅ Final: {event.text}")

```

This configuration uses the default VAD parameters (`threshold=0.5`, `speech_continuation_ms=200` milliseconds). Adjust these values to tune detection sensitivity for specific acoustic environments.

## Testing and Validation

The repository validates this behavior through targeted test suites. [`tests/test_vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/test_vad_iterator.py) confirms that the iterator correctly respects silence thresholds and emits buffered audio only after true silence gaps. [`tests/test_speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/test_speculative_turns.py) constructs synthetic `VADAudio` streams to verify that `VADHandler` generates `VADOutItem` instances with accurate turn IDs and audio boundaries, ensuring reliable live transcription under various edge cases.

## Summary

- **Live transcription during active speech detection** relies on **VADHandler** in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) to buffer audio chunks only when the **VADIterator** detects speech probability exceeding the configured `threshold`.
- The soft-end mechanism uses `speech_continuation_ms` to maintain buffering during brief silences, preventing fragmentation of natural utterances.
- Completed segments are packaged as **VADOutItem** objects with unique turn identifiers and revision numbers before transmission to STT handlers.
- The system emits **PartialTranscription** events for real-time feedback and **Transcription** objects for final results, enabling responsive user interfaces while the pipeline continues capturing audio.
- Fine-tuning occurs through **VADHandlerArguments** in [`src/speech_to_speech/arguments_classes/vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/vad_arguments.py), allowing precise adjustment of detection sensitivity and silence tolerance.

## Frequently Asked Questions

### How does the VADHandler determine when to stop buffering audio?

The **VADHandler** stops buffering when the **VADIterator** reports inactive speech status for a duration exceeding `speech_continuation_ms` (default 200ms), or when it receives a `VADAudio` message with mode set to `"final"`. At this point, it concatenates all buffered frames into a single audio segment wrapped in a `VADOutItem` and resets the buffer for the next turn.

### What is the difference between PartialTranscription and Transcription events?

**PartialTranscription** objects represent interim recognition results emitted while the STT engine is still processing the audio segment, allowing applications to display live text updates. **Transcription** objects indicate final, confirmed recognition results for a complete utterance, typically sent once the VAD has confirmed the end of speech and the ASR model has finished inference.

### Can I adjust how sensitive the voice activity detection is?

Yes, sensitivity is controlled through the **VADHandlerArguments** class in [`src/speech_to_speech/arguments_classes/vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/vad_arguments.py). Lower the `threshold` value (default 0.5) to detect quieter speech, or increase `speech_continuation_ms` to allow longer pauses without triggering a new segment. These parameters can be passed as kwargs when calling `build_s2s_pipeline()`.

### Where is the turn identification logic implemented?

Turn identification is managed within [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py), where the **VADHandler** generates UUID-based **turn_id** values and incrementing revision numbers for each speech segment. This metadata travels with the `VADOutItem` through the pipeline, allowing downstream components to correlate partial and final transcriptions with specific conversation turns.