How Live Transcription Works During Active Speech Detection in Hugging Face Speech-to-Speech

Live transcription during active speech detection in the huggingface/speech-to-speech repository works by buffering audio chunks only when the Voice Activity Detection (VAD) component detects speech above a probability threshold, then emitting complete segments to the STT pipeline once the configured silence duration (speech_continuation_ms) is exceeded.

The huggingface/speech-to-speech library implements a real-time audio processing pipeline that coordinates Voice Activity Detection (VAD) with Automatic Speech Recognition (ASR) to enable efficient live transcription. Understanding the interaction between the VAD subsystem and the speech-to-text pipeline reveals how the system achieves near real-time responsiveness while maintaining accuracy.

Audio Ingestion and Stream Initialization

The pipeline begins with the LocalAudioStreamer in src/speech_to_speech/connections/local_audio_streamer.py, which captures raw PCM audio from the microphone. Instead of passing raw bytes downstream, the streamer wraps each chunk in a VADAudio message object. This structure carries a mode flag indicating whether the audio represents ongoing speech (progressive mode) or an explicit end-of-turn signal (final mode), allowing the pipeline to handle audio as discrete semantic events rather than continuous streams.

Voice Activity Detection Implementation

The core detection logic resides in VADHandler (src/speech_to_speech/VAD/vad_handler.py), a subclass of BaseHandler that receives the stream of VADAudio objects. Internally, the handler creates a VADIterator (defined in src/speech_to_speech/VAD/vad_iterator.py) that analyzes each audio chunk using a lightweight PyTorch model to determine speech probability.

Configuration and Thresholds

Detection sensitivity is governed by VADHandlerArguments in src/speech_to_speech/arguments_classes/vad_arguments.py:

class VADHandlerArguments:
    threshold: float = 0.5          # probability above which speech is considered active

    speech_continuation_ms: int = 200  # silence window that still belongs to the same utterance

Soft-End Buffering Strategy

The VADIterator yields a boolean speech flag for each audio chunk. When the probability drops below threshold, the iterator implements a soft-end grace period: if the elapsed silence remains shorter than speech_continuation_ms, the handler continues buffering audio as though speech were still active. This prevents premature segmentation during natural conversational pauses. Only when silence persists beyond the configured window does the handler finalize the segment.

Turn Lifecycle and Transcription Triggering

Upon determining that a speech segment is complete—either through silence timeout or receiving a final mode signal—the VADHandler emits a VADOutItem (defined in src/speech_to_speech/pipeline/messages.py). This data structure contains:

  • The concatenated audio bytes for the complete utterance
  • A unique turn_id (UUID) identifying the conversation turn
  • An incrementing revision number to track segment iterations

The VADOutItem travels downstream to the STT handler (such as WhisperSTTHandler or FasterWhisperSTTHandler), which converts the buffered audio into text. The pipeline distinguishes between PartialTranscription events—emitted during active recognition to provide real-time feedback—and final Transcription objects sent when inference completes.

End-to-End Pipeline Configuration

The following example demonstrates initializing the complete pipeline with VAD, Whisper-based STT, a language model, and TTS capabilities:

from speech_to_speech.s2s_pipeline import build_s2s_pipeline
from speech_to_speech.arguments_classes.vad_arguments import VADHandlerArguments
from speech_to_speech.arguments_classes.whisper_stt_arguments import WhisperSTTArguments
from speech_to_speech.arguments_classes.chat_completions_language_model_arguments import ChatCompletionsLanguageModelArguments
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments

# Initialize with default VAD settings

pipeline = build_s2s_pipeline(
    vad_handler_kwargs=VADHandlerArguments(),
    stt_handler_kwargs=WhisperSTTArguments(),
    lm_handler_kwargs=ChatCompletionsLanguageModelArguments(),
    tts_handler_kwargs=Qwen3TTSArguments(),
)

# Process live audio stream

for event in pipeline.run():
    if isinstance(event, PartialTranscription):
        print(f"⏳ Partial: {event.text}")
    elif isinstance(event, Transcription):
        print(f"✅ Final: {event.text}")

This configuration uses the default VAD parameters (threshold=0.5, speech_continuation_ms=200 milliseconds). Adjust these values to tune detection sensitivity for specific acoustic environments.

Testing and Validation

The repository validates this behavior through targeted test suites. tests/test_vad_iterator.py confirms that the iterator correctly respects silence thresholds and emits buffered audio only after true silence gaps. tests/test_speculative_turns.py constructs synthetic VADAudio streams to verify that VADHandler generates VADOutItem instances with accurate turn IDs and audio boundaries, ensuring reliable live transcription under various edge cases.

Summary

  • Live transcription during active speech detection relies on VADHandler in src/speech_to_speech/VAD/vad_handler.py to buffer audio chunks only when the VADIterator detects speech probability exceeding the configured threshold.
  • The soft-end mechanism uses speech_continuation_ms to maintain buffering during brief silences, preventing fragmentation of natural utterances.
  • Completed segments are packaged as VADOutItem objects with unique turn identifiers and revision numbers before transmission to STT handlers.
  • The system emits PartialTranscription events for real-time feedback and Transcription objects for final results, enabling responsive user interfaces while the pipeline continues capturing audio.
  • Fine-tuning occurs through VADHandlerArguments in src/speech_to_speech/arguments_classes/vad_arguments.py, allowing precise adjustment of detection sensitivity and silence tolerance.

Frequently Asked Questions

How does the VADHandler determine when to stop buffering audio?

The VADHandler stops buffering when the VADIterator reports inactive speech status for a duration exceeding speech_continuation_ms (default 200ms), or when it receives a VADAudio message with mode set to "final". At this point, it concatenates all buffered frames into a single audio segment wrapped in a VADOutItem and resets the buffer for the next turn.

What is the difference between PartialTranscription and Transcription events?

PartialTranscription objects represent interim recognition results emitted while the STT engine is still processing the audio segment, allowing applications to display live text updates. Transcription objects indicate final, confirmed recognition results for a complete utterance, typically sent once the VAD has confirmed the end of speech and the ASR model has finished inference.

Can I adjust how sensitive the voice activity detection is?

Yes, sensitivity is controlled through the VADHandlerArguments class in src/speech_to_speech/arguments_classes/vad_arguments.py. Lower the threshold value (default 0.5) to detect quieter speech, or increase speech_continuation_ms to allow longer pauses without triggering a new segment. These parameters can be passed as kwargs when calling build_s2s_pipeline().

Where is the turn identification logic implemented?

Turn identification is managed within src/speech_to_speech/VAD/vad_handler.py, where the VADHandler generates UUID-based turn_id values and incrementing revision numbers for each speech segment. This metadata travels with the VADOutItem through the pipeline, allowing downstream components to correlate partial and final transcriptions with specific conversation turns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →