How to Implement Streaming Speech Processing with VAD, ASR, and TTS
The bojieli/ai-agent-book repository provides a production-ready reference implementation for streaming speech processing that combines Voice Activity Detection (VAD), Whisper-based Automatic Speech Recognition (ASR), and Text-to-Speech (TTS) interruption handling with real-time barge-in capabilities.
Building a responsive voice AI agent requires seamlessly integrating audio segmentation, transcription, and playback management. This guide examines the streaming speech processing implementation found in the ai-agent-book repository, which demonstrates how to process live audio streams with minimal latency while handling user interruptions gracefully.
Voice Activity Detection and Audio Segmentation
The foundation of low-latency speech processing lies in accurate voice activity detection that minimizes buffering delays.
Energy-Based VAD Implementation
In /chapter6/streaming-speech/whisper_baseline.py, the energy_vad_events function (lines 44-75) analyzes raw audio samples to detect speech boundaries. The implementation computes short-term RMS energy frames to identify three critical markers:
- Speech start: The initial frame where energy exceeds the threshold
- Acoustic endpoint: The frame where speech energy drops
- Decision frame: A later frame confirming reliable silence gap
This separation between acoustic endpoint and decision frame is crucial for streaming architectures. By keeping these distinct, the system avoids introducing a premature 600 ms latency that would otherwise occur if it waited for silence confirmation before processing.
Segment Isolation Strategy
Each detected segment is extracted and processed independently, allowing the ASR component to begin transcription as soon as a valid speech segment is identified, rather than waiting for the entire audio file to complete.
ASR Implementation with OpenAI Whisper
The repository implements a LocalWhisper wrapper that optimizes Whisper for streaming scenarios through lazy initialization and reuse.
Lazy Model Loading
The LocalWhisper class (lines 84-95 in whisper_baseline.py) loads the open-source Whisper model only once upon first invocation and reuses it for every subsequent VAD segment. This pattern prevents the expensive overhead of repeatedly loading model weights into memory, which is essential for real-time processing pipelines.
Segment Transcription Pipeline
The run_whisper_baseline helper function (lines 97-138) orchestrates the full transcription workflow:
- Loads the full audio file
- Runs VAD to obtain segment boundaries via
energy_vad_events - Writes each segment to a temporary WAV file
- Transcribes each segment using the cached
LocalWhisperinstance - Aggregates timing statistics into a
BaselineResultdataclass
from pathlib import Path
from chapter6.streaming_speech.whisper_baseline import run_whisper_baseline, LocalWhisper
audio_path = Path("sample.wav")
transcriber = LocalWhisper(model="small")
result = run_whisper_baseline(audio_path, transcriber)
print("Full transcript:", result.transcript)
print("Segment timings (s):", result.segment_asr_seconds)
Real-Time Barge-In and TTS Interruption
For duplex conversation flows, the system must detect when a user interrupts the assistant's speech and respond immediately.
DuplexInterruptionManager Design
The DuplexInterruptionManager class in /chapter6/streaming-speech/interruption_manager.py continuously monitors incoming user audio while TTS output is playing. This component maintains state about ongoing playback and tracks consecutive frames of user speech energy to distinguish between background noise and intentional interruptions.
Energy Calculation and Threshold Detection
The manager relies on the versatile calculate_energy routine (lines 1-100), which accepts raw PCM bytes, numpy arrays, or Python lists. The process_audio_chunk method uses this to compute RMS energy and compares it against a configurable VAD threshold across a configurable number of consecutive frames. Only when both conditions are met does the system register a valid barge-in event, preventing false positives from transient audio spikes.
Handling Interruptions
When the energy surpasses the threshold for the required consecutive frames, the manager invokes handle_barge_in (lines 76-133), which executes three critical operations:
- Cancels the pending TTS audio stream immediately
- Truncates the most recent assistant turn in the dialogue context
- Emits a re-planning payload that downstream agents consume to regenerate a contextual response
from chapter6.streaming_speech.interruption_manager import DuplexInterruptionManager
def on_barge_in(event):
print("Barge‑in detected:", event)
def on_replan(payload):
print("Re‑plan triggered:", payload)
mgr = DuplexInterruptionManager(vad_threshold=0.02,
consecutive_frames_required=2,
on_barge_in=on_barge_in,
on_replan=on_replan)
# Simulate TTS playback start
mgr.start_playback(initial_audio_stream=[b"...tts chunk..."])
# Simulate incoming user audio chunks (numpy arrays here)
import numpy as np
user_chunk = np.random.randn(1600) * 0.01 # low energy, no speech
print(mgr.process_audio_chunk(user_chunk))
# A loud chunk that crosses the threshold
loud_chunk = np.random.randn(1600) * 0.5
print(mgr.process_audio_chunk(loud_chunk))
Integration Workflow
A typical streaming session implemented according to the ai-agent-book architecture follows this sequence:
- Load audio input from microphone or file source
- Run
energy_vad_eventsto obtain precise speech segments with acoustic endpoints - For each segment, call
LocalWhisper.transcribeor the higher-levelrun_whisper_baseline - While the system speaks, instantiate
DuplexInterruptionManagerand start playback withstart_playback - Feed each incoming user chunk to
process_audio_chunk - On barge-in detection, the manager's callbacks stop the TTS engine, update dialogue state, and trigger a new LLM planning step
This architecture ensures continuous, low-latency speech processing where VAD isolates speech efficiently, Whisper provides high-quality ASR through model reuse, and the interruption manager maintains natural duplex conversation dynamics.
Summary
- VAD Optimization: The
energy_vad_eventsfunction inwhisper_baseline.pyseparates acoustic endpoints from decision frames to eliminate 600 ms processing delays. - Efficient ASR:
LocalWhisperimplements lazy loading to cache the Whisper model, avoiding reload overhead across multiple segments. - Real-Time Interruption:
DuplexInterruptionManageruses configurable energy thresholds and consecutive frame counting to detect barge-in without false positives. - Pipeline Integration: The combination of VAD segmentation, cached ASR, and barge-in handling creates a robust streaming speech processing pipeline suitable for production voice agents.
Frequently Asked Questions
How does the VAD implementation avoid the 600ms latency penalty?
According to the source code in whisper_baseline.py, the energy_vad_events function distinguishes between the acoustic endpoint (where speech energy drops) and the decision frame (where silence is confirmed). By processing the acoustic endpoint immediately for ASR while separately tracking the decision frame for segmentation boundaries, the system avoids waiting for silence confirmation before beginning transcription, eliminating the typical 600 ms buffer delay found in simpler VAD implementations.
What audio input formats does the energy calculation support?
The calculate_energy routine in interruption_manager.py is designed to handle multiple input types transparently. It accepts raw PCM bytes, numpy arrays, and standard Python lists, automatically detecting the input format and computing RMS energy accordingly. This flexibility allows the DuplexInterruptionManager to process audio from various sources—microphone streams, file buffers, or networked packets—without preprocessing conversion.
How is the Whisper model optimized for repeated transcription tasks?
The LocalWhisper class implements a lazy initialization pattern where the model weights are loaded into memory only upon the first call to transcribe. Subsequent VAD segments reuse this cached model instance, which is critical for streaming applications where loading a multi-gigabyte model for every speech segment would introduce unacceptable latency. This approach maintains the full accuracy of the Whisper model while achieving real-time performance.
What happens to the dialogue context when a barge-in is detected?
When DuplexInterruptionManager detects a valid interruption via handle_barge_in, it performs three synchronized operations: it cancels the active TTS audio stream to stop playback immediately, truncates the most recent assistant turn in the conversation history to remove the interrupted content, and emits a re-planning payload that signals downstream LLM agents to regenerate a response based on the new user input and updated context.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →