# How to Implement Streaming Speech Processing with VAD, ASR, and TTS

> Implement streaming speech processing with VAD, ASR, and TTS. Explore a production-ready reference implementation for real-time barge-in capabilities in the bojieli ai-agent-book repository.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: how-to-guide
- Published: 2026-08-17

---

**The bojieli/ai-agent-book repository provides a production-ready reference implementation for streaming speech processing that combines Voice Activity Detection (VAD), Whisper-based Automatic Speech Recognition (ASR), and Text-to-Speech (TTS) interruption handling with real-time barge-in capabilities.**

Building a responsive voice AI agent requires seamlessly integrating audio segmentation, transcription, and playback management. This guide examines the streaming speech processing implementation found in the **ai-agent-book** repository, which demonstrates how to process live audio streams with minimal latency while handling user interruptions gracefully.

## Voice Activity Detection and Audio Segmentation

The foundation of low-latency speech processing lies in accurate voice activity detection that minimizes buffering delays.

### Energy-Based VAD Implementation

In [`/chapter6/streaming-speech/whisper_baseline.py`](https://github.com/bojieli/ai-agent-book/blob/main//chapter6/streaming-speech/whisper_baseline.py), the `energy_vad_events` function (lines 44-75) analyzes raw audio samples to detect speech boundaries. The implementation computes short-term **RMS energy frames** to identify three critical markers:

- **Speech start**: The initial frame where energy exceeds the threshold
- **Acoustic endpoint**: The frame where speech energy drops
- **Decision frame**: A later frame confirming reliable silence gap

This separation between acoustic endpoint and decision frame is crucial for streaming architectures. By keeping these distinct, the system avoids introducing a premature **600 ms latency** that would otherwise occur if it waited for silence confirmation before processing.

### Segment Isolation Strategy

Each detected segment is extracted and processed independently, allowing the ASR component to begin transcription as soon as a valid speech segment is identified, rather than waiting for the entire audio file to complete.

## ASR Implementation with OpenAI Whisper

The repository implements a **LocalWhisper** wrapper that optimizes Whisper for streaming scenarios through lazy initialization and reuse.

### Lazy Model Loading

The `LocalWhisper` class (lines 84-95 in [`whisper_baseline.py`](https://github.com/bojieli/ai-agent-book/blob/main/whisper_baseline.py)) loads the open-source Whisper model only once upon first invocation and reuses it for every subsequent VAD segment. This pattern prevents the expensive overhead of repeatedly loading model weights into memory, which is essential for real-time processing pipelines.

### Segment Transcription Pipeline

The `run_whisper_baseline` helper function (lines 97-138) orchestrates the full transcription workflow:

1. Loads the full audio file
2. Runs VAD to obtain segment boundaries via `energy_vad_events`
3. Writes each segment to a temporary WAV file
4. Transcribes each segment using the cached `LocalWhisper` instance
5. Aggregates timing statistics into a `BaselineResult` dataclass

```python
from pathlib import Path
from chapter6.streaming_speech.whisper_baseline import run_whisper_baseline, LocalWhisper

audio_path = Path("sample.wav")
transcriber = LocalWhisper(model="small")
result = run_whisper_baseline(audio_path, transcriber)

print("Full transcript:", result.transcript)
print("Segment timings (s):", result.segment_asr_seconds)

```

## Real-Time Barge-In and TTS Interruption

For duplex conversation flows, the system must detect when a user interrupts the assistant's speech and respond immediately.

### DuplexInterruptionManager Design

The `DuplexInterruptionManager` class in [`/chapter6/streaming-speech/interruption_manager.py`](https://github.com/bojieli/ai-agent-book/blob/main//chapter6/streaming-speech/interruption_manager.py) continuously monitors incoming user audio while TTS output is playing. This component maintains state about ongoing playback and tracks consecutive frames of user speech energy to distinguish between background noise and intentional interruptions.

### Energy Calculation and Threshold Detection

The manager relies on the versatile `calculate_energy` routine (lines 1-100), which accepts raw PCM bytes, numpy arrays, or Python lists. The `process_audio_chunk` method uses this to compute RMS energy and compares it against a **configurable VAD threshold** across a **configurable number of consecutive frames**. Only when both conditions are met does the system register a valid barge-in event, preventing false positives from transient audio spikes.

### Handling Interruptions

When the energy surpasses the threshold for the required consecutive frames, the manager invokes `handle_barge_in` (lines 76-133), which executes three critical operations:

- Cancels the pending TTS audio stream immediately
- Truncates the most recent assistant turn in the dialogue context
- Emits a **re-planning payload** that downstream agents consume to regenerate a contextual response

```python
from chapter6.streaming_speech.interruption_manager import DuplexInterruptionManager

def on_barge_in(event):
    print("Barge‑in detected:", event)

def on_replan(payload):
    print("Re‑plan triggered:", payload)

mgr = DuplexInterruptionManager(vad_threshold=0.02,
                                consecutive_frames_required=2,
                                on_barge_in=on_barge_in,
                                on_replan=on_replan)

# Simulate TTS playback start

mgr.start_playback(initial_audio_stream=[b"...tts chunk..."])

# Simulate incoming user audio chunks (numpy arrays here)

import numpy as np
user_chunk = np.random.randn(1600) * 0.01  # low energy, no speech

print(mgr.process_audio_chunk(user_chunk))

# A loud chunk that crosses the threshold

loud_chunk = np.random.randn(1600) * 0.5
print(mgr.process_audio_chunk(loud_chunk))

```

## Integration Workflow

A typical streaming session implemented according to the ai-agent-book architecture follows this sequence:

1. Load audio input from microphone or file source
2. Run `energy_vad_events` to obtain precise speech segments with acoustic endpoints
3. For each segment, call `LocalWhisper.transcribe` or the higher-level `run_whisper_baseline`
4. While the system speaks, instantiate `DuplexInterruptionManager` and start playback with `start_playback`
5. Feed each incoming user chunk to `process_audio_chunk`
6. On barge-in detection, the manager's callbacks stop the TTS engine, update dialogue state, and trigger a new LLM planning step

This architecture ensures **continuous, low-latency speech processing** where VAD isolates speech efficiently, Whisper provides high-quality ASR through model reuse, and the interruption manager maintains natural duplex conversation dynamics.

## Summary

- **VAD Optimization**: The `energy_vad_events` function in [`whisper_baseline.py`](https://github.com/bojieli/ai-agent-book/blob/main/whisper_baseline.py) separates acoustic endpoints from decision frames to eliminate 600 ms processing delays.
- **Efficient ASR**: `LocalWhisper` implements lazy loading to cache the Whisper model, avoiding reload overhead across multiple segments.
- **Real-Time Interruption**: `DuplexInterruptionManager` uses configurable energy thresholds and consecutive frame counting to detect barge-in without false positives.
- **Pipeline Integration**: The combination of VAD segmentation, cached ASR, and barge-in handling creates a robust streaming speech processing pipeline suitable for production voice agents.

## Frequently Asked Questions

### How does the VAD implementation avoid the 600ms latency penalty?

According to the source code in [`whisper_baseline.py`](https://github.com/bojieli/ai-agent-book/blob/main/whisper_baseline.py), the `energy_vad_events` function distinguishes between the **acoustic endpoint** (where speech energy drops) and the **decision frame** (where silence is confirmed). By processing the acoustic endpoint immediately for ASR while separately tracking the decision frame for segmentation boundaries, the system avoids waiting for silence confirmation before beginning transcription, eliminating the typical 600 ms buffer delay found in simpler VAD implementations.

### What audio input formats does the energy calculation support?

The `calculate_energy` routine in [`interruption_manager.py`](https://github.com/bojieli/ai-agent-book/blob/main/interruption_manager.py) is designed to handle multiple input types transparently. It accepts raw PCM bytes, numpy arrays, and standard Python lists, automatically detecting the input format and computing RMS energy accordingly. This flexibility allows the `DuplexInterruptionManager` to process audio from various sources—microphone streams, file buffers, or networked packets—without preprocessing conversion.

### How is the Whisper model optimized for repeated transcription tasks?

The `LocalWhisper` class implements a **lazy initialization pattern** where the model weights are loaded into memory only upon the first call to `transcribe`. Subsequent VAD segments reuse this cached model instance, which is critical for streaming applications where loading a multi-gigabyte model for every speech segment would introduce unacceptable latency. This approach maintains the full accuracy of the Whisper model while achieving real-time performance.

### What happens to the dialogue context when a barge-in is detected?

When `DuplexInterruptionManager` detects a valid interruption via `handle_barge_in`, it performs three synchronized operations: it cancels the active TTS audio stream to stop playback immediately, truncates the most recent assistant turn in the conversation history to remove the interrupted content, and emits a re-planning payload that signals downstream LLM agents to regenerate a response based on the new user input and updated context.