How to Implement VAD with DeepFilterNet Audio Enhancement in Speech-to-Speech

The Speech-to-Speech repository provides a VADHandler class that integrates Silero VAD for voice activity detection and optionally applies DeepFilterNet audio enhancement by setting audio_enhancement=True during setup, which triggers noise reduction via the _apply_audio_enhancement method before emitting VADAudio events.

The huggingface/speech-to-speech repository offers a production-ready pipeline for real-time voice processing that combines Voice Activity Detection (VAD) with DeepFilterNet audio enhancement. This guide explains how to implement VAD with DeepFilterNet audio enhancement using the library's modular handler architecture, referencing actual source file paths and configuration patterns found in src/speech_to_speech/VAD/vad_handler.py and related modules.

Architecture Overview

The VAD stack consists of two primary components that coordinate to detect speech and optionally enhance audio quality before downstream processing.

VADIterator: Core Speech Detection Engine

The VADIterator class in src/speech_to_speech/VAD/vad_iterator.py serves as the low-level VAD engine. It consumes raw audio chunks, maintains a pre-speech buffer, and determines when speech segments start and end using the Silero VAD model loaded via torch.hub.load.

Key characteristics include:

  • Trigger management: Maintains a self.triggered flag that indicates active utterances
  • Buffer handling: Manages silence padding and speech-pad configurations
  • Output format: Returns a list of torch.Tensor chunks when speech ends, otherwise None

VADHandler: Orchestration and Enhancement Layer

The VADHandler class in src/speech_to_speech/VAD/vad_handler.py orchestrates the iterator, manages turn metadata, and optionally runs DeepFilterNet on finalized audio segments. This handler is instantiated per session and receives a should_listen event to control processing flow.

Critical implementation details:

  • Optional dependency: DeepFilterNet import is guarded with a try/except block that sets a HAS_DF flag
  • Graceful degradation: If the import fails, the library logs a warning and continues without enhancement capabilities
  • Event emission: Produces VADAudio events (defined in src/speech_to_speech/pipeline/messages.py) containing either raw or enhanced NumPy arrays

Enabling DeepFilterNet Enhancement

To implement audio enhancement, you must first install the optional dependency and then configure the handler during initialization.

Installation Requirements

Install DeepFilterNet with PyTorch support:

pip install "deep-filter-net[torch]"

Handler Configuration

Activate enhancement by passing audio_enhancement=True to the setup method:

from speech_to_speech.VAD.vad_handler import VADHandler

handler = VADHandler()
handler.setup(
    should_listen=listen_event,
    thresh=0.6,
    sample_rate=16000,
    min_silence_ms=64,
    min_speech_ms=384,
    audio_enhancement=True,  # Enable DeepFilterNet

)

The setup method (lines 44-51 in vad_handler.py) attempts to import df.enhance and initialize the model via init_df(). If the import fails, HAS_DF becomes False and the audio_enhancement flag is automatically cleared to prevent runtime errors.

Runtime Execution Flow

When processing audio streams, the handler applies enhancement at the final speech boundary before emitting events.

The Enhancement Pipeline

When audio_enhancement=True and a speech segment completes, the handler invokes _apply_audio_enhancement (lines 110-133 in vad_handler.py). This method performs the following operations:

  1. Resampling: Converts audio to DeepFilterNet's internal sample rate (self.df_state.sr()) if necessary using torchaudio.functional.resample
  2. Enhancement: Processes through enhance(self.enhanced_model, self.df_state, audio_tensor)
  3. Restoration: Resamples back to the original sample rate
  4. Output: Returns the enhanced NumPy array that replaces the original in the VADAudio event
import torchaudio

if self.sample_rate != self.df_state.sr():
    audio_float32 = torchaudio.functional.resample(
        torch.from_numpy(array), orig_freq=self.sample_rate, new_freq=self.df_state.sr()
    )
    enhanced = enhance(self.enhanced_model, self.df_state, audio_float32.unsqueeze(0))
    enhanced = torchaudio.functional.resample(
        enhanced, orig_freq=self.df_state.sr(), new_freq=self.sample_rate
    )
else:
    enhanced = enhance(self.enhanced_model, self.df_state, torch.from_numpy(array))
return enhanced.numpy().squeeze()

Integration Patterns

You can integrate this functionality through the high-level pipeline or directly in custom implementations.

Using the S2SPipeline

When using the built-in S2SPipeline, pass the enhancement flag via constructor or CLI:

python -m speech_to_speech --audio-enhancement

Custom Handler Implementation

For custom flows, instantiate VADHandler directly and process byte chunks:

from speech_to_speech.VAD.vad_handler import VADHandler
from threading import Event

listen_event = Event()
listen_event.set()

handler = VADHandler()
handler.setup(
    should_listen=listen_event,
    audio_enhancement=True,
)

# Process simulated audio stream

for chunk in audio_byte_chunks:
    for output in handler.process(chunk):
        if isinstance(output, VADAudio):
            # output.audio contains enhanced NumPy array

            process_audio(output.audio)

Summary

  • VADIterator (src/speech_to_speech/VAD/vad_iterator.py) handles low-level Silero VAD inference and buffer management using torch.hub.load
  • VADHandler (src/speech_to_speech/VAD/vad_handler.py) coordinates turn-taking and optionally applies DeepFilterNet enhancement via _apply_audio_enhancement
  • DeepFilterNet integration requires pip install "deep-filter-net[torch]" and setting audio_enhancement=True in setup()
  • Graceful fallback occurs when DeepFilterNet is unavailable—the handler logs a warning and continues without enhancement
  • Resampling is handled automatically when input sample rates differ from DeepFilterNet's internal rate
  • Enhanced output is delivered as a VADAudio event containing a NumPy array ready for downstream STT or LLM processing

Frequently Asked Questions

What happens if DeepFilterNet is not installed but I set audio_enhancement=True?

The handler catches the import error in setup(), sets HAS_DF to False, logs a warning, and automatically disables enhancement for that session. The VAD functionality continues working normally without noise reduction.

How does the VAD handler manage different sample rates?

The _apply_audio_enhancement method checks if self.sample_rate matches self.df_state.sr(). If they differ, it uses torchaudio.functional.resample to convert the audio to DeepFilterNet's required rate before processing, then resamples back to the original rate before returning.

Can I use VAD without audio enhancement?

Yes. Simply omit the audio_enhancement parameter or set it to False when calling handler.setup(). The VADHandler works independently of DeepFilterNet and will emit VADAudio events containing the original audio buffers without modification.

Where is the enhanced audio stored in the output structure?

The enhanced audio is stored in the audio attribute of the VADAudio object emitted by the handler. This object is defined in src/speech_to_speech/pipeline/messages.py and contains the NumPy array along with metadata such as the mode flag ("final" or "progressive") and turn identifiers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →