How to Implement VAD with DeepFilterNet Audio Enhancement in Speech-to-Speech
The Speech-to-Speech repository provides a VADHandler class that integrates Silero VAD for voice activity detection and optionally applies DeepFilterNet audio enhancement by setting audio_enhancement=True during setup, which triggers noise reduction via the _apply_audio_enhancement method before emitting VADAudio events.
The huggingface/speech-to-speech repository offers a production-ready pipeline for real-time voice processing that combines Voice Activity Detection (VAD) with DeepFilterNet audio enhancement. This guide explains how to implement VAD with DeepFilterNet audio enhancement using the library's modular handler architecture, referencing actual source file paths and configuration patterns found in src/speech_to_speech/VAD/vad_handler.py and related modules.
Architecture Overview
The VAD stack consists of two primary components that coordinate to detect speech and optionally enhance audio quality before downstream processing.
VADIterator: Core Speech Detection Engine
The VADIterator class in src/speech_to_speech/VAD/vad_iterator.py serves as the low-level VAD engine. It consumes raw audio chunks, maintains a pre-speech buffer, and determines when speech segments start and end using the Silero VAD model loaded via torch.hub.load.
Key characteristics include:
- Trigger management: Maintains a
self.triggeredflag that indicates active utterances - Buffer handling: Manages silence padding and speech-pad configurations
- Output format: Returns a list of
torch.Tensorchunks when speech ends, otherwiseNone
VADHandler: Orchestration and Enhancement Layer
The VADHandler class in src/speech_to_speech/VAD/vad_handler.py orchestrates the iterator, manages turn metadata, and optionally runs DeepFilterNet on finalized audio segments. This handler is instantiated per session and receives a should_listen event to control processing flow.
Critical implementation details:
- Optional dependency: DeepFilterNet import is guarded with a try/except block that sets a
HAS_DFflag - Graceful degradation: If the import fails, the library logs a warning and continues without enhancement capabilities
- Event emission: Produces
VADAudioevents (defined insrc/speech_to_speech/pipeline/messages.py) containing either raw or enhanced NumPy arrays
Enabling DeepFilterNet Enhancement
To implement audio enhancement, you must first install the optional dependency and then configure the handler during initialization.
Installation Requirements
Install DeepFilterNet with PyTorch support:
pip install "deep-filter-net[torch]"
Handler Configuration
Activate enhancement by passing audio_enhancement=True to the setup method:
from speech_to_speech.VAD.vad_handler import VADHandler
handler = VADHandler()
handler.setup(
should_listen=listen_event,
thresh=0.6,
sample_rate=16000,
min_silence_ms=64,
min_speech_ms=384,
audio_enhancement=True, # Enable DeepFilterNet
)
The setup method (lines 44-51 in vad_handler.py) attempts to import df.enhance and initialize the model via init_df(). If the import fails, HAS_DF becomes False and the audio_enhancement flag is automatically cleared to prevent runtime errors.
Runtime Execution Flow
When processing audio streams, the handler applies enhancement at the final speech boundary before emitting events.
The Enhancement Pipeline
When audio_enhancement=True and a speech segment completes, the handler invokes _apply_audio_enhancement (lines 110-133 in vad_handler.py). This method performs the following operations:
- Resampling: Converts audio to DeepFilterNet's internal sample rate (
self.df_state.sr()) if necessary usingtorchaudio.functional.resample - Enhancement: Processes through
enhance(self.enhanced_model, self.df_state, audio_tensor) - Restoration: Resamples back to the original sample rate
- Output: Returns the enhanced NumPy array that replaces the original in the
VADAudioevent
import torchaudio
if self.sample_rate != self.df_state.sr():
audio_float32 = torchaudio.functional.resample(
torch.from_numpy(array), orig_freq=self.sample_rate, new_freq=self.df_state.sr()
)
enhanced = enhance(self.enhanced_model, self.df_state, audio_float32.unsqueeze(0))
enhanced = torchaudio.functional.resample(
enhanced, orig_freq=self.df_state.sr(), new_freq=self.sample_rate
)
else:
enhanced = enhance(self.enhanced_model, self.df_state, torch.from_numpy(array))
return enhanced.numpy().squeeze()
Integration Patterns
You can integrate this functionality through the high-level pipeline or directly in custom implementations.
Using the S2SPipeline
When using the built-in S2SPipeline, pass the enhancement flag via constructor or CLI:
python -m speech_to_speech --audio-enhancement
Custom Handler Implementation
For custom flows, instantiate VADHandler directly and process byte chunks:
from speech_to_speech.VAD.vad_handler import VADHandler
from threading import Event
listen_event = Event()
listen_event.set()
handler = VADHandler()
handler.setup(
should_listen=listen_event,
audio_enhancement=True,
)
# Process simulated audio stream
for chunk in audio_byte_chunks:
for output in handler.process(chunk):
if isinstance(output, VADAudio):
# output.audio contains enhanced NumPy array
process_audio(output.audio)
Summary
- VADIterator (
src/speech_to_speech/VAD/vad_iterator.py) handles low-level Silero VAD inference and buffer management usingtorch.hub.load - VADHandler (
src/speech_to_speech/VAD/vad_handler.py) coordinates turn-taking and optionally applies DeepFilterNet enhancement via_apply_audio_enhancement - DeepFilterNet integration requires
pip install "deep-filter-net[torch]"and settingaudio_enhancement=Trueinsetup() - Graceful fallback occurs when DeepFilterNet is unavailable—the handler logs a warning and continues without enhancement
- Resampling is handled automatically when input sample rates differ from DeepFilterNet's internal rate
- Enhanced output is delivered as a
VADAudioevent containing a NumPy array ready for downstream STT or LLM processing
Frequently Asked Questions
What happens if DeepFilterNet is not installed but I set audio_enhancement=True?
The handler catches the import error in setup(), sets HAS_DF to False, logs a warning, and automatically disables enhancement for that session. The VAD functionality continues working normally without noise reduction.
How does the VAD handler manage different sample rates?
The _apply_audio_enhancement method checks if self.sample_rate matches self.df_state.sr(). If they differ, it uses torchaudio.functional.resample to convert the audio to DeepFilterNet's required rate before processing, then resamples back to the original rate before returning.
Can I use VAD without audio enhancement?
Yes. Simply omit the audio_enhancement parameter or set it to False when calling handler.setup(). The VADHandler works independently of DeepFilterNet and will emit VADAudio events containing the original audio buffers without modification.
Where is the enhanced audio stored in the output structure?
The enhanced audio is stored in the audio attribute of the VADAudio object emitted by the handler. This object is defined in src/speech_to_speech/pipeline/messages.py and contains the NumPy array along with metadata such as the mode flag ("final" or "progressive") and turn identifiers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →