# How to Implement VAD with DeepFilterNet Audio Enhancement in Speech-to-Speech

> Learn to implement VAD with DeepFilterNet audio enhancement in Speech-to-Speech. Discover how the VADHandler simplifies voice activity detection and noise reduction for cleaner audio.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**The Speech-to-Speech repository provides a `VADHandler` class that integrates Silero VAD for voice activity detection and optionally applies DeepFilterNet audio enhancement by setting `audio_enhancement=True` during setup, which triggers noise reduction via the `_apply_audio_enhancement` method before emitting `VADAudio` events.**

The huggingface/speech-to-speech repository offers a production-ready pipeline for real-time voice processing that combines Voice Activity Detection (VAD) with DeepFilterNet audio enhancement. This guide explains how to implement VAD with DeepFilterNet audio enhancement using the library's modular handler architecture, referencing actual source file paths and configuration patterns found in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) and related modules.

## Architecture Overview

The VAD stack consists of two primary components that coordinate to detect speech and optionally enhance audio quality before downstream processing.

### VADIterator: Core Speech Detection Engine

The **`VADIterator`** class in [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py) serves as the low-level VAD engine. It consumes raw audio chunks, maintains a pre-speech buffer, and determines when speech segments start and end using the Silero VAD model loaded via `torch.hub.load`.

Key characteristics include:
- **Trigger management**: Maintains a `self.triggered` flag that indicates active utterances
- **Buffer handling**: Manages silence padding and speech-pad configurations
- **Output format**: Returns a list of `torch.Tensor` chunks when speech ends, otherwise `None`

### VADHandler: Orchestration and Enhancement Layer

The **`VADHandler`** class in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) orchestrates the iterator, manages turn metadata, and optionally runs DeepFilterNet on finalized audio segments. This handler is instantiated per session and receives a `should_listen` event to control processing flow.

Critical implementation details:
- **Optional dependency**: DeepFilterNet import is guarded with a try/except block that sets a `HAS_DF` flag
- **Graceful degradation**: If the import fails, the library logs a warning and continues without enhancement capabilities
- **Event emission**: Produces `VADAudio` events (defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py)) containing either raw or enhanced NumPy arrays

## Enabling DeepFilterNet Enhancement

To implement audio enhancement, you must first install the optional dependency and then configure the handler during initialization.

### Installation Requirements

Install DeepFilterNet with PyTorch support:

```bash
pip install "deep-filter-net[torch]"

```

### Handler Configuration

Activate enhancement by passing `audio_enhancement=True` to the `setup` method:

```python
from speech_to_speech.VAD.vad_handler import VADHandler

handler = VADHandler()
handler.setup(
    should_listen=listen_event,
    thresh=0.6,
    sample_rate=16000,
    min_silence_ms=64,
    min_speech_ms=384,
    audio_enhancement=True,  # Enable DeepFilterNet

)

```

The `setup` method (lines 44-51 in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py)) attempts to import `df.enhance` and initialize the model via `init_df()`. If the import fails, `HAS_DF` becomes `False` and the `audio_enhancement` flag is automatically cleared to prevent runtime errors.

## Runtime Execution Flow

When processing audio streams, the handler applies enhancement at the final speech boundary before emitting events.

### The Enhancement Pipeline

When `audio_enhancement=True` and a speech segment completes, the handler invokes `_apply_audio_enhancement` (lines 110-133 in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py)). This method performs the following operations:

1. **Resampling**: Converts audio to DeepFilterNet's internal sample rate (`self.df_state.sr()`) if necessary using `torchaudio.functional.resample`
2. **Enhancement**: Processes through `enhance(self.enhanced_model, self.df_state, audio_tensor)`
3. **Restoration**: Resamples back to the original sample rate
4. **Output**: Returns the enhanced NumPy array that replaces the original in the `VADAudio` event

```python
import torchaudio

if self.sample_rate != self.df_state.sr():
    audio_float32 = torchaudio.functional.resample(
        torch.from_numpy(array), orig_freq=self.sample_rate, new_freq=self.df_state.sr()
    )
    enhanced = enhance(self.enhanced_model, self.df_state, audio_float32.unsqueeze(0))
    enhanced = torchaudio.functional.resample(
        enhanced, orig_freq=self.df_state.sr(), new_freq=self.sample_rate
    )
else:
    enhanced = enhance(self.enhanced_model, self.df_state, torch.from_numpy(array))
return enhanced.numpy().squeeze()

```

## Integration Patterns

You can integrate this functionality through the high-level pipeline or directly in custom implementations.

### Using the S2SPipeline

When using the built-in `S2SPipeline`, pass the enhancement flag via constructor or CLI:

```bash
python -m speech_to_speech --audio-enhancement

```

### Custom Handler Implementation

For custom flows, instantiate `VADHandler` directly and process byte chunks:

```python
from speech_to_speech.VAD.vad_handler import VADHandler
from threading import Event

listen_event = Event()
listen_event.set()

handler = VADHandler()
handler.setup(
    should_listen=listen_event,
    audio_enhancement=True,
)

# Process simulated audio stream

for chunk in audio_byte_chunks:
    for output in handler.process(chunk):
        if isinstance(output, VADAudio):
            # output.audio contains enhanced NumPy array

            process_audio(output.audio)

```

## Summary

- **VADIterator** ([`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py)) handles low-level Silero VAD inference and buffer management using `torch.hub.load`
- **VADHandler** ([`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py)) coordinates turn-taking and optionally applies DeepFilterNet enhancement via `_apply_audio_enhancement`
- **DeepFilterNet integration** requires `pip install "deep-filter-net[torch]"` and setting `audio_enhancement=True` in `setup()`
- **Graceful fallback** occurs when DeepFilterNet is unavailable—the handler logs a warning and continues without enhancement
- **Resampling** is handled automatically when input sample rates differ from DeepFilterNet's internal rate
- **Enhanced output** is delivered as a `VADAudio` event containing a NumPy array ready for downstream STT or LLM processing

## Frequently Asked Questions

### What happens if DeepFilterNet is not installed but I set audio_enhancement=True?

The handler catches the import error in `setup()`, sets `HAS_DF` to `False`, logs a warning, and automatically disables enhancement for that session. The VAD functionality continues working normally without noise reduction.

### How does the VAD handler manage different sample rates?

The `_apply_audio_enhancement` method checks if `self.sample_rate` matches `self.df_state.sr()`. If they differ, it uses `torchaudio.functional.resample` to convert the audio to DeepFilterNet's required rate before processing, then resamples back to the original rate before returning.

### Can I use VAD without audio enhancement?

Yes. Simply omit the `audio_enhancement` parameter or set it to `False` when calling `handler.setup()`. The `VADHandler` works independently of DeepFilterNet and will emit `VADAudio` events containing the original audio buffers without modification.

### Where is the enhanced audio stored in the output structure?

The enhanced audio is stored in the `audio` attribute of the `VADAudio` object emitted by the handler. This object is defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py) and contains the NumPy array along with metadata such as the mode flag (`"final"` or `"progressive"`) and turn identifiers.