# Meetily VAD Architecture: How Voice Activity Detection Filters Audio for Transcription

> Explore Meetily's VAD architecture. Learn how its Silero-based pipeline dynamically segments speech, filters audio, and delivers voice-active segments to Whisper for efficient transcription.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: architecture
- Published: 2026-08-02

---

**Meetily employs a Silero-based Voice Activity Detection pipeline that processes mixed 48 kHz audio streams at 16 kHz resolution, dynamically segmenting speech using a 2000 ms redemption time to deliver only voice-active audio to the Whisper transcription engine.**

Meetily is an open-source meeting transcription application built with Rust and Tauri. Its **Voice Activity Detection (VAD) architecture** serves as the critical gateway between raw audio capture and the Whisper transcription engine, ensuring that only valid speech reaches the GPU-intensive inference stage.

## Core Components of the VAD Pipeline

### Dual-Stream Audio Mixing

The pipeline ingests two parallel audio sources: the **microphone stream** capturing user speech and the **system audio stream** recording shared computer sound. These 48 kHz sources are mixed into a single continuous flow before entering the VAD processor located in [`/frontend/src-tauri/src/audio/vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main//frontend/src-tauri/src/audio/vad.rs)【/frontend/src-tauri/src/audio/vad.rs†L33-L65】.

### The VAD Processor

At the heart of the architecture sits the `VadProcessor` struct, which resamples the mixed 48 kHz input to **16 kHz**—the required sample rate for Silero VAD. It processes audio in discrete **30 ms chunks** (480 samples), analyzing each frame for voice presence. The processor is instantiated with specific redemption timing parameters that define how long to wait after speech ends before finalizing a segment.

## How VAD Shapes the Transcription Path

### Speech-Only Segmentation

The VAD stage converts continuous mixed audio into discrete **speech segments** by stripping silence and background noise. This segmentation occurs in [`/frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main//frontend/src-tauri/src/audio/pipeline.rs)【/frontend/src-tauri/src/audio/pipeline.rs†L678-L735】, where the system identifies contiguous voice regions and discards non-speech intervals. By preventing silent audio from reaching Whisper, the architecture eliminates transcription hallucinations that occur when models process noise as words.

### Dynamic Chunking Strategy

Unlike traditional time-based chunking that splits audio into fixed 2-second blocks, Meetily's VAD-driven approach uses **adaptive boundaries** dictated by actual speech detection. The configuration setting `target_chunk_duration_ms` is ignored during VAD processing; instead, the system relies on a **redemption time** of 2000 ms to bridge natural pauses in conversation【/frontend/src-tauri/src/audio/pipeline.rs†L722-L731】. This creates longer, linguistically coherent segments that preserve sentence context better than rigid time slices.

### Pre-Filtering for Whisper

Once the VAD processor emits a validated speech segment, the pipeline dispatches it directly to the `whisper_engine` without additional silence pruning. The system logs each segment's duration and sample count before transmission【/frontend/src-tauri/src/audio/pipeline.rs†L841-L858】, ensuring complete traceability. This guarantees that Whisper receives exclusively high-quality speech data, significantly reducing GPU/CPU load while improving transcription accuracy.

## Implementation Details

Creating the VAD processor involves specifying the input sample rate, target VAD sample rate, and redemption timing:

```rust
let vad = VadProcessor::new(
    input_sample_rate,            // 48 kHz mixed audio
    VAD_SAMPLE_RATE,              // 16 kHz required by Silero VAD
    VAD_REDEMPTION_TIME_MS,       // 2000 ms bridges natural pauses
)?;

```

The pipeline initialization switches between standard chunking and VAD-driven processing:

```rust
info!("VAD-driven audio pipeline started");
pipeline.run_vad().await?;
info!("VAD-driven audio pipeline ended");

```

When a segment is ready, the pipeline forwards it to the transcription worker:

```rust
info!(
    "📤 Sending VAD segment: {:.1}ms, {} samples",
    segment.duration_ms(),
    segment.sample_count()
);
whisper_engine.transcribe(segment).await?;

```

### Key Source Files

- **[`/frontend/src-tauri/src/audio/vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main//frontend/src-tauri/src/audio/vad.rs)**: Contains the Silero-based VAD implementation, resampling logic, and segment detection algorithms【/frontend/src-tauri/src/audio/vad.rs†L33-L65】.
- **[`/frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main//frontend/src-tauri/src/audio/pipeline.rs)**: Orchestrates the audio flow, integrates VAD processing, and manages the handoff to Whisper【/frontend/src-tauri/src/audio/pipeline.rs†L678-L735】【/frontend/src-tauri/src/audio/pipeline.rs†L841-L858】.
- **[`/frontend/src-tauri/src/audio/transcription/worker.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main//frontend/src-tauri/src/audio/transcription/worker.rs)**: Receives VAD-filtered segments and executes the Whisper inference.
- **[`/frontend/src-tauri/src/audio/retranscription.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main//frontend/src-tauri/src/audio/retranscription.rs)**: Implements file-wide VAD processing for replay and retranscription scenarios.

## Summary

- Meetily's VAD architecture processes mixed 48 kHz microphone and system audio through a Silero-based detector running at 16 kHz.
- The system uses **30 ms chunks** (480 samples) to identify speech with a **2000 ms redemption time** that bridges natural conversation pauses.
- **Dynamic chunking** replaces fixed-time segmentation, creating linguistically coherent speech segments that improve transcription context.
- Only validated speech segments reach the Whisper engine, reducing computational overhead and preventing silence-induced hallucinations.
- The implementation spans [`vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/vad.rs) for detection logic and [`pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/pipeline.rs) for integration and orchestration.

## Frequently Asked Questions

### How does Meetily handle audio from both microphone and system sources?

Meetily captures two parallel streams—user microphone input and system audio output—mixes them at 48 kHz, and feeds the combined stream into the VAD processor. This ensures that transcription captures both sides of a conversation while applying uniform speech detection to the mixed signal.

### What is the "redemption time" in Meetily's VAD implementation?

The redemption time is a 2000 ms buffer period that extends speech segments beyond the last detected voice activity. This bridges natural pauses in speech, preventing the system from fragmenting sentences into separate chunks when speakers pause briefly between words or phrases.

### Why does the VAD processor resample audio to 16 kHz?

The Silero VAD model requires 16 kHz input for optimal voice detection accuracy. Meetily's `VadProcessor` handles real-time resampling from the 48 kHz mixed source down to 16 kHz, processing audio in 30 ms windows (480 samples) to maintain low latency while meeting the model's specifications.

### How does VAD improve Whisper transcription performance?

By filtering out silence and non-speech audio before it reaches the transcription engine, VAD reduces the volume of data processed by Whisper, lowering GPU/CPU utilization. This pre-filtering also prevents the model from generating hallucinated text during silent periods, resulting in more accurate and reliable transcripts.