How Does Meetily's Voice Activity Detection (VAD) Filter Audio Before Whisper Transcription?
Meetily uses a multi-stage VAD pipeline that resamples 48 kHz mixed audio to 16 kHz, applies Silero VAD frame-by-frame, merges speech segments with a 2000 ms redemption time, and applies energy-based filtering before sending only valid speech chunks to Whisper.
Meetily's open-source AI meeting assistant relies on precise Voice Activity Detection to reduce computational overhead and prevent transcription hallucinations. The audio pipeline in Zackriya-Solutions/meetily cleanly separates raw capture from transcription by validating every audio chunk through a Rust-based VAD module before it ever reaches the Whisper engine.
The VAD Audio Pipeline Architecture
The complete flow from microphone capture to Whisper transcription follows eight distinct stages, each implemented in specific source files within the Tauri backend.
Step 1: Audio Capture at 48 kHz
Mixed audio from microphone and system streams enters the pipeline at 48 kHz PCM in pipeline.rs. The VAD processor receives this raw buffer immediately after mixing.
// pipeline.rs – VAD-driven pipeline initialization
let vad = VADProcessor::new(
input_sample_rate, // 48 kHz mixed audio
VAD_REDEMPTION_TIME_MS, // 2000 ms pause-bridging
)?;
info!(
"VAD-driven pipeline started – segments will be sent directly to Whisper"
);
See the initialization comment and log at lines 722-724 in pipeline.rs【/frontend/src-tauri/src/audio/pipeline.rs#L722-L724】.
Step 2: VAD Processor Configuration
The VADProcessor constructor in vad.rs hardcodes 16 kHz as the target sample rate required by the Silero VAD model. It configures:
- Input sample rate: 48 kHz (from mixer)
- Target sample rate: 16 kHz (Silero requirement)
- Chunk size: 30 ms → 480 samples at 16 kHz
- Redemption time: 2000 ms to bridge natural speech pauses
// vad.rs – sample rate constants and chunk configuration
// Lines 33-38: 16 kHz target rate constant
// Lines 64-68: 30ms frame = 480 samples comment
// Lines 48-52: Redemption time from retranscription.rs
【/frontend/src-tauri/src/audio/vad.rs#L33-L38】【/frontend/src-tauri/src/audio/vad.rs#L64-L68】【/frontend/src-tauri/src/audio/retranscription.rs#L48-L52】
Step 3: Resampling from 48 kHz to 16 kHz
The handle_resample method internally converts all incoming audio to 16 kHz before VAD analysis. This resampling happens transparently within the processor.
The "Handles resampling" comment at lines 86-90 in vad.rs documents this behavior【/frontend/src-tauri/src/audio/vad.rs#L86-L90】.
Step 4: Frame-Wise Silero VAD Detection
Audio is processed in 30 ms frames fed sequentially to the Silero VAD model. The model emits transition events indicating speech start and end points.
Log messages at lines 226-239 in vad.rs capture these transitions for debugging【/frontend/src-tauri/src/audio/vad.rs#L226-L239】.
Step 5: Speech Segment Accumulation
Detected speech frames are accumulated into contiguous segments. The redemption time (default 2000 ms) prevents over-segmentation by merging brief pauses within ongoing utterances.
This means a speaker's "um" or short breath won't split their sentence into multiple transcription requests.
Step 6: Energy-Based Safety Filter
Before any segment leaves the VAD module, it undergoes RMS and peak energy validation. Low-energy segments that slipped through VAD classification are dropped here to avoid Whisper hallucinations on noise.
// vad.rs – energy validation logs at lines 313-321
// Segments below energy threshold are skipped with explanatory logging
【/frontend/src-tauri/src/audio/vad.rs#L313-L321】
Step 7: Direct Stream to Whisper
Validated speech segments bypass time-based chunking entirely. The VAD determines exact boundaries, and segments are sent as independent audio chunks directly to the Whisper engine.
The send-segment log at lines 841-853 in pipeline.rs confirms this direct handoff【/frontend/src-tauri/src/audio/pipeline.rs#L841-L853】.
Step 8: Final Flush on Recording Stop
When capture ends, any buffered speech remaining in the VAD processor is flushed and sent to Whisper. This ensures no trailing words are lost.
The flush handling spans lines 784-903 in pipeline.rs【/frontend/src-tauri/src/audio/pipeline.rs#L784-L903】.
Processing Audio Through the VAD Loop
The core processing loop in pipeline.rs demonstrates how chunks flow through the system:
// Inside the processing loop (pipeline.rs)
match vad.process(&mixed_audio_chunk) {
Ok(segments) => {
for seg in segments {
// Drop very short segments (< threshold)
// Valid segments sent directly to Whisper
}
}
Err(e) => {
error!("VAD processing error: {}", e);
}
}
Why This VAD Design Matters for Whisper Transcription
| Design Choice | Benefit |
|---|---|
| 16 kHz resampling | Matches Silero's trained sample rate; avoids model mismatch |
| 30 ms frame size | Balances latency and detection accuracy |
| 2000 ms redemption time | Preserves natural speech flow without over-segmenting |
| Energy gating | Prevents Whisper hallucinations on borderline noise |
| VAD-driven boundaries | Eliminates arbitrary time cuts that split words |
According to the Meetily source code, this approach "dramatically reduc[es] unnecessary processing and improv[es] transcription accuracy" compared to naive time-based chunking.
Meetily VAD Configuration Constants
| Constant | Value | Location | Purpose |
|---|---|---|---|
VAD_SAMPLE_RATE |
16,000 Hz | vad.rs:33 |
Silero model requirement |
VAD_FRAME_SIZE_MS |
30 ms | vad.rs:64 |
Frame analysis window |
VAD_CHUNK_SIZE |
480 samples | vad.rs:66 |
30 ms at 16 kHz |
VAD_REDEMPTION_TIME_MS |
2000 ms | retranscription.rs:48 |
Pause-bridging threshold |
Summary
- VAD filters audio before Whisper by validating every chunk through Silero VAD at 16 kHz
- Resampling from 48 kHz mixed audio happens transparently in
handle_resample - Redemption time of 2000 ms merges natural pauses without splitting utterances
- Energy gating provides a second safety layer against noise hallucinations
- Direct segment streaming replaces arbitrary time-based chunking for cleaner transcription boundaries
Frequently Asked Questions
What sample rate does Meetily's VAD use internally?
Meetily's VAD internally processes audio at 16 kHz, regardless of input source. The 48 kHz mixed stream from pipeline.rs is downsampled in vad.rs before Silero analysis. This 16 kHz rate is hardcoded as a constant because Silero VAD models are specifically trained on 16 kHz speech.
Why does Meetily use a 2000 ms redemption time?
The 2000 ms redemption time prevents natural speech pauses from creating false segment boundaries. Without this buffer, a speaker's brief hesitation or breath would split one sentence into multiple transcription requests, fragmenting context and increasing Whisper API calls. The value is configurable but defaults to 2 seconds as defined in retranscription.rs.
How does Meetily prevent Whisper from transcribing noise?
Two mechanisms protect against noise transcription: frame-wise Silero VAD filters obvious silence at the 30 ms level, and energy-based safety filtering applies RMS/peak validation before any segment leaves the VAD module. Segments failing either check are logged and discarded, never reaching Whisper.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →