How Does Meetily's Voice Activity Detection (VAD) Filter Audio Before Whisper Transcription?

Meetily uses a multi-stage VAD pipeline that resamples 48 kHz mixed audio to 16 kHz, applies Silero VAD frame-by-frame, merges speech segments with a 2000 ms redemption time, and applies energy-based filtering before sending only valid speech chunks to Whisper.

Meetily's open-source AI meeting assistant relies on precise Voice Activity Detection to reduce computational overhead and prevent transcription hallucinations. The audio pipeline in Zackriya-Solutions/meetily cleanly separates raw capture from transcription by validating every audio chunk through a Rust-based VAD module before it ever reaches the Whisper engine.

The VAD Audio Pipeline Architecture

The complete flow from microphone capture to Whisper transcription follows eight distinct stages, each implemented in specific source files within the Tauri backend.

Step 1: Audio Capture at 48 kHz

Mixed audio from microphone and system streams enters the pipeline at 48 kHz PCM in pipeline.rs. The VAD processor receives this raw buffer immediately after mixing.

// pipeline.rs – VAD-driven pipeline initialization
let vad = VADProcessor::new(
    input_sample_rate,          // 48 kHz mixed audio
    VAD_REDEMPTION_TIME_MS,    // 2000 ms pause-bridging
)?;
info!(
    "VAD-driven pipeline started – segments will be sent directly to Whisper"
);

See the initialization comment and log at lines 722-724 in pipeline.rs【/frontend/src-tauri/src/audio/pipeline.rs#L722-L724】.

Step 2: VAD Processor Configuration

The VADProcessor constructor in vad.rs hardcodes 16 kHz as the target sample rate required by the Silero VAD model. It configures:

  • Input sample rate: 48 kHz (from mixer)
  • Target sample rate: 16 kHz (Silero requirement)
  • Chunk size: 30 ms → 480 samples at 16 kHz
  • Redemption time: 2000 ms to bridge natural speech pauses
// vad.rs – sample rate constants and chunk configuration
// Lines 33-38: 16 kHz target rate constant
// Lines 64-68: 30ms frame = 480 samples comment
// Lines 48-52: Redemption time from retranscription.rs

【/frontend/src-tauri/src/audio/vad.rs#L33-L38】【/frontend/src-tauri/src/audio/vad.rs#L64-L68】【/frontend/src-tauri/src/audio/retranscription.rs#L48-L52】

Step 3: Resampling from 48 kHz to 16 kHz

The handle_resample method internally converts all incoming audio to 16 kHz before VAD analysis. This resampling happens transparently within the processor.

The "Handles resampling" comment at lines 86-90 in vad.rs documents this behavior【/frontend/src-tauri/src/audio/vad.rs#L86-L90】.

Step 4: Frame-Wise Silero VAD Detection

Audio is processed in 30 ms frames fed sequentially to the Silero VAD model. The model emits transition events indicating speech start and end points.

Log messages at lines 226-239 in vad.rs capture these transitions for debugging【/frontend/src-tauri/src/audio/vad.rs#L226-L239】.

Step 5: Speech Segment Accumulation

Detected speech frames are accumulated into contiguous segments. The redemption time (default 2000 ms) prevents over-segmentation by merging brief pauses within ongoing utterances.

This means a speaker's "um" or short breath won't split their sentence into multiple transcription requests.

Step 6: Energy-Based Safety Filter

Before any segment leaves the VAD module, it undergoes RMS and peak energy validation. Low-energy segments that slipped through VAD classification are dropped here to avoid Whisper hallucinations on noise.

// vad.rs – energy validation logs at lines 313-321
// Segments below energy threshold are skipped with explanatory logging

【/frontend/src-tauri/src/audio/vad.rs#L313-L321】

Step 7: Direct Stream to Whisper

Validated speech segments bypass time-based chunking entirely. The VAD determines exact boundaries, and segments are sent as independent audio chunks directly to the Whisper engine.

The send-segment log at lines 841-853 in pipeline.rs confirms this direct handoff【/frontend/src-tauri/src/audio/pipeline.rs#L841-L853】.

Step 8: Final Flush on Recording Stop

When capture ends, any buffered speech remaining in the VAD processor is flushed and sent to Whisper. This ensures no trailing words are lost.

The flush handling spans lines 784-903 in pipeline.rs【/frontend/src-tauri/src/audio/pipeline.rs#L784-L903】.

Processing Audio Through the VAD Loop

The core processing loop in pipeline.rs demonstrates how chunks flow through the system:

// Inside the processing loop (pipeline.rs)
match vad.process(&mixed_audio_chunk) {
    Ok(segments) => {
        for seg in segments {
            // Drop very short segments (< threshold)
            // Valid segments sent directly to Whisper
        }
    }
    Err(e) => {
        error!("VAD processing error: {}", e);
    }
}

Why This VAD Design Matters for Whisper Transcription

Design Choice Benefit
16 kHz resampling Matches Silero's trained sample rate; avoids model mismatch
30 ms frame size Balances latency and detection accuracy
2000 ms redemption time Preserves natural speech flow without over-segmenting
Energy gating Prevents Whisper hallucinations on borderline noise
VAD-driven boundaries Eliminates arbitrary time cuts that split words

According to the Meetily source code, this approach "dramatically reduc[es] unnecessary processing and improv[es] transcription accuracy" compared to naive time-based chunking.

Meetily VAD Configuration Constants

Constant Value Location Purpose
VAD_SAMPLE_RATE 16,000 Hz vad.rs:33 Silero model requirement
VAD_FRAME_SIZE_MS 30 ms vad.rs:64 Frame analysis window
VAD_CHUNK_SIZE 480 samples vad.rs:66 30 ms at 16 kHz
VAD_REDEMPTION_TIME_MS 2000 ms retranscription.rs:48 Pause-bridging threshold

Summary

  • VAD filters audio before Whisper by validating every chunk through Silero VAD at 16 kHz
  • Resampling from 48 kHz mixed audio happens transparently in handle_resample
  • Redemption time of 2000 ms merges natural pauses without splitting utterances
  • Energy gating provides a second safety layer against noise hallucinations
  • Direct segment streaming replaces arbitrary time-based chunking for cleaner transcription boundaries

Frequently Asked Questions

What sample rate does Meetily's VAD use internally?

Meetily's VAD internally processes audio at 16 kHz, regardless of input source. The 48 kHz mixed stream from pipeline.rs is downsampled in vad.rs before Silero analysis. This 16 kHz rate is hardcoded as a constant because Silero VAD models are specifically trained on 16 kHz speech.

Why does Meetily use a 2000 ms redemption time?

The 2000 ms redemption time prevents natural speech pauses from creating false segment boundaries. Without this buffer, a speaker's brief hesitation or breath would split one sentence into multiple transcription requests, fragmenting context and increasing Whisper API calls. The value is configurable but defaults to 2 seconds as defined in retranscription.rs.

How does Meetily prevent Whisper from transcribing noise?

Two mechanisms protect against noise transcription: frame-wise Silero VAD filters obvious silence at the 30 ms level, and energy-based safety filtering applies RMS/peak validation before any segment leaves the VAD module. Segments failing either check are logged and discarded, never reaching Whisper.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →