Voice Activity Detection (VAD) Implementation in Meetily: How It Filters Audio for Real-Time Transcription

Meetily uses a three-layer VAD architecture built on the silero-rs library that resamples audio to 16 kHz, buffers 30 ms chunks, and applies a 400 ms redemption time to bridge natural pauses before sending speech-only segments to Whisper for transcription.

Meetily’s real-time transcription engine relies on a tightly-coupled Voice Activity Detection (VAD) subsystem that isolates speech before feeding it to the Whisper model. According to the Zackriya-Solutions/meetily source code, this integration reduces the audio volume sent to transcription by approximately 70% while preserving full mixed audio for recordings.

The VAD Engine: ContinuousVadProcessor

The core VAD implementation lives in frontend/src-tauri/src/audio/vad.rs within the ContinuousVadProcessor struct. This Rust module wraps the Silero VAD library and handles resampling, chunking, and speech segmentation.

Initialization and Configuration

When the audio pipeline starts, it instantiates the processor with the capture device’s sample rate and a configurable redemption time (400 ms on all platforms). The implementation sets strict thresholds to avoid fragmenting short speech bursts:

// From vad.rs [L46-L57]
config.redemption_time = Duration::from_millis(redemption_time_ms as u64);
config.min_speech_time = Duration::from_millis(250);

The ContinuousVadProcessor::new() constructor (lines 32-50 of vad.rs) establishes these parameters and prepares the Silero session for 16 kHz processing.

Resampling and Chunk Processing

If the capture device does not run at 16 kHz, the process_audio method resamples incoming samples using resample_to_16k (source: lines 88-94 and 12-15 of vad.rs). The engine buffers audio until a full 30 ms window (480 samples at 16 kHz) is available before feeding it to the Silero session (lines 99-108).

Speech Segmentation and Redemption Time

The VAD emits complete speech segments by tracking transitions between speech and silence. A configurable redemption time of 400 ms bridges natural pauses in conversation, preventing over-segmentation while ensuring responsiveness. Completed segments are converted into SpeechSegment structs and queued for the pipeline.

Audio Pipeline Integration

The AudioPipeline struct in frontend/src-tauri/src/audio/pipeline.rs orchestrates the flow between raw audio capture and VAD-filtered output.

Mixing Microphone and System Audio

Raw microphone and system audio streams accumulate in AudioMixerRingBuffer. When sufficient samples exist, ProfessionalAudioMixer::mix_window produces a mixed 48 kHz signal. This mixed audio is immediately forwarded to the VAD processor:

// From pipeline.rs [L35-L67] (excerpt)
let speech_segments = self.vad_processor.process_audio(&mixed)?;

Segment Filtering and Dispatch

For each SpeechSegment returned by the VAD, the pipeline applies a minimum length filter (≥ 800 samples, approximately 50 ms at 16 kHz) to remove noise artifacts. Valid segments are repackaged as AudioChunk structs with a fixed 16 kHz sample rate and dispatched to the transcription channel (lines 38-56 of pipeline.rs):

if seg.samples.len() >= 800 {
    let transcription_chunk = AudioChunk {
        data: seg.samples,
        sample_rate: 16000,
        timestamp: seg.start_timestamp_ms / 1000.0,
        chunk_id: self.chunk_id_counter,
        device_type: DeviceType::Microphone,
    };
    self.transcription_sender.send(transcription_chunk)?;
}

Simultaneously, the unfiltered mixed audio forwards to the recording saver, ensuring WAV files contain the complete session while transcripts reflect only speech.

Shutdown Flushing

When the pipeline terminates, flush_remaining_audio (lines 100-128) triggers vad_processor.flush() to force-close any open speech segments and emit final audio chunks, preventing data loss at the end of sessions.

Transcription Worker Consumption

The transcription worker in frontend/src-tauri/src/audio/transcription/worker.rs receives VAD-generated AudioChunks via channel and forwards them directly to Whisper. The source code explicitly notes that no additional VAD processing occurs at this stage:

// From worker.rs [L420-L422]
// Skip VAD processing here since the pipeline already extracted speech using VAD

This design separates concerns: the pipeline handles audio preprocessing and speech isolation, while the worker focuses solely on inference.

Configuration Parameters and Performance Impact

Meetily’s VAD configuration balances latency reductions with transcription accuracy:

Parameter Value Purpose
Target Sample Rate 16 kHz Required by Silero VAD model
Chunk Size 30 ms (480 samples) Real-time processing window
Redemption Time 400 ms Bridges natural speech pauses
Minimum Speech Time 250 ms Filters out short noise bursts
Segment Threshold 800 samples (~50 ms) Final noise gate before transcription

This architecture ensures that only approximately 30% of captured audio (speech segments) reaches the Whisper model, significantly reducing compute overhead while the full mixed stream records to disk.

Summary

  • Meetily implements VAD through the ContinuousVadProcessor in frontend/src-tauri/src/audio/vad.rs, wrapping the Silero VAD library.
  • The AudioPipeline in pipeline.rs mixes microphone and system audio, processes it through the VAD, and dispatches speech-only chunks to transcription.
  • A 400 ms redemption time bridges natural pauses to prevent over-segmentation.
  • The transcription worker deliberately skips VAD processing because the pipeline already filtered the audio.
  • On shutdown, flush_remaining_audio ensures all buffered speech segments reach the transcription engine.

Frequently Asked Questions

What VAD library does Meetily use?

Meetily uses the Silero VAD library via the silero-rs Rust bindings, as implemented in frontend/src-tauri/src/audio/vad.rs. This neural network-based detector runs locally and processes 16 kHz audio chunks in real-time.

How does Meetily handle different audio sample rates?

If the capture device does not provide 16 kHz audio, the ContinuousVadProcessor automatically resamples incoming data using resample_to_16k before feeding it to the Silero session. This ensures consistent VAD performance regardless of input hardware.

What is redemption time in Meetily's VAD?

The redemption time is a 400 ms buffer window that extends speech segments slightly beyond detected silence. This bridges natural conversational pauses (like breaths between sentences) to prevent the transcript from splitting into excessive fragments. It is configurable during ContinuousVadProcessor initialization.

Why does the transcription worker skip VAD processing?

The transcription worker in worker.rs skips VAD because the audio pipeline already performed speech extraction. The comment at lines 420-422 clarifies this design: "Skip VAD processing here since the pipeline already extracted speech using VAD." This prevents duplicate processing and reduces latency in the transcription thread.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →