How Meetily's VAD Filters Audio Before Transcription: A Deep Dive into the Silero-Based Pipeline

Meetily uses a continuous Silero-based Voice Activity Detection (VAD) system that processes 30 ms audio frames, applies redemption time bridging, minimum speech duration checks, and energy threshold filtering to isolate clean speech before sending it to Whisper/Parakeet transcription engines.

The VAD (Voice Activity Detection) pipeline in Meetily sits at the critical boundary between raw audio capture and transcription. Implemented in Rust within the Tauri backend, this system ensures that only high-quality, speech-dense audio reaches computationally expensive speech-to-text models. According to the Meetily source code, the VAD operates in frontend/src-tauri/src/audio/vad.rs and is orchestrated by frontend/src-tauri/src/audio/pipeline.rs.

How the VAD Pipeline Processes Raw Audio

Audio Capture and Mixing

Audio enters the pipeline from two sources: microphone input and system audio (for capturing remote meeting participants). The pipeline.rs module mixes these streams into a single mono buffer called mixed_with_gain:

// pipeline.rs – after mixing microphone + system audio
match self.vad_processor.process_audio(&mixed_with_gain) {
    Ok(speech_segments) => { /* forward each segment to transcription */ }
    Err(e) => warn!("VAD error: {}", e),
}

This mixed buffer is then handed to the VAD processor (see lines 35-36 in pipeline.rs).

Core VAD Processing in vad.rs

The ContinuousVadProcessor struct in vad.rs handles four critical transformations:

  1. Resampling: Any non-16 kHz input is resampled using linear interpolation with a simple low-pass filter (lines 12-30 in vad.rs)
  2. Frame chunking: Audio is buffered and fed to the Silero VadSession in 30 ms frames (480 samples at 16 kHz)
  3. Transition detection: The session emits VadTransition events (SpeechStart, SpeechEnd)
  4. Segment creation: Each SpeechEnd transition produces a SpeechSegment containing raw samples, timestamps, and confidence scores

Configurable VAD Thresholds and Filtering Rules

Silero Session Configuration

The VAD initializes with strict thresholds in ContinuousVadProcessor::new (lines 43-56 in vad.rs):

let mut config = VadConfig::default();
config.positive_speech_threshold = 0.50;  // Confidence to trigger speech start
config.negative_speech_threshold = 0.35;  // Confidence to trigger speech end
config.pre_speech_pad = Duration::from_millis(500);   // Lead-in capture
config.post_speech_pad = Duration::from_millis(500);  // Tail capture
config.min_speech_time = Duration::from_millis(250);  // Minimum segment length
config.redemption_time = Duration::from_millis(2000); // Pause bridging

Post-VAD Energy Filtering

After Silero detection, segments undergo additional filtering in extract_speech_16k (lines 103-116 in vad.rs):

  • RMS threshold: Segments with RMS < 0.2 are discarded as silence
  • Peak threshold: Segments with peak amplitude < 0.2 are dropped as noise
// Energy filtering prevents empty transcripts and hallucinations
let rms = (samples.iter().map(|s| s * s).sum::<f32>() / samples.len() as f32).sqrt();
let peak = samples.iter().map(|s| s.abs()).fold(0.0f32, f32::max);

if rms < 0.2 || peak < 0.2 {
    return None; // Filter out low-energy segments
}

From VAD Segments to Transcription Chunks

Size Enforcement in the Pipeline

Before forwarding to the transcription worker, pipeline.rs enforces a minimum sample count (lines 40-44):

if segment.samples.len() >= 800 {  // ~50 ms minimum at 16 kHz
    let chunk = AudioChunk {
        data: segment.samples,
        sample_rate: 16_000,
        timestamp: segment.start_timestamp_ms / 1000.0,
        chunk_id: self.chunk_id_counter,
        device_type: DeviceType::Microphone,
    };
    self.transcription_sender.send(chunk).unwrap();
    self.chunk_id_counter += 1;
}

Transcription Worker Integration

The worker.rs module receives these AudioChunk objects and dispatches them to the configured engine (Whisper or Parakeet). Because VAD filtering already occurred, the transcription engine operates only on clean, speech-dense audio.

Practical Code Examples

Using the VAD Utility Directly

use meetily::frontend::src_tauri::src::audio::vad::{get_speech_chunks, SpeechSegment};

fn filter_audio(samples_16k: &[f32]) -> Result<Vec<SpeechSegment>> {
    // 2000 ms redemption window bridges natural pauses in conversation
    get_speech_chunks(samples_16k, 2000)
}

Real-Time VAD Processing with Custom Sample Rates

use crate::audio::vad::ContinuousVadProcessor;

fn process_live_audio() -> Result<()> {
    // Initialize with 48 kHz input; internal resampling to 16 kHz
    let mut vad = ContinuousVadProcessor::new(48000, 2000)?;

    // Process audio in real-time chunks
    for chunk in audio_stream.chunks(480) { // 30 ms @ 48 kHz
        let segments = vad.process_audio(chunk)?;
        for seg in segments {
            send_to_transcription(seg);
        }
    }

    // Ensure final utterances aren't lost
    let trailing = vad.flush()?;
    process_final_segments(trailing);
    Ok(())
}

Why This Filtering Architecture Matters

Filter Stage Purpose Transcription Impact
Redemption time (2000 ms) Bridges natural pauses, prevents utterance fragmentation Reduces tiny segments that Whisper rejects
Minimum speech time (250 ms) Enforces Whisper's length requirements Eliminates sub-100 ms hallucination-prone inputs
Energy thresholds (RMS/peak < 0.2) Removes silence and background noise Prevents empty transcripts and false positives
Chunk size check (≥800 samples) Final pipeline safety net Guarantees transcription worker receives viable audio

Summary

  • Meetily's VAD uses Silero in 30 ms frames with configurable redemption time and speech duration thresholds
  • Resampling and energy filtering occur in vad.rs before segments reach transcription
  • Pipeline enforcement in pipeline.rs adds a final size-based filter at the transcription boundary
  • Clean audio output dramatically reduces Whisper/Parakeet errors and computational waste

Frequently Asked Questions

What sample rate does Meetily's VAD require?

The VAD internally requires 16 kHz, but ContinuousVadProcessor accepts any input sample rate. The resample_to_16k function in vad.rs (lines 12-30) performs linear interpolation with low-pass filtering to avoid aliasing when resampling non-16 kHz sources like 48 kHz system audio.

Why does Meetily use a 2000 ms redemption time?

The 2000 ms redemption window bridges natural pauses in human speech, preventing a single utterance from being split into multiple fragments. Without this, Whisper would receive sub-optimal segments that often produce truncated or hallucinated transcriptions. This value is configurable in VadConfig::redemption_time.

How does Meetily prevent transcription of pure silence?

Three layers filter silence: (1) Silero's positive_speech_threshold (0.50) confidence gate, (2) post-detection RMS and peak energy thresholds (both < 0.2), and (3) the minimum 250 ms speech duration requirement. Segments failing any check are discarded before reaching the transcription worker.

What happens to audio when a meeting ends abruptly?

The flush() method in ContinuousVadProcessor (lines 62-84 in vad.rs) guarantees that any trailing speech still in the VAD buffer is emitted as final SpeechSegment objects. This prevents loss of the speaker's final words when the audio stream terminates unexpectedly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →