How Meetily's VAD Filters Audio Before Transcription: A Deep Dive into the Silero-Based Pipeline
Meetily uses a continuous Silero-based Voice Activity Detection (VAD) system that processes 30 ms audio frames, applies redemption time bridging, minimum speech duration checks, and energy threshold filtering to isolate clean speech before sending it to Whisper/Parakeet transcription engines.
The VAD (Voice Activity Detection) pipeline in Meetily sits at the critical boundary between raw audio capture and transcription. Implemented in Rust within the Tauri backend, this system ensures that only high-quality, speech-dense audio reaches computationally expensive speech-to-text models. According to the Meetily source code, the VAD operates in frontend/src-tauri/src/audio/vad.rs and is orchestrated by frontend/src-tauri/src/audio/pipeline.rs.
How the VAD Pipeline Processes Raw Audio
Audio Capture and Mixing
Audio enters the pipeline from two sources: microphone input and system audio (for capturing remote meeting participants). The pipeline.rs module mixes these streams into a single mono buffer called mixed_with_gain:
// pipeline.rs – after mixing microphone + system audio
match self.vad_processor.process_audio(&mixed_with_gain) {
Ok(speech_segments) => { /* forward each segment to transcription */ }
Err(e) => warn!("VAD error: {}", e),
}
This mixed buffer is then handed to the VAD processor (see lines 35-36 in pipeline.rs).
Core VAD Processing in vad.rs
The ContinuousVadProcessor struct in vad.rs handles four critical transformations:
- Resampling: Any non-16 kHz input is resampled using linear interpolation with a simple low-pass filter (lines 12-30 in vad.rs)
- Frame chunking: Audio is buffered and fed to the Silero
VadSessionin 30 ms frames (480 samples at 16 kHz) - Transition detection: The session emits
VadTransitionevents (SpeechStart,SpeechEnd) - Segment creation: Each
SpeechEndtransition produces aSpeechSegmentcontaining raw samples, timestamps, and confidence scores
Configurable VAD Thresholds and Filtering Rules
Silero Session Configuration
The VAD initializes with strict thresholds in ContinuousVadProcessor::new (lines 43-56 in vad.rs):
let mut config = VadConfig::default();
config.positive_speech_threshold = 0.50; // Confidence to trigger speech start
config.negative_speech_threshold = 0.35; // Confidence to trigger speech end
config.pre_speech_pad = Duration::from_millis(500); // Lead-in capture
config.post_speech_pad = Duration::from_millis(500); // Tail capture
config.min_speech_time = Duration::from_millis(250); // Minimum segment length
config.redemption_time = Duration::from_millis(2000); // Pause bridging
Post-VAD Energy Filtering
After Silero detection, segments undergo additional filtering in extract_speech_16k (lines 103-116 in vad.rs):
- RMS threshold: Segments with RMS < 0.2 are discarded as silence
- Peak threshold: Segments with peak amplitude < 0.2 are dropped as noise
// Energy filtering prevents empty transcripts and hallucinations
let rms = (samples.iter().map(|s| s * s).sum::<f32>() / samples.len() as f32).sqrt();
let peak = samples.iter().map(|s| s.abs()).fold(0.0f32, f32::max);
if rms < 0.2 || peak < 0.2 {
return None; // Filter out low-energy segments
}
From VAD Segments to Transcription Chunks
Size Enforcement in the Pipeline
Before forwarding to the transcription worker, pipeline.rs enforces a minimum sample count (lines 40-44):
if segment.samples.len() >= 800 { // ~50 ms minimum at 16 kHz
let chunk = AudioChunk {
data: segment.samples,
sample_rate: 16_000,
timestamp: segment.start_timestamp_ms / 1000.0,
chunk_id: self.chunk_id_counter,
device_type: DeviceType::Microphone,
};
self.transcription_sender.send(chunk).unwrap();
self.chunk_id_counter += 1;
}
Transcription Worker Integration
The worker.rs module receives these AudioChunk objects and dispatches them to the configured engine (Whisper or Parakeet). Because VAD filtering already occurred, the transcription engine operates only on clean, speech-dense audio.
Practical Code Examples
Using the VAD Utility Directly
use meetily::frontend::src_tauri::src::audio::vad::{get_speech_chunks, SpeechSegment};
fn filter_audio(samples_16k: &[f32]) -> Result<Vec<SpeechSegment>> {
// 2000 ms redemption window bridges natural pauses in conversation
get_speech_chunks(samples_16k, 2000)
}
Real-Time VAD Processing with Custom Sample Rates
use crate::audio::vad::ContinuousVadProcessor;
fn process_live_audio() -> Result<()> {
// Initialize with 48 kHz input; internal resampling to 16 kHz
let mut vad = ContinuousVadProcessor::new(48000, 2000)?;
// Process audio in real-time chunks
for chunk in audio_stream.chunks(480) { // 30 ms @ 48 kHz
let segments = vad.process_audio(chunk)?;
for seg in segments {
send_to_transcription(seg);
}
}
// Ensure final utterances aren't lost
let trailing = vad.flush()?;
process_final_segments(trailing);
Ok(())
}
Why This Filtering Architecture Matters
| Filter Stage | Purpose | Transcription Impact |
|---|---|---|
| Redemption time (2000 ms) | Bridges natural pauses, prevents utterance fragmentation | Reduces tiny segments that Whisper rejects |
| Minimum speech time (250 ms) | Enforces Whisper's length requirements | Eliminates sub-100 ms hallucination-prone inputs |
| Energy thresholds (RMS/peak < 0.2) | Removes silence and background noise | Prevents empty transcripts and false positives |
| Chunk size check (≥800 samples) | Final pipeline safety net | Guarantees transcription worker receives viable audio |
Summary
- Meetily's VAD uses Silero in 30 ms frames with configurable redemption time and speech duration thresholds
- Resampling and energy filtering occur in
vad.rsbefore segments reach transcription - Pipeline enforcement in
pipeline.rsadds a final size-based filter at the transcription boundary - Clean audio output dramatically reduces Whisper/Parakeet errors and computational waste
Frequently Asked Questions
What sample rate does Meetily's VAD require?
The VAD internally requires 16 kHz, but ContinuousVadProcessor accepts any input sample rate. The resample_to_16k function in vad.rs (lines 12-30) performs linear interpolation with low-pass filtering to avoid aliasing when resampling non-16 kHz sources like 48 kHz system audio.
Why does Meetily use a 2000 ms redemption time?
The 2000 ms redemption window bridges natural pauses in human speech, preventing a single utterance from being split into multiple fragments. Without this, Whisper would receive sub-optimal segments that often produce truncated or hallucinated transcriptions. This value is configurable in VadConfig::redemption_time.
How does Meetily prevent transcription of pure silence?
Three layers filter silence: (1) Silero's positive_speech_threshold (0.50) confidence gate, (2) post-detection RMS and peak energy thresholds (both < 0.2), and (3) the minimum 250 ms speech duration requirement. Segments failing any check are discarded before reaching the transcription worker.
What happens to audio when a meeting ends abruptly?
The flush() method in ContinuousVadProcessor (lines 62-84 in vad.rs) guarantees that any trailing speech still in the VAD buffer is emitted as final SpeechSegment objects. This prevents loss of the speaker's final words when the audio stream terminates unexpectedly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →