How Meetily's Audio Pipeline Uses VAD to Reduce Whisper Processing

Meetily filters silence using a Voice Activity Detection (VAD) processor before audio reaches the Whisper engine, reducing the transcription workload by approximately 70% while maintaining speech fidelity.

Meetily is an open-source meeting assistant built with Rust and Tauri that implements a sophisticated audio pipeline to optimize transcription performance. Instead of sending raw audio directly to OpenAI's Whisper model, the system employs a Voice Activity Detection (VAD) layer that segments speech in real-time. This architecture ensures that only valid speech segments undergo transcription, dramatically cutting computational overhead.

Three-Layer Audio Architecture

Meetily's audio handling follows a pipeline architecture designed to minimize unnecessary processing. The system processes audio through three distinct layers before reaching the transcription engine.

  • Capture and Pre‑Processing: Raw microphone and system audio streams are captured, normalized, and resampled into AudioChunk structures.
  • Mixing and Ring‑Buffer: An AudioMixerRingBuffer aligns dual streams into fixed‑size 50 ms windows, combining them without aggressive ducking.
  • VAD‑Driven Segmentation: The ContinuousVadProcessor analyzes 30 ms frames (480 samples at 16 kHz) and forwards only speech segments to Whisper.

From Capture to Pipeline

The journey begins in AudioCapture::process_audio_data, which constructs AudioChunk instances containing normalized f32 samples, sample rate, timestamp, and device type metadata. These chunks enter the processing pipeline via the RecordingState channel.

In frontend/src-tauri/src/audio/pipeline.rs lines 1224‑1244, the system forwards chunks to the pipeline state:

// Conceptual flow - actual implementation handles channel synchronization
state.send_audio_chunk(AudioChunk {
    data: normalized_samples,
    sample_rate: 48000,
    timestamp: elapsed_seconds,
    chunk_id: sequence_number,
    device_type: DeviceType::Microphone,
})?;

Ring‑Buffer Mixing Strategy

Before VAD analysis, Meetily handles scenarios with both microphone and system audio (e.g., capturing meeting participants and local output). The AudioMixerRingBuffer, initialized at lines 25‑47 in pipeline.rs, stores incoming samples until both streams contain sufficient data for a 50 ms mixing window.

The ProfessionalAudioMixer combines streams using soft‑scaling to prevent clipping, implemented in the mix_window function (lines 54‑86):

// From pipeline.rs - mixing logic extracts aligned windows
let mic_window = ring_buffer.extract_window(DeviceType::Microphone)?;
let sys_window = ring_buffer.extract_window(DeviceType::SystemAudio)?;
let mixed = mixer.mix_window(&mic_window, &sys_window)?;

This produces a single mixed_with_gain buffer ready for voice detection.

Continuous VAD Processing

The core optimization occurs in ContinuousVadProcessor defined in frontend/src-tauri/src/audio/vad.rs. The processor resamples audio to 16 kHz and analyzes it in 30 ms frames.

Configuration happens during initialization (lines 32‑70), where key thresholds are set:

  • Redemption time: Default 400 ms, overridden to 2000 ms in the pipeline to bridge conversational pauses (lines 48‑52)
  • Minimum speech duration: 250 ms (lines 55‑59)

When the pipeline invokes self.vad_processor.process_audio (lines 335‑367 in pipeline.rs), the VAD evaluates energy levels and speech probability. Only segments exceeding min_speech_time emit as SpeechSegment structures.

Feeding Whisper Selectively

Upon detecting valid speech, the pipeline constructs a new AudioChunk at 16 kHz and transmits it via the transcription_sender channel (lines 444‑452 in pipeline.rs):

// Segment emitted from VAD becomes transcription input
let transcription_chunk = AudioChunk {
    data: speech_segment.samples,
    sample_rate: 16000,  // Whisper optimal rate
    timestamp: segment_start_time,
    chunk_id: segment_id,
    device_type: DeviceType::Mixed,
};
transcription_sender.send(transcription_chunk)?;

This channel feeds directly into the Whisper inference engine, which receives only condensed speech rather than continuous audio streams.

Performance Impact on Transcription

By discarding silence at the VAD layer, Meetily avoids sending thousands of seconds of idle audio to Whisper. Empirical logging in process_audio and flush functions indicates the VAD eliminates approximately 70% of raw samples before transcription.

This reduction translates directly to:

  • Lower GPU/CPU utilization during meetings
  • Faster transcription turnaround times
  • Reduced energy consumption on client devices

Using the VAD for Batch Processing

Developers can leverage the VAD independently for offline audio processing. The get_speech_chunks utility function creates a ContinuousVadProcessor with configurable redemption windows:

use meetily::frontend::src_tauri::audio::vad::{get_speech_chunks, SpeechSegment};

// Load 48 kHz mono audio from file
let raw_audio_48k: Vec<f32> = load_wav("meeting.wav")?;
let segments: Vec<SpeechSegment> = get_speech_chunks(&raw_audio_48k, 2000)?;

// Process only speech segments through Whisper
for segment in segments {
    whisper.transcribe(&segment.samples, 16000).await?;
}

Summary

  • Meetily implements a three‑layer audio pipeline that preprocesses, mixes, and filters audio before transcription.
  • The ContinuousVadProcessor in vad.rs analyzes 30 ms frames at 16 kHz, requiring 250 ms minimum speech duration with configurable redemption time (2000 ms default).
  • Ring‑buffer mixing in pipeline.rs aligns dual audio streams into 50 ms windows before VAD analysis.
  • Only validated SpeechSegment structures reach the Whisper engine via dedicated channels.
  • This architecture achieves approximately 70% reduction in audio data processed by Whisper.

Frequently Asked Questions

How does Meetily handle overlapping speech from multiple participants?

The AudioMixerRingBuffer stores microphone and system audio separately until both buffers contain sufficient samples for a 50 ms window. The ProfessionalAudioMixer sums these streams using soft‑scaling rather than aggressive ducking, preserving all speech content while avoiding digital clipping. The combined output then undergoes VAD analysis as a single stream.

What happens if someone pauses mid‑sentence?

The ContinuousVadProcessor implements a redemption time mechanism that bridges short pauses. Configured to 2000 ms in the pipeline (lines 48‑52 of vad.rs), this setting allows the VAD to continue the current speech segment through brief silences, preventing fragmentation of natural speech patterns while still filtering out extended silence.

Can developers adjust the VAD sensitivity?

Yes. When initializing ContinuousVadProcessor or calling get_speech_chunks, developers can modify the min_speech_time parameter (default 250 ms) and redemption_time (default 2000 ms in pipeline usage). Lowering the minimum speech time captures shorter utterances but may increase false positives, while adjusting redemption time controls how aggressively the system splits sentences at pauses.

What audio format does Whisper receive from the pipeline?

Whisper receives AudioChunk structures containing f32 samples at 16 kHz sample rate, regardless of the source input rate. The VAD processor handles resampling internally (see vad.rs lines 32‑70), ensuring Whisper always operates at its optimal input frequency while receiving only speech‑segmented data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →