How Meetily's Ring Buffer Mixing Algorithm Synchronizes Microphone and System Audio Streams

Meetily's ring buffer mixing algorithm aligns asynchronous microphone and system audio streams into fixed 600ms windows using a dual-queue buffering system that zero-pads missing data to prevent glitches and maintain transcription accuracy.

Meetily, developed by Zackriya-Solutions, captures both user microphone input and desktop system audio for meeting transcription. Because these independent streams arrive with variable latency and chunk sizes, Meetily's ring buffer mixing algorithm synchronizes them before mixing, ensuring the transcription pipeline receives perfectly aligned audio data regardless of network jitter or capture delays.

Core Ring Buffer Architecture

The synchronization mechanism centers on AudioMixerRingBuffer, implemented in frontend/src-tauri/src/audio/pipeline.rs. This struct maintains separate buffers for each audio source and manages the temporal alignment required for clean mixing.

The AudioMixerRingBuffer Structure

The ring buffer uses two VecDeque<f32> queues to store incoming samples:

  • mic_buffer: Holds raw microphone samples from the user's audio input device
  • system_buffer: Stores system audio samples captured from the desktop output
  • window_size_samples: Defines the fixed window duration (600ms at the pipeline sample rate) for mixing operations
  • max_buffer_size: Enforces a safety limit of approximately 400ms to prevent uncontrolled memory growth when one stream lags

The buffer initializes with a calculated window size based on the pipeline sample rate:

let window_ms = 600.0;  // 600ms window for mixing
let window_size_samples = (sample_rate as f32 * window_ms / 1000.0) as usize;

Safety Limits and Memory Management

The max_buffer_size field acts as a circuit breaker against memory exhaustion. When either buffer exceeds this limit due to persistent jitter or capture failures, the oldest samples are automatically dropped with a logged warning, ensuring the application remains responsive even during degraded audio capture conditions.

Buffering and Synchronization Logic

The ring buffer operates on a continuous flow of AudioChunk structures, routing samples to the appropriate queue and extracting aligned windows only when sufficient data exists for processing.

Ingesting Audio Chunks

When the audio capture threads generate new chunks, the pipeline calls add_samples to route data into the correct buffer:

self.ring_buffer.add_samples(chunk.device_type.clone(), chunk.data);

This method appends incoming f32 samples to either mic_buffer or system_buffer based on the DeviceType enum (defined in recording_state.rs). The non-blocking ingestion allows capture threads to operate independently of the mixing frequency.

Mixing Window Detection

The can_mix() method determines when sufficient data exists to generate a synchronized output window:

fn can_mix(&self) -> bool {
    self.mic_buffer.len() >= self.window_size_samples ||
    self.system_buffer.len() >= self.window_size_samples
}

This design uses an OR condition rather than AND, meaning mixing proceeds as soon as either stream accumulates a full 600ms window. This prevents pipeline stalls when one source experiences temporary dropouts, with the missing stream compensated via zero-padding during extraction.

Alignment via Zero-Padding

The extract_window() method performs the critical synchronization by draining samples and padding deficits:

  1. Drain available samples up to window_size_samples from each buffer
  2. Zero-pad the depleted buffer when insufficient samples exist, preventing "last-sample-hold" artifacts that would cause audio distortion
  3. Return aligned tuples ready for the mixing stage
let mic_window = if self.mic_buffer.len() >= self.window_size_samples {
    self.mic_buffer.drain(0..self.window_size_samples).collect()
} else if !self.mic_buffer.is_empty() {
    let mut padded = self.mic_buffer.drain(..).collect::<Vec<f32>>();
    padded.resize(self.window_size_samples, 0.0);
    padded
} else {
    vec![0.0; self.window_size_samples]
};

This zero-padding strategy ensures continuous output without audible clicks or gaps, maintaining synchronization even when one stream temporarily fails or arrives significantly later than the other.

Professional Audio Mixing and Output

Once aligned, the windows pass to ProfessionalAudioMixer::mix_window, which combines the streams using soft clipping to prevent distortion:

let sys_scaled = sys * 1.0;  // System audio scaling factor
let sum = mic + sys_scaled;
let mixed_sample = if sum.abs() > 1.0 { sum / sum.abs() } else { sum };

The microphone audio enters this stage already normalized to -23 LUFS from the capture preprocessing (handled in audio_processing.rs), requiring no additional post-gain. The mixed output then routes simultaneously to:

  • The VAD (Voice Activity Detection) processor for transcription via Whisper
  • The recording sender for WAV file persistence

Integration with the Audio Pipeline

The complete data flow through Meetily's audio pipeline demonstrates how the ring buffer integrates with capture and processing stages:

  1. Capture: AudioCapture::process_audio_data in audio_capture.rs generates AudioChunk structures after resampling and mono conversion
  2. Transmission: Chunks flow through the shared audio_sender channel to the pipeline thread
  3. Buffering: AudioMixerRingBuffer::add_samples receives chunks and accumulates them in the appropriate queue
  4. Extraction: When can_mix() returns true, extract_window() produces synchronized 600ms windows
  5. Mixing: ProfessionalAudioMixer::mix_window combines the aligned audio
  6. Output: Mixed audio feeds the VAD processor for transcription and the recording writer for file storage

This architecture decouples capture timing from processing timing, allowing the microphone and system audio capture threads to operate at different cadences while guaranteeing synchronized output for the transcription engine.

Summary

  • Dual-queue buffering: AudioMixerRingBuffer maintains separate VecDeque<f32> buffers for microphone and system audio in pipeline.rs
  • Fixed window extraction: The algorithm aligns streams into 600ms windows, calculating sample counts based on the pipeline sample rate
  • Zero-padding synchronization: Missing data is padded with silence rather than waited for, preventing pipeline stalls and audio glitches
  • Soft clipping mixing: ProfessionalAudioMixer::mix_window combines streams with distortion protection for clean output
  • Safety mechanisms: max_buffer_size limits prevent memory exhaustion during network jitter or device failures

Frequently Asked Questions

How does Meetily handle when one audio stream lags behind the other?

Meetily's ring buffer mixing algorithm handles lag through its max_buffer_size safety limit and zero-padding strategy. When one buffer exceeds approximately 400ms while waiting for the other stream, the oldest samples are dropped with a logged warning. During mixing, the extract_window() method zero-pads any deficit in the lagging stream, ensuring continuous output without blocking the transcription pipeline.

What prevents audio glitches when system audio is temporarily unavailable?

The can_mix() method uses an OR condition to trigger mixing as soon as either stream fills a 600ms window, rather than waiting for both. When extract_window() encounters an empty or partial buffer, it resizes the vector with 0.0 values to fill the window completely. This prevents "last-sample-hold" artifacts that would otherwise create clicking sounds or DC offset shifts in the output.

Why does the mixer use zero-padding instead of waiting for both streams?

Zero-padding maintains real-time processing guarantees essential for live transcription. Waiting for both streams would introduce unpredictable latency spikes whenever one capture device buffers or drops packets. By padding with silence, Meetily ensures the Whisper transcription model receives consistent 600ms windows at regular intervals, preventing buffer underruns in the downstream audio processing chain.

Where is the ring buffer mixing algorithm implemented in the Meetily codebase?

The core implementation resides in frontend/src-tauri/src/audio/pipeline.rs, containing the AudioMixerRingBuffer struct and ProfessionalAudioMixer implementation. Supporting definitions for AudioChunk and DeviceType appear in recording_state.rs, while audio_capture.rs generates the input chunks and audio_processing.rs handles preliminary resampling and mono conversion before data reaches the ring buffer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →