How Meetily Synchronizes Microphone and System Audio Streams During Capture

Meetily synchronizes independent microphone and system audio streams by buffering them in separate ring buffers, extracting aligned 50ms windows with zero-padding for drift compensation, and mixing them sample-by-sample at 48kHz.

Meetily is an open-source meeting transcription tool that captures both local microphone input and system audio output simultaneously. Because these streams originate from different devices with independent clocks, the application must align them before transcription to prevent echo and drift. According to the Zackriya-Solutions/meetily source code, the application implements a ring-buffer mixer (AudioMixerRingBuffer) in frontend/src-tauri/src/audio/pipeline.rs to coordinate these asynchronous inputs.

Ring-Buffer Architecture for Dual-Stream Storage

The core synchronization primitive is the AudioMixerRingBuffer struct, which maintains separate queues for each audio source while enforcing size limits to prevent unbounded memory growth during network or device jitter.

The AudioMixerRingBuffer Structure

// frontend/src-tauri/src/audio/pipeline.rs
struct AudioMixerRingBuffer {
    mic_buffer: VecDeque<f32>,
    system_buffer: VecDeque<f32>,
    window_size_samples: usize,   // e.g. 50ms of audio
    max_buffer_size: usize,       // safety limit (≈400ms)
}

Each incoming chunk from the microphone or system audio is appended to its dedicated VecDeque. The window_size_samples defines the granularity of synchronization (typically 50ms), while max_buffer_size acts as a safety valve against memory overflow.

Asynchronous Sample Ingestion

Capture threads for each device operate independently, feeding raw PCM data into the ring buffer through the add_samples method. This design accepts data asynchronously without blocking either capture thread.

fn add_samples(&mut self, device_type: DeviceType, samples: Vec<f32>) {
    match device_type {
        DeviceType::Microphone => self.mic_buffer.extend(samples),
        DeviceType::System     => self.system_buffer.extend(samples),
    }
    // overflow handling & periodic diagnostics …
}

The DeviceType enum distinguishes between Microphone and System inputs, ensuring samples route to the correct queue while preserving chronological order within each stream.

Window-Based Synchronization Logic

Synchronization occurs at the window level rather than sample-by-sample, allowing the pipeline to tolerate minor timing variations between devices.

Triggering Mix with can_mix

The pipeline checks readiness using the can_mix method, which returns true when either buffer contains a full window:

fn can_mix(&self) -> bool {
    self.mic_buffer.len() >= self.window_size_samples ||
    self.system_buffer.len() >= self.window_size_samples
}

This logical OR operation ensures processing continues even if one stream lags temporarily, preventing the faster stream from blocking indefinitely.

Extracting Aligned Windows with Zero-Padding

When mixing is triggered, extract_window drains samples from both buffers and aligns them into matching vectors:

fn extract_window(&mut self) -> Option<(Vec<f32>, Vec<f32>)> {
    // …if a buffer is too short, the remaining samples are drained
    // and the rest of the window is zero-padded (silence) …
}

If the microphone buffer contains fewer samples than required, the missing samples are replaced with zeros (silence). The same applies to the system audio buffer. This zero-padding strategy prevents audible artifacts like clicks or pops that would result from sample-holding or abrupt truncations.

Sample-Level Mixing and Clipping

Once aligned windows are extracted, the ProfessionalAudioMixer::mix_window method combines them:

let mixed_clean = self.mixer.mix_window(&mic_window, &sys_window);

This implementation adds the two streams sample-by-sample at 48kHz, applies headroom scaling, and performs soft clipping when the sum exceeds ±1.0. The result is a single, perfectly aligned frame ready for downstream voice activity detection (VAD) and Whisper transcription.

The Pipeline Execution Loop

The synchronization logic executes within AudioPipeline::run, which coordinates capture threads and transcription:

while self.ring_buffer.can_mix() {
    if let Some((mic_window, sys_window)) = self.ring_buffer.extract_window() {
        let mixed = self.mixer.mix_window(&mic_window, &sys_window);
        // ► VAD → Whisper (transcription)
        // ► Recording sender (WAV file)
    }
}

This loop runs continuously, pulling mixed audio from the ring buffer and forwarding it to both the real-time transcription path and the optional file recording sink.

End-to-End Data Flow

The complete synchronization path follows these stages:

  1. Capture – AudioCapture reads raw PCM from devices, optionally resamples to 48kHz, and sends AudioChunk structs to the shared RecordingState.
  2. State forwarding – RecordingState forwards chunks via an mpsc::UnboundedSender to the pipeline.
  3. Ring-buffer accumulation – AudioPipeline stores samples in the respective mic_buffer or system_buffer queues.
  4. Windowed alignment – When either buffer reaches the 50ms threshold, both buffers are drained and zero-padded if necessary to create time-aligned windows.
  5. Mixed output – The mixed frame emits simultaneously to the VAD-driven transcription engine and the WAV file recording channel.

Implementation Example

To initialize the synchronized pipeline in a Tauri application:

use frontend::src_tauri::audio::{
    AudioPipelineManager, RecordingState,
};
use tokio::sync::mpsc;
use std::sync::Arc;

// Create shared recording state
let state = Arc::new(RecordingState::new());

// Channels for transcription and file saving
let (trans_tx, _trans_rx) = mpsc::unbounded_channel();
let (record_tx, _record_rx) = mpsc::unbounded_channel();

// Initialize pipeline manager
let mut manager = AudioPipelineManager::new();
manager.start(
    state.clone(),
    trans_tx,
    0, // target chunk duration
    48_000, // sample rate
    Some(record_tx),
    "Built‑in Microphone".into(),
    InputDeviceKind::Wired,
    "BlackHole 2ch".into(), // macOS virtual audio device
    InputDeviceKind::Virtual,
)?;

The AudioPipelineManager internally instantiates AudioMixerRingBuffer and manages the capture threads, automatically handling synchronization without additional UI code.

Summary

  • Meetily uses a ring-buffer architecture (AudioMixerRingBuffer) with separate VecDeque queues for microphone and system audio to handle asynchronous capture threads.
  • Synchronization occurs at 50ms window boundaries, with the can_mix method triggering processing when either buffer fills, and extract_window aligning the data.
  • Zero-padding compensates for clock drift and jitter by replacing missing samples with silence, preventing audio artifacts.
  • The ProfessionalAudioMixer combines aligned windows with soft clipping to produce a single 48kHz stream for transcription.
  • All synchronization logic resides in frontend/src-tauri/src/audio/pipeline.rs, integrated into the main processing loop at lines 667-785.

Frequently Asked Questions

How does Meetily handle clock drift between microphone and system audio?

Meetily tolerates clock drift through its zero-padding strategy in AudioMixerRingBuffer::extract_window. When one stream lags behind the other, the missing samples for that window are filled with zeros (silence) rather than waiting indefinitely. This ensures the mixed output maintains consistent timing without introducing gaps or clicks, as the faster stream continues processing aligned windows.

What happens if one audio stream stops transmitting temporarily?

If a capture thread stops producing samples, the max_buffer_size limit (approximately 400ms) prevents memory overflow in the active buffer. When extract_window is called, the stalled stream's window is zero-padded to match the expected window_size_samples. The pipeline continues mixing and producing output, simply treating the missing stream as silence until data resumes.

Why does Meetily use 50ms windows for synchronization?

The 50ms window size balances latency and processing efficiency. Smaller windows reduce transcription delay but increase CPU overhead from mixing operations and thread synchronization. The 50ms duration (2400 samples at 48kHz) provides sufficient data for the downstream Voice Activity Detection (VAD) algorithm to make accurate decisions while maintaining imperceptible latency for real-time transcription.

Where is the synchronization logic located in the Meetily source code?

The core synchronization implementation is in frontend/src-tauri/src/audio/pipeline.rs. Key components include the AudioMixerRingBuffer struct (lines 16-23), the add_samples method (lines 49-64), the mixing logic (lines 54-84), and the main processing loop (lines 667-785). Device-specific capture logic resides in frontend/src-tauri/src/audio/capture/microphone.rs and frontend/src-tauri/src/audio/capture/system.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →