# How Meetily Synchronizes Microphone and System Audio Streams During Capture

> Discover how Meetily synchronizes microphone and system audio streams with precise buffering, drift compensation, and sample-by-sample mixing. Learn the technical details.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: internals
- Published: 2026-07-30

---

**Meetily synchronizes independent microphone and system audio streams by buffering them in separate ring buffers, extracting aligned 50ms windows with zero-padding for drift compensation, and mixing them sample-by-sample at 48kHz.**

Meetily is an open-source meeting transcription tool that captures both local microphone input and system audio output simultaneously. Because these streams originate from different devices with independent clocks, the application must align them before transcription to prevent echo and drift. According to the Zackriya-Solutions/meetily source code, the application implements a **ring-buffer mixer** (`AudioMixerRingBuffer`) in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs) to coordinate these asynchronous inputs.

## Ring-Buffer Architecture for Dual-Stream Storage

The core synchronization primitive is the `AudioMixerRingBuffer` struct, which maintains separate queues for each audio source while enforcing size limits to prevent unbounded memory growth during network or device jitter.

### The AudioMixerRingBuffer Structure

```rust
// frontend/src-tauri/src/audio/pipeline.rs
struct AudioMixerRingBuffer {
    mic_buffer: VecDeque<f32>,
    system_buffer: VecDeque<f32>,
    window_size_samples: usize,   // e.g. 50ms of audio
    max_buffer_size: usize,       // safety limit (≈400ms)
}

```

Each incoming chunk from the microphone or system audio is appended to its dedicated `VecDeque`. The `window_size_samples` defines the granularity of synchronization (typically 50ms), while `max_buffer_size` acts as a safety valve against memory overflow.

## Asynchronous Sample Ingestion

Capture threads for each device operate independently, feeding raw PCM data into the ring buffer through the `add_samples` method. This design accepts data asynchronously without blocking either capture thread.

```rust
fn add_samples(&mut self, device_type: DeviceType, samples: Vec<f32>) {
    match device_type {
        DeviceType::Microphone => self.mic_buffer.extend(samples),
        DeviceType::System     => self.system_buffer.extend(samples),
    }
    // overflow handling & periodic diagnostics …
}

```

The `DeviceType` enum distinguishes between `Microphone` and `System` inputs, ensuring samples route to the correct queue while preserving chronological order within each stream.

## Window-Based Synchronization Logic

Synchronization occurs at the window level rather than sample-by-sample, allowing the pipeline to tolerate minor timing variations between devices.

### Triggering Mix with can_mix

The pipeline checks readiness using the `can_mix` method, which returns `true` when either buffer contains a full window:

```rust
fn can_mix(&self) -> bool {
    self.mic_buffer.len() >= self.window_size_samples ||
    self.system_buffer.len() >= self.window_size_samples
}

```

This logical OR operation ensures processing continues even if one stream lags temporarily, preventing the faster stream from blocking indefinitely.

### Extracting Aligned Windows with Zero-Padding

When mixing is triggered, `extract_window` drains samples from both buffers and aligns them into matching vectors:

```rust
fn extract_window(&mut self) -> Option<(Vec<f32>, Vec<f32>)> {
    // …if a buffer is too short, the remaining samples are drained
    // and the rest of the window is zero-padded (silence) …
}

```

If the microphone buffer contains fewer samples than required, the missing samples are replaced with zeros (silence). The same applies to the system audio buffer. This **zero-padding strategy** prevents audible artifacts like clicks or pops that would result from sample-holding or abrupt truncations.

## Sample-Level Mixing and Clipping

Once aligned windows are extracted, the `ProfessionalAudioMixer::mix_window` method combines them:

```rust
let mixed_clean = self.mixer.mix_window(&mic_window, &sys_window);

```

This implementation adds the two streams sample-by-sample at **48kHz**, applies headroom scaling, and performs soft clipping when the sum exceeds ±1.0. The result is a single, perfectly aligned frame ready for downstream voice activity detection (VAD) and Whisper transcription.

## The Pipeline Execution Loop

The synchronization logic executes within `AudioPipeline::run`, which coordinates capture threads and transcription:

```rust
while self.ring_buffer.can_mix() {
    if let Some((mic_window, sys_window)) = self.ring_buffer.extract_window() {
        let mixed = self.mixer.mix_window(&mic_window, &sys_window);
        // ► VAD → Whisper (transcription)
        // ► Recording sender (WAV file)
    }
}

```

This loop runs continuously, pulling mixed audio from the ring buffer and forwarding it to both the real-time transcription path and the optional file recording sink.

## End-to-End Data Flow

The complete synchronization path follows these stages:

1. **Capture** – `AudioCapture` reads raw PCM from devices, optionally resamples to 48kHz, and sends `AudioChunk` structs to the shared `RecordingState`.
2. **State forwarding** – `RecordingState` forwards chunks via an `mpsc::UnboundedSender` to the pipeline.
3. **Ring-buffer accumulation** – `AudioPipeline` stores samples in the respective `mic_buffer` or `system_buffer` queues.
4. **Windowed alignment** – When either buffer reaches the 50ms threshold, both buffers are drained and zero-padded if necessary to create time-aligned windows.
5. **Mixed output** – The mixed frame emits simultaneously to the VAD-driven transcription engine and the WAV file recording channel.

## Implementation Example

To initialize the synchronized pipeline in a Tauri application:

```rust
use frontend::src_tauri::audio::{
    AudioPipelineManager, RecordingState,
};
use tokio::sync::mpsc;
use std::sync::Arc;

// Create shared recording state
let state = Arc::new(RecordingState::new());

// Channels for transcription and file saving
let (trans_tx, _trans_rx) = mpsc::unbounded_channel();
let (record_tx, _record_rx) = mpsc::unbounded_channel();

// Initialize pipeline manager
let mut manager = AudioPipelineManager::new();
manager.start(
    state.clone(),
    trans_tx,
    0, // target chunk duration
    48_000, // sample rate
    Some(record_tx),
    "Built‑in Microphone".into(),
    InputDeviceKind::Wired,
    "BlackHole 2ch".into(), // macOS virtual audio device
    InputDeviceKind::Virtual,
)?;

```

The `AudioPipelineManager` internally instantiates `AudioMixerRingBuffer` and manages the capture threads, automatically handling synchronization without additional UI code.

## Summary

- Meetily uses a **ring-buffer architecture** (`AudioMixerRingBuffer`) with separate `VecDeque` queues for microphone and system audio to handle asynchronous capture threads.
- Synchronization occurs at **50ms window boundaries**, with the `can_mix` method triggering processing when either buffer fills, and `extract_window` aligning the data.
- **Zero-padding** compensates for clock drift and jitter by replacing missing samples with silence, preventing audio artifacts.
- The **ProfessionalAudioMixer** combines aligned windows with soft clipping to produce a single 48kHz stream for transcription.
- All synchronization logic resides in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs), integrated into the main processing loop at lines 667-785.

## Frequently Asked Questions

### How does Meetily handle clock drift between microphone and system audio?

Meetily tolerates clock drift through its **zero-padding strategy** in `AudioMixerRingBuffer::extract_window`. When one stream lags behind the other, the missing samples for that window are filled with zeros (silence) rather than waiting indefinitely. This ensures the mixed output maintains consistent timing without introducing gaps or clicks, as the faster stream continues processing aligned windows.

### What happens if one audio stream stops transmitting temporarily?

If a capture thread stops producing samples, the `max_buffer_size` limit (approximately 400ms) prevents memory overflow in the active buffer. When `extract_window` is called, the stalled stream's window is zero-padded to match the expected `window_size_samples`. The pipeline continues mixing and producing output, simply treating the missing stream as silence until data resumes.

### Why does Meetily use 50ms windows for synchronization?

The **50ms window size** balances latency and processing efficiency. Smaller windows reduce transcription delay but increase CPU overhead from mixing operations and thread synchronization. The 50ms duration (2400 samples at 48kHz) provides sufficient data for the downstream Voice Activity Detection (VAD) algorithm to make accurate decisions while maintaining imperceptible latency for real-time transcription.

### Where is the synchronization logic located in the Meetily source code?

The core synchronization implementation is in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs). Key components include the `AudioMixerRingBuffer` struct (lines 16-23), the `add_samples` method (lines 49-64), the mixing logic (lines 54-84), and the main processing loop (lines 667-785). Device-specific capture logic resides in [`frontend/src-tauri/src/audio/capture/microphone.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/capture/microphone.rs) and [`frontend/src-tauri/src/audio/capture/system.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/capture/system.rs).