# How Meetily Synchronizes Microphone and System Audio in AudioMixerRingBuffer

> Learn how Meetily synchronizes microphone and system audio using a dual ring buffer. Discover buffering, window extraction, and zero-padding techniques for perfect time alignment.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: internals
- Published: 2026-07-30

---

**Meetily aligns microphone and system audio by buffering both streams in a dual `VecDeque` ring buffer, extracting fixed 50 ms windows, and zero-padding any shortfall to maintain time alignment before mixing.**

Capturing desktop audio alongside microphone input creates a classic synchronization problem: the two streams arrive on independent threads with unpredictable jitter. Meetily solves this in its Rust audio pipeline using a specialized `AudioMixerRingBuffer` that marshals asynchronous inputs into perfectly aligned windows for transcription and recording.

## Ring Buffer Architecture

The synchronization engine is defined in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs) as the `AudioMixerRingBuffer` struct. It maintains two independent FIFO queues to decouple capture threads from the mixing stage:

```rust
struct AudioMixerRingBuffer {
    mic_buffer: VecDeque<f32>,
    system_buffer: VecDeque<f32>,
    window_size_samples: usize,   // 50 ms of audio at pipeline sample rate
    max_buffer_size: usize,       // Safety cap (≈ 400 ms)
}

```

Each buffer holds raw **f32** PCM samples. The constructor **[AudioMixerRingBuffer::new](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs#L25-L47)** calculates `window_size_samples` based on the pipeline’s sample rate—typically 48 kHz resulting in a 50 ms window—while `max_buffer_size` prevents unbounded memory growth during stalls.

## Ingesting Asynchronous Audio Streams

When an `AudioChunk` arrives from either the microphone or system capture thread, the pipeline calls **[add_samples](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs#L49-L85)**:

```rust
self.ring_buffer.add_samples(chunk.device_type.clone(), chunk.data);

```

This method appends the incoming slice to the appropriate `VecDeque`, logs buffer health metrics, and enforces the safety limit. If a buffer exceeds `max_buffer_size`, the oldest samples are dropped with a warning, prioritizing recent audio over historical data to minimize latency.

## Permissive Synchronization Logic

Unlike strict synchronizers that wait for both sources, Meetily uses a permissive readiness check implemented in **[can_mix](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs#L87-L90)**:

```rust
fn can_mix(&self) -> bool {
    self.mic_buffer.len() >= self.window_size_samples ||
    self.system_buffer.len() >= self.window_size_samples
}

```

This approach ensures the pipeline never stalls because one stream is lagging. As soon as either buffer contains a full 50 ms window, the system proceeds to extraction, preventing jitter in one capture from freezing the transcription thread.

## Extracting Time-Aligned Windows

When `can_mix()` returns true, the pipeline repeatedly calls **[extract_window](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs#L92-L141)** to produce synchronized pairs:

```rust
if let Some((mic_window, sys_window)) = self.ring_buffer.extract_window() {
    // Both vectors have exactly window_size_samples elements
    let mixed = self.mixer.mix_window(&mic_window, &sys_window);
}

```

The method drains **exactly** `window_size_samples` from each deque. If a source has fewer samples available—due to jitter, dropouts, or asymmetric startup delays—it drains the available data and pads the remainder with zeros (silence). Zero-padding is intentionally chosen over sample-holding to avoid repeating audio artifacts and to preserve the absolute timeline.

The result is two aligned `Vec<f32>` instances representing the same 50 ms of real-time audio, regardless of network jitter or thread scheduling variance between the capture sources.

## Downstream Mixing and Processing

Once aligned, the windows are handed to `ProfessionalAudioMixer::mix_window`, which sums the streams (system audio is pre-scaled to 70 % headroom). Inside **[AudioPipeline::run](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs#L22-L33)**, the mixed output is distributed to:

- **VAD → Whisper**: The Voice Activity Detection module receives the time-aligned mix to generate accurate speech segments for transcription.
- **Recording Saver**: The [`recording_saver.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/recording_saver.rs) module writes the synchronized PCM to the final WAV export, ensuring that user voice and desktop audio never drift out of sync in the saved file.

## Summary

- **AudioMixerRingBuffer** uses dual `VecDeque<f32>` queues in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs) to decouple capture threads.
- **Permissive gating** via `can_mix()` allows the pipeline to proceed as soon as one buffer is ready, preventing deadlocks.
- **Zero-padding** during `extract_window()` guarantees that every output frame has identical length and timeline, even when one source drops samples.
- The **50 ms window size** provides a deterministic frame rate for the VAD and transcription engines while maintaining low latency.

## Frequently Asked Questions

### What happens if the microphone or system audio temporarily drops?

The `extract_window` method drains whatever samples are currently available in the depleted buffer and pads the remainder with zeros (silence) to reach `window_size_samples`. This **zero-padding** strategy preserves the absolute timeline and prevents one slow stream from skewing the mix relative to the other.

### How does the ring buffer handle different sample rates between devices?

The pipeline resamples both inputs to a common rate—typically 48 kHz—before they enter the `AudioMixerRingBuffer`. Because both queues store `f32` PCM samples at the same clock rate, the synchronization logic operates on sample-aligned indices regardless of the original capture hardware.

### Why does Meetily use a permissive `can_mix` logic instead of waiting for both buffers?

Waiting for both buffers to fill would make the pipeline vulnerable to jitter in either capture thread. By returning true when **either** buffer reaches the threshold, `can_mix()` ensures that real-time transcription and recording continue uninterrupted; the lagging side simply receives silence padding until it catches up.

### Where is the main synchronization implementation located?

The `AudioMixerRingBuffer` struct and its synchronization methods—including `new()`, `add_samples()`, `can_mix()`, and `extract_window()`—are implemented in **[`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs)**, specifically between lines 25 and 141.