How Meetily Synchronizes Microphone and System Audio Streams During Capture
Meetily synchronizes independent microphone and system audio streams by buffering them in separate ring buffers, extracting aligned 50ms windows with zero-padding for drift compensation, and mixing them sample-by-sample at 48kHz.
Meetily is an open-source meeting transcription tool that captures both local microphone input and system audio output simultaneously. Because these streams originate from different devices with independent clocks, the application must align them before transcription to prevent echo and drift. According to the Zackriya-Solutions/meetily source code, the application implements a ring-buffer mixer (AudioMixerRingBuffer) in frontend/src-tauri/src/audio/pipeline.rs to coordinate these asynchronous inputs.
Ring-Buffer Architecture for Dual-Stream Storage
The core synchronization primitive is the AudioMixerRingBuffer struct, which maintains separate queues for each audio source while enforcing size limits to prevent unbounded memory growth during network or device jitter.
The AudioMixerRingBuffer Structure
// frontend/src-tauri/src/audio/pipeline.rs
struct AudioMixerRingBuffer {
mic_buffer: VecDeque<f32>,
system_buffer: VecDeque<f32>,
window_size_samples: usize, // e.g. 50ms of audio
max_buffer_size: usize, // safety limit (≈400ms)
}
Each incoming chunk from the microphone or system audio is appended to its dedicated VecDeque. The window_size_samples defines the granularity of synchronization (typically 50ms), while max_buffer_size acts as a safety valve against memory overflow.
Asynchronous Sample Ingestion
Capture threads for each device operate independently, feeding raw PCM data into the ring buffer through the add_samples method. This design accepts data asynchronously without blocking either capture thread.
fn add_samples(&mut self, device_type: DeviceType, samples: Vec<f32>) {
match device_type {
DeviceType::Microphone => self.mic_buffer.extend(samples),
DeviceType::System => self.system_buffer.extend(samples),
}
// overflow handling & periodic diagnostics …
}
The DeviceType enum distinguishes between Microphone and System inputs, ensuring samples route to the correct queue while preserving chronological order within each stream.
Window-Based Synchronization Logic
Synchronization occurs at the window level rather than sample-by-sample, allowing the pipeline to tolerate minor timing variations between devices.
Triggering Mix with can_mix
The pipeline checks readiness using the can_mix method, which returns true when either buffer contains a full window:
fn can_mix(&self) -> bool {
self.mic_buffer.len() >= self.window_size_samples ||
self.system_buffer.len() >= self.window_size_samples
}
This logical OR operation ensures processing continues even if one stream lags temporarily, preventing the faster stream from blocking indefinitely.
Extracting Aligned Windows with Zero-Padding
When mixing is triggered, extract_window drains samples from both buffers and aligns them into matching vectors:
fn extract_window(&mut self) -> Option<(Vec<f32>, Vec<f32>)> {
// …if a buffer is too short, the remaining samples are drained
// and the rest of the window is zero-padded (silence) …
}
If the microphone buffer contains fewer samples than required, the missing samples are replaced with zeros (silence). The same applies to the system audio buffer. This zero-padding strategy prevents audible artifacts like clicks or pops that would result from sample-holding or abrupt truncations.
Sample-Level Mixing and Clipping
Once aligned windows are extracted, the ProfessionalAudioMixer::mix_window method combines them:
let mixed_clean = self.mixer.mix_window(&mic_window, &sys_window);
This implementation adds the two streams sample-by-sample at 48kHz, applies headroom scaling, and performs soft clipping when the sum exceeds ±1.0. The result is a single, perfectly aligned frame ready for downstream voice activity detection (VAD) and Whisper transcription.
The Pipeline Execution Loop
The synchronization logic executes within AudioPipeline::run, which coordinates capture threads and transcription:
while self.ring_buffer.can_mix() {
if let Some((mic_window, sys_window)) = self.ring_buffer.extract_window() {
let mixed = self.mixer.mix_window(&mic_window, &sys_window);
// ► VAD → Whisper (transcription)
// ► Recording sender (WAV file)
}
}
This loop runs continuously, pulling mixed audio from the ring buffer and forwarding it to both the real-time transcription path and the optional file recording sink.
End-to-End Data Flow
The complete synchronization path follows these stages:
- Capture –
AudioCapturereads raw PCM from devices, optionally resamples to 48kHz, and sendsAudioChunkstructs to the sharedRecordingState. - State forwarding –
RecordingStateforwards chunks via anmpsc::UnboundedSenderto the pipeline. - Ring-buffer accumulation –
AudioPipelinestores samples in the respectivemic_bufferorsystem_bufferqueues. - Windowed alignment – When either buffer reaches the 50ms threshold, both buffers are drained and zero-padded if necessary to create time-aligned windows.
- Mixed output – The mixed frame emits simultaneously to the VAD-driven transcription engine and the WAV file recording channel.
Implementation Example
To initialize the synchronized pipeline in a Tauri application:
use frontend::src_tauri::audio::{
AudioPipelineManager, RecordingState,
};
use tokio::sync::mpsc;
use std::sync::Arc;
// Create shared recording state
let state = Arc::new(RecordingState::new());
// Channels for transcription and file saving
let (trans_tx, _trans_rx) = mpsc::unbounded_channel();
let (record_tx, _record_rx) = mpsc::unbounded_channel();
// Initialize pipeline manager
let mut manager = AudioPipelineManager::new();
manager.start(
state.clone(),
trans_tx,
0, // target chunk duration
48_000, // sample rate
Some(record_tx),
"Built‑in Microphone".into(),
InputDeviceKind::Wired,
"BlackHole 2ch".into(), // macOS virtual audio device
InputDeviceKind::Virtual,
)?;
The AudioPipelineManager internally instantiates AudioMixerRingBuffer and manages the capture threads, automatically handling synchronization without additional UI code.
Summary
- Meetily uses a ring-buffer architecture (
AudioMixerRingBuffer) with separateVecDequequeues for microphone and system audio to handle asynchronous capture threads. - Synchronization occurs at 50ms window boundaries, with the
can_mixmethod triggering processing when either buffer fills, andextract_windowaligning the data. - Zero-padding compensates for clock drift and jitter by replacing missing samples with silence, preventing audio artifacts.
- The ProfessionalAudioMixer combines aligned windows with soft clipping to produce a single 48kHz stream for transcription.
- All synchronization logic resides in
frontend/src-tauri/src/audio/pipeline.rs, integrated into the main processing loop at lines 667-785.
Frequently Asked Questions
How does Meetily handle clock drift between microphone and system audio?
Meetily tolerates clock drift through its zero-padding strategy in AudioMixerRingBuffer::extract_window. When one stream lags behind the other, the missing samples for that window are filled with zeros (silence) rather than waiting indefinitely. This ensures the mixed output maintains consistent timing without introducing gaps or clicks, as the faster stream continues processing aligned windows.
What happens if one audio stream stops transmitting temporarily?
If a capture thread stops producing samples, the max_buffer_size limit (approximately 400ms) prevents memory overflow in the active buffer. When extract_window is called, the stalled stream's window is zero-padded to match the expected window_size_samples. The pipeline continues mixing and producing output, simply treating the missing stream as silence until data resumes.
Why does Meetily use 50ms windows for synchronization?
The 50ms window size balances latency and processing efficiency. Smaller windows reduce transcription delay but increase CPU overhead from mixing operations and thread synchronization. The 50ms duration (2400 samples at 48kHz) provides sufficient data for the downstream Voice Activity Detection (VAD) algorithm to make accurate decisions while maintaining imperceptible latency for real-time transcription.
Where is the synchronization logic located in the Meetily source code?
The core synchronization implementation is in frontend/src-tauri/src/audio/pipeline.rs. Key components include the AudioMixerRingBuffer struct (lines 16-23), the add_samples method (lines 49-64), the mixing logic (lines 54-84), and the main processing loop (lines 667-785). Device-specific capture logic resides in frontend/src-tauri/src/audio/capture/microphone.rs and frontend/src-tauri/src/audio/capture/system.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →