How Meetily's Ring Buffer Synchronizes Microphone and System Audio Streams
Meetily uses a dual-channel ring buffer implementation called AudioMixerRingBuffer to time-align independent microphone and system audio streams into a single synchronized track for recording and transcription.
The open-source Meetily application (Zackriya-Solutions/meetily) captures microphone input and system output on separate threads, which naturally produces asynchronous data arrival. The ring buffer decouples capture timing from processing, ensuring both streams mix cleanly regardless of jitter or latency differences.
Core Architecture of AudioMixerRingBuffer
The synchronization logic resides in frontend/src-tauri/src/audio/pipeline.rs, where the AudioMixerRingBuffer struct maintains two independent VecDeque<f32> buffers—one for microphone samples and one for system audio.
Buffer Initialization and Capacity Planning
When instantiated via AudioMixerRingBuffer::new(sample_rate), the buffer calculates two critical thresholds:
- Fixed mixing window: Default approximately 50ms of audio data
- Maximum buffer size: Approximately 400ms to absorb temporary bursts or pauses
// Inside AudioMixerRingBuffer::new()
let mixing_window_samples = (sample_rate as f32 * 0.05) as usize; // 50ms
let max_buffer_size = (sample_rate as f32 * 0.4) as usize; // 400ms
These values prevent overflow while maintaining low latency, allowing the buffer to tolerate significant jitter without dropping data.
Dual-Stream Ingestion Strategy
Incoming audio chunks arrive via add_samples(device_type, samples), which appends data to the appropriate buffer and trims overflow beyond the maximum size. The function emits warnings when truncation occurs, indicating potential data loss from extreme latency spikes.
// Feeding audio into the ring buffer
self.ring_buffer.add_samples(chunk.device_type.clone(), chunk.data);
Each AudioChunk contains a DeviceType discriminator (Microphone or System) that routes samples to the correct internal queue.
The Synchronization Mechanism
The ring buffer employs a time-aligned staging strategy that guarantees both streams occupy identical window lengths during mixing, regardless of individual capture rates.
Mix-Readiness Detection
The can_mix() method returns true when either buffer contains at least one full mixing window. This permissive threshold allows processing to begin as soon as one stream is ready, while the other continues accumulating data.
// Inside AudioPipeline::run
while self.ring_buffer.can_mix() {
// Process synchronized windows
}
This approach prevents the faster stream from stalling while waiting for the slower one, reducing end-to-end latency.
Window Extraction and Zero-Padding
The extract_window() method pulls synchronized slices from both buffers. When a buffer lacks sufficient samples, the function zero-pads the missing frames with silence rather than delaying processing or reusing stale data.
if let Some((mic_win, sys_win)) = self.ring_buffer.extract_window() {
let mixed = self.mixer.mix_window(&mic_win, &sys_win);
// mixed contains time-aligned audio from both sources
}
Zero-padding maintains consistent window sizes and prevents audio artifacts like repeated samples or phase misalignment. The extracted windows always match the 50ms mixing window duration, ensuring sample-accurate synchronization.
Audio Mixing and Pipeline Integration
Once synchronized windows are extracted, the pipeline blends them into a single track while preventing distortion.
ProfessionalAudioMixer Implementation
The ProfessionalAudioMixer::mix_window function performs additive mixing with soft scaling. When the combined sum exceeds ±1.0, the mixer scales the entire window down proportionally rather than hard-clipping individual samples, preserving dynamic range and avoiding digital distortion.
Continuous Processing Loop
Inside AudioPipeline::run, a tight loop repeatedly checks can_mix(), extracts windows, and blends them until buffer levels drop below the mixing threshold. This design ensures tight synchronization across varying arrival rates:
while self.ring_buffer.can_mix() {
if let Some((mic_win, sys_win)) = self.ring_buffer.extract_window() {
let mixed = self.mixer.mix_window(&mic_win, &sys_win);
// Send to VAD for transcription and WAV recorder for file output
}
}
The mixed output feeds simultaneously into the Voice Activity Detection (VAD) processor for real-time transcription and the WAV recorder for persistent storage.
Summary
- Dual-buffer design: Separate
VecDeque<f32>instances for microphone and system audio decouple capture threads from processing. - Jitter tolerance: 400ms maximum buffer capacity absorbs temporary latency spikes without data loss.
- Zero-padding strategy: Missing samples are filled with silence to maintain time alignment rather than causing artifacts.
- Soft scaling: The mixer prevents clipping through proportional scaling rather than hard limiting.
- Continuous drainage: The processing loop consumes buffers as fast as data arrives, minimizing latency while ensuring synchronization.
Frequently Asked Questions
What happens when one audio stream lags behind the other?
The ring buffer continues processing using zero-padding for the lagging stream. When extract_window() detects insufficient samples in one buffer, it fills the remainder of the 50ms window with silence values. This allows the leading stream to progress without stalling, maintaining real-time performance while the slower stream catches up.
How does the ring buffer prevent audio clipping during mixing?
The ProfessionalAudioMixer implements soft scaling within mix_window(). If the additive combination of microphone and system samples exceeds the ±1.0 floating-point range, the mixer calculates a scaling factor to reduce the entire window proportionally. This preserves the relative dynamics of both sources while eliminating digital distortion that would occur from hard clipping.
What is the maximum latency the buffer can tolerate?
The buffer tolerates approximately 400ms of latency per stream before triggering overflow warnings. This max_buffer_size value (calculated as sample_rate * 0.4) accommodates significant jitter from system scheduling or audio driver variability. Beyond this threshold, add_samples() truncates excess data and logs warnings to indicate potential synchronization degradation.
How are the microphone and system audio streams captured separately?
Meetily uses independent capture threads in audio_capture.rs that instantiate separate AudioCapture instances for each device type. Each thread preprocesses raw PCM data into AudioChunk structures tagged with DeviceType::Microphone or DeviceType::System. These chunks flow through distinct channels into the pipeline, where the ring buffer recombines them into synchronized windows for mixing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →