How Meetily's Audio Pipeline Handles Simultaneous Microphone and System Audio Capture with Ducking
Meetily captures microphone and system audio through independent back-ends, synchronizes them via a dual ring buffer with 50 ms windows, and applies soft scaling ducking to prevent clipping while preserving voice clarity.
Meetily, an open-source meeting assistant from Zackriya-Solutions/meetily, records both microphone input and system audio output simultaneously to create comprehensive meeting transcripts. The Rust-based audio pipeline in frontend/src-tauri/src/audio/ employs a sophisticated ring-buffer synchronization mechanism and proportional scaling algorithm to balance speech and background audio without aggressive compression artifacts.
Dual Capture Architecture with Independent Back-Ends
The pipeline initiates two separate capture instances through the start_recording Tauri command defined in recording_commands.rs. Each stream operates independently to prevent cross-contamination between microphone input and system playback.
Microphone Capture
The microphone stream is handled by capture/microphone.rs, which captures raw PCM data from the selected input device. According to the Meetily source code, this back-end creates an AudioCapture instance with DeviceType::Microphone, applying device-specific sample rates and channel configurations before forwarding chunks to the shared pipeline.
System Audio Capture
System audio is captured via capture/system.rs, which implements platform-specific loopback mechanisms (ScreenCaptureKit on macOS, WASAPI on Windows). This stream uses DeviceType::System and similarly routes raw chunks to the mixer, ensuring that application audio and speaker output are preserved alongside voice input.
Ring Buffer Synchronization Strategy
At the core of the simultaneous capture system lies the AudioMixerRingBuffer struct defined in pipeline.rs. This component maintains two VecDeque<f32> queues—mic_buffer and system_buffer—that accumulate samples until a fixed mixing window is available.
The synchronization logic in AudioMixerRingBuffer::add_samples and AudioMixerRingBuffer::extract_window manages temporal alignment through the following mechanism:
- Window-based extraction: The pipeline waits until both buffers contain at least
window_size_samples(default approximately 50 ms) before mixing - Silence padding: If one stream lags due to device latency differences, the extractor pads the missing samples with silence to maintain window alignment
- Atomic handling: Each extraction yields a synchronized pair of floating-point vectors ready for the mixing stage
This approach ensures that microphone speech and system audio remain temporally aligned even when hardware capture latencies differ between devices.
Soft Scaling Ducking Algorithm
Rather than implementing aggressive traditional ducking (which creates audible "pumping" artifacts), Meetily employs a soft scaling approach through ProfessionalAudioMixer::mix_window in pipeline.rs.
The mixing algorithm processes each sample pair as follows:
// System audio currently passes at unity gain (placeholder for future 70% reduction)
let sys_scaled = sys * 1.0;
// Microphone remains at full level
let sum = mic + sys_scaled;
// Soft scaling prevents hard clipping
let mixed_sample = if sum.abs() > 1.0 {
sum / sum.abs() // Proportionally scale back to ±1.0 range
} else {
sum
};
Key characteristics of this ducking implementation:
- Proportional attenuation: When the summed signal exceeds the ±1.0 floating-point range, the entire mix is scaled back proportionally rather than hard-limiting individual channels
- Voice preservation: The microphone stream receives priority through the proportional calculation, ensuring speech clarity over background audio
- Distortion prevention: The
sum / sum.abs()operation maintains the relative balance between sources while eliminating digital clipping
Pre-Processing and Resampling Pipeline
Before reaching the ring buffer, each capture stream undergoes format normalization. The AudioCapture::new constructor in pipeline.rs initializes a persistent SincFixedIn resampler to convert all inputs to the pipeline's target 48 kHz sample rate.
Microphone-specific enhancements (applied in audio_processing.rs):
- Noise suppression: Real-time suppression of keyboard clicks and ambient room noise
- High-pass filtering: Removal of sub-80 Hz rumble and DC offset
- Loudness normalization: Automatic gain control to standardize input levels before mixing
System audio bypasses these enhancement stages to preserve the original fidelity of shared media and participant voices from remote calls.
End-to-End Recording Workflow
The complete simultaneous capture workflow operates as follows:
- Initialization: The
start_recordingcommand creates twoAudioCaptureinstances (microphone and system) with sharedrecording_stateand channel senders - Data ingestion: Each capture invokes
process_audio_datawhen audio callbacks fire, forwarding resampled mono chunks toAudioMixerRingBuffer::add_samples - Window extraction: When both buffers reach the 50 ms threshold,
extract_windowyields synchronized(mic_win, sys_win)pairs - Mixing:
ProfessionalAudioMixer::mix_windowcombines the streams using soft scaling ducking - Persistence: The mixed buffer routes to
RecordingSaverfor disk storage and simultaneously feeds the VAD-driven transcription pipeline
// Rust implementation of dual capture initialization
let mic_capture = AudioCapture::new(
mic_device,
recording_state.clone(),
mic_sample_rate,
mic_channels,
DeviceType::Microphone,
Some(recording_sender.clone()),
);
let sys_capture = AudioCapture::new(
sys_device,
recording_state.clone(),
sys_sample_rate,
sys_channels,
DeviceType::System,
Some(recording_sender),
);
Summary
- Meetily captures microphone and system audio through parallel
AudioCaptureinstances incapture/microphone.rsandcapture/system.rs - The
AudioMixerRingBufferinpipeline.rssynchronizes streams using dualVecDeque<f32>queues with 50 ms windows and silence padding for lag compensation - Soft scaling ducking in
ProfessionalAudioMixer::mix_windowprevents clipping through proportional attenuation (sum / sum.abs()) rather than hard limiting - Microphone audio receives exclusive noise suppression, high-pass filtering, and normalization through
audio_processing.rsbefore mixing - The architecture preserves voice clarity while preventing distortion, producing balanced recordings suitable for automatic transcription
Frequently Asked Questions
How does Meetily prevent audio clipping when mixing microphone and system audio?
The pipeline implements proportional scaling in ProfessionalAudioMixer::mix_window. When the sum of microphone and system samples exceeds the ±1.0 floating-point range, the algorithm divides by the absolute value of the sum (sum / sum.abs()). This scales the entire mixed signal back into range while preserving the relative balance between sources, avoiding the harsh distortion caused by hard clipping or zero-threshold limiting.
What happens if the microphone and system audio streams fall out of sync?
The AudioMixerRingBuffer::extract_window method detects latency differences through buffer depth monitoring. If one stream delivers samples slower than the other, the extractor pads the underflowing buffer with silence (zero-valued samples) until both queues contain sufficient data for the 50 ms extraction window. This zero-padding mechanism ensures temporal alignment without dropping samples or introducing pitch artifacts.
Why does Meetily use soft scaling instead of traditional audio ducking?
Traditional ducking applies aggressive compression or level reduction to background audio whenever voice is detected, creating audible "pumping" artifacts that fatigue listeners. Meetily's soft scaling approach only attenuates the mixed signal when necessary to prevent digital clipping, maintaining the natural dynamic relationship between speech and system audio. This preserves the intelligibility of quiet background conversations or subtle audio cues while protecting against overload.
Which audio processing steps are applied exclusively to microphone input?
According to the AudioCapture::new implementation in pipeline.rs, only microphone streams (identified by DeviceType::Microphone) are routed through the noise suppressor, high-pass filter, and loudness normalizer defined in audio_processing.rs. System audio bypasses these stages to maintain the original fidelity of computer playback, ensuring that music, videos, and remote participant audio retain their intended frequency characteristics and dynamic range.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →