How Meetily's Audio Pipeline Handles Synchronized Mixing of Microphone and System Audio

Meetily captures microphone and system audio simultaneously through a Rust-based pipeline that synchronizes streams with a ring buffer, applies professional soft-clipping mixing, and routes the combined output to both file recording and speech transcription.

Meetily's audio architecture solves the complex problem of real-time synchronized mixing for meeting recordings. The system must capture two independent audio sources—your voice and your computer's output—blend them seamlessly, and feed clean results to both storage and OpenAI Whisper transcription. The core implementation resides in frontend/src-tauri/src/audio/pipeline.rs as part of Meetily's Tauri-based Rust audio subsystem.

Three-Stage Pipeline Architecture

The audio processing flow divides cleanly into capture, synchronization, and mixing stages.

1. Device-Specific Capture and Preprocessing

Each audio stream originates from an AudioCapture instance. The microphone path undergoes substantial enhancement while system audio passes through with minimal modification.

Microphone processing chain:

  • High-pass filter (removes rumble)
  • RNNoise neural noise suppression
  • EBU-R128 loudness normalization to -23 LUFS

Both streams convert to mono and resample to the pipeline's native 48 kHz when necessary. The processed chunks forward as AudioChunk structs to the shared pipeline via state.send_audio_chunk.

// Simplified excerpt from AudioCapture::process_audio_data
fn process_audio_data(&mut self, raw_samples: Vec<f32>) {
    // Enhancement chain for microphone only
    let processed = if self.device_kind == DeviceKind::Microphone {
        self.apply_noise_suppression(
            self.apply_highpass(
                self.normalize_loudness(raw_samples)
            )
        )
    } else {
        raw_samples // System audio: no enhancement
    };
    
    let chunk = AudioChunk {
        samples: processed,
        timestamp: Instant::now(),
        sample_rate: self.target_sample_rate,
    };
    self.state.send_audio_chunk(chunk);
}

The implementation spans lines 86-118 in pipeline.rs, where AudioCapture::process_audio_data orchestrates this chain.

2. Ring-Buffer Synchronization

The AudioPipeline receives raw chunks through an unbounded Tokio channel. An AudioMixerRingBuffer manages timing alignment using two independent VecDeque buffers—one per audio source.

// Core ring buffer structure
pub struct AudioMixerRingBuffer {
    mic_buffer: VecDeque<f32>,
    sys_buffer: VecDeque<f32>,
    window_size: usize,      // ~50ms at 48kHz = 2400 samples
    max_buffer_size: usize,  // 400ms cap for jitter absorption
}

The buffer implements three critical operations:

  • add_samples – Appends incoming chunks to the appropriate deque
  • can_mix – Returns true when both buffers contain sufficient samples for a full window
  • extract_window – Removes and returns synchronized Vec<f32> windows, padding with silence if one stream lags

The 400 ms maximum buffer size (8× the 50 ms mixing window) specifically absorbs bursty system audio delivery common on macOS Core Audio. When data is missing, the extractor pads with zeros rather than repeating samples—eliminating click artifacts.

3. Professional Mixing and Downstream Routing

Once synchronized, ProfessionalAudioMixer::mix_window combines the streams:

fn mix_window(&mut self, mic_window: &[f32], sys_window: &[f32]) -> Vec<f32> {
    let max_len = mic_window.len().max(sys_window.len());
    let mut mixed = Vec::with_capacity(max_len);
    
    for i in 0..max_len {
        let mic = mic_window.get(i).copied().unwrap_or(0.0);
        let sys = sys_window.get(i).copied().unwrap_or(0.0);
        
        // System audio attenuated to 70%, mic at full level
        let sum = mic + sys * 0.7;
        
        // Soft clipping: proportional scaling prevents hard distortion
        let mixed_sample = if sum.abs() > 1.0 {
            sum / sum.abs()  // Scales to ±1.0 boundary
        } else {
            sum
        };
        mixed.push(mixed_sample);
    }
    mixed
}

No additional gain applies post-mix. The microphone already carries broadcast-normalized loudness; system audio carries intentional attenuation. The soft-clipping algorithm preserves fidelity by scaling proportionally rather than truncating—avoiding the "radio-break" distortion of hard clipping.

The AudioPipeline::run loop orchestrates downstream routing:

// From pipeline.rs run loop (simplified)
async fn run(&mut self) {
    while let Some(chunk) = self.audio_rx.recv().await {
        self.ring_buffer.add_samples(chunk);
        
        while self.ring_buffer.can_mix() {
            let (mic_win, sys_win) = self.ring_buffer.extract_window();
            let mixed = self.mixer.mix_window(&mic_win, &sys_win);
            
            // Parallel downstream paths
            self.vad.process_for_transcription(&mixed);  // Speech → Whisper
            self.recording_sender.send(mixed);           // Full mix → WAV
        }
    }
}

This creates clean separation: transcription receives voice-only segments post-VAD, while recordings preserve the complete microphone-plus-system blend.

Handling Real-World Audio Challenges

Sample-Rate Mismatches

When capture devices report rates other than 48 kHz, AudioCapture instantiates a persistent Rubato resampler. The implementation buffers variable-size chunks until accumulating 512 samples—Rubato's preferred block size—then processes while preserving RMS energy across chunk boundaries.

Latency and Jitter Compensation

Mechanism Implementation Purpose
50 ms mixing window Fixed-size extraction Synchronization granularity
400 ms max buffer max_buffer_size cap Absorbs delivery variance
Silence padding unwrap_or(0.0) in extraction Prevents repetition artifacts

Clipping Avoidance

The mix_window soft-scaling algorithm guarantees output stays within ±1.0 regardless of input levels. This is essential because:

  • Microphone normalization targets -23 LUFS
  • System audio may contain peaks near 0 dBFS
  • Summation without scaling would clip frequently

Pipeline Initialization from Tauri Commands

The mixing pipeline launches from TypeScript-callable Rust commands:

#[tauri::command]
async fn start_recording(
    app: tauri::AppHandle,
    mic_device_name: Option<String>,
    system_device_name: Option<String>,
) -> Result<(), String> {
    let state = Arc::new(RecordingState::new());
    let (trans_tx, _) = mpsc::unbounded_channel();
    let (rec_tx, _) = mpsc::unbounded_channel();
    
    let mut manager = AudioPipelineManager::new();
    manager.start(
        state,
        trans_tx,
        20,           // Target chunk duration ms
        48_000,       // Unified sample rate
        Some(rec_tx),
        mic_device_name.unwrap_or_default(),
        DeviceKind::Microphone,
        system_device_name.unwrap_or_default(),
        DeviceKind::SystemOutput,
    )?;
    Ok(())
}

The AudioPipelineManager in recording_manager.rs handles lifecycle and device enumeration from the devices/ directory (WASAPI, CoreAudio, ALSA implementations).

Summary

  • Dual capture: AudioCapture instances handle microphone and system audio with device-appropriate processing chains
  • Synchronization: AudioMixerRingBuffer aligns timing with 50 ms windows and 400 ms jitter absorption
  • Professional mixing: ProfessionalAudioMixer::mix_window applies 70% system attenuation and soft clipping
  • Clean routing: Mixed audio writes to WAV files while VAD-filtered speech feeds Whisper transcription
  • Robust handling: Rubato resampling, silence padding, and proportional clipping prevent common audio pipeline failures

Frequently Asked Questions

How does Meetily prevent audio desynchronization between microphone and system sources?

The AudioMixerRingBuffer holds both streams in independent VecDeque buffers until sufficient samples accumulate for a complete 50 ms window. If one source delivers faster, the slower stream accumulates until synchronized extraction becomes possible. Missing data pads with silence rather than stretching or repeating samples.

Why does Meetily reduce system audio to 70% during mixing?

Full-level system audio would frequently trigger the soft-clipping threshold when combined with normalized microphone input. The 70% attenuation preserves headroom for typical content while maintaining audibility. This ratio is hardcoded in ProfessionalAudioMixer::mix_window and could expose configuration in future versions.

What happens when system audio temporarily stops (mute, no playback)?

The ring buffer's extract_window method returns zeros via unwrap_or(0.0) when deque elements are exhausted. This silence padding prevents audible artifacts that would occur from sample repetition or DC offset introduction. The microphone stream continues unaffected through the mixing process.

Does Meetily's audio pipeline support sample rates other than 48 kHz?

Yes—AudioCapture detects mismatches and creates persistent Rubato resamplers. The implementation buffers incoming chunks until reaching 512-sample blocks suitable for processing, preserving energy characteristics across resampling boundaries. Output always normalizes to 48 kHz for downstream consistency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →