# How Meetily Handles Simultaneous Microphone and System Audio Capture with Intelligent Mixing

> Discover how Meetily captures and intelligently mixes microphone and system audio simultaneously, creating professional recordings with advanced normalization and Whisper integration.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: how-to-guide
- Published: 2026-07-28

---

**Meetily captures microphone and system audio as independent streams, synchronizes them through a ring-buffer mechanism, and mixes them with soft-clipping normalization to create a professional-grade recording while routing speech-only segments to Whisper for transcription.**

Meetily, an open-source meeting transcription tool developed by Zackriya-Solutions, implements a sophisticated Rust-based audio pipeline to manage **simultaneous microphone and system audio capture**. The system processes dual audio streams in real-time within the Tauri backend, applying intelligent mixing algorithms that prevent hard clipping and maintain broadcast-level loudness standards. This architecture ensures that both the user's voice and system sounds are captured cleanly, synchronized precisely, and blended into a single coherent track saved as a WAV file.

## Three-Stage Pipeline Architecture

The core audio processing logic resides in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs), where the `AudioPipeline` orchestrates a three-stage workflow: device-specific capture with preprocessing, ring-buffer synchronization, and professional mixing with downstream routing.

### Stage 1: Device-Specific Capture and Preprocessing

Each audio stream originates from an `AudioCapture` instance that handles device-specific initialization and real-time processing. The microphone stream undergoes significant enhancement before entering the mixing pipeline, while system audio is processed with minimal modification.

The microphone signal passes through a **high-pass filter**, **RNNoise suppression** for noise reduction, and **EBU-R128 loudness normalization** targeting -23 LUFS (broadcast standard). Both streams are converted to mono and resampled to the pipeline's native **48 kHz** sample rate when necessary. The `AudioCapture::process_audio_data` method applies this enhancement chain and forwards the processed chunks as `AudioChunk` structures via `state.send_audio_chunk` to the shared pipeline state.

System audio capture bypasses the enhancement filters but undergoes the same format standardization, ensuring both streams share identical specifications before mixing.

### Stage 2: Ring-Buffer Synchronization

The `AudioPipeline` receives raw chunks from both devices through an unbounded Tokio channel. An `AudioMixerRingBuffer` stores microphone and system samples in two independent `VecDeque` buffers until each accumulates sufficient data for a fixed-size mixing window of approximately **50 milliseconds**.

The ring buffer implements defensive programming against stream drift and dropouts. When one stream temporarily lacks data, the `extract_window` method pads the missing window with silence (zeros) rather than repeating previous samples, eliminating audible artifacts. The buffer caps its `max_buffer_size` at **400 milliseconds** (eight times the mixing window) to absorb bursty system-audio delivery common in macOS Core Audio while preventing runaway memory growth.

Key methods include `add_samples` for ingestion, `can_mix` to verify window readiness, and `extract_window` for synchronized retrieval.

### Stage 3: Professional Mixing and Downstream Routing

Once the ring buffer reports a full window via `can_mix`, the `ProfessionalAudioMixer` combines the microphone and system buffers in `mix_window`. The mixer applies a **70% amplitude reduction** to system audio (scaling by 0.7) to prevent background sounds from overwhelming the primary microphone signal.

**Soft-clipping protection** ensures the mixed signal never exceeds the ±1.0 floating-point range. If the sum of microphone and scaled system samples exceeds this boundary, the mixer proportionally scales the result back to the unit range, preventing the harsh distortion associated with hard clipping while preserving relative dynamics.

After mixing, the pipeline applies **no additional gain**—the microphone signal has already been normalized to broadcast level and the system audio attenuated. The mixed chunk is then dispatched to two destinations:
- The **VAD processor** ([`vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/vad.rs)), which extracts speech-only segments for Whisper transcription
- The **recording sender**, which writes the final mixed audio to a WAV file containing both sources

This flow executes within the `AudioPipeline::run` loop, which coordinates the ring buffer, mixer, VAD, and recording components in sequence.

## Handling Sample-Rate Mismatches and Latency

### Dynamic Resampling with RMS Preservation

When input devices operate at sample rates other than 48 kHz, `AudioCapture` detects the mismatch and instantiates a persistent **Rubato resampler**. This resampler buffers variable-size input chunks until accumulating a 512-sample block for processing, preserving RMS energy across chunk boundaries to maintain consistent loudness levels during format conversion.

### Jitter Absorption for System Audio

The **400-millisecond maximum buffer size** specifically addresses latency variations in system audio capture. On platforms like macOS where Core Audio may deliver bursts of samples irregularly, the 8x safety margin (8 × 50ms window) ensures the mixer always has sufficient data to generate continuous output without blocking the microphone stream.

## Implementation Example

The following excerpt demonstrates pipeline initialization from the Tauri command layer:

```rust
// In frontend/src-tauri/src/audio/recording_manager.rs (simplified)
#[tauri::command]
async fn start_recording(
    app: tauri::AppHandle,
    mic_device_name: Option<String>,
    system_device_name: Option<String>,
    meeting_name: Option<String>,
) -> Result<(), String> {
    let state = RecordingState::new();
    let (trans_tx, trans_rx) = mpsc::unbounded_channel();
    let recording_sender = Some(mpsc::unbounded_channel().0);

    let mut manager = AudioPipelineManager::new();
    manager.start(
        Arc::new(state),
        trans_tx,
        20,           // target_chunk_duration_ms
        48_000,       // sample_rate
        recording_sender,
        mic_device_name.unwrap_or_default(),
        detect_mic_kind(),
        system_device_name.unwrap_or_default(),
        detect_system_kind(),
    )?;
    Ok(())
}

```

The mixing implementation applies soft-clipping logic as shown in this simplified version of `ProfessionalAudioMixer::mix_window`:

```rust
fn mix_window(&mut self, mic_window: &[f32], sys_window: &[f32]) -> Vec<f32> {
    let max_len = mic_window.len().max(sys_window.len());
    let mut mixed = Vec::with_capacity(max_len);
    for i in 0..max_len {
        let mic = mic_window.get(i).copied().unwrap_or(0.0);
        let sys = sys_window.get(i).copied().unwrap_or(0.0) * 0.7; // 70% system level
        
        let sum = mic + sys;
        // Soft clipping: proportional scaling if exceeding unit range
        let mixed_sample = if sum.abs() > 1.0 { 
            sum / sum.abs() 
        } else { 
            sum 
        };
        mixed.push(mixed_sample);
    }
    mixed
}

```

Ring-buffer extraction and mixing occur in the main pipeline loop:

```rust
// Inside AudioPipeline::run loop
if let Some((mic_win, sys_win)) = self.ring_buffer.extract_window() {
    let mixed = self.mixer.mix_window(&mic_win, &sys_win);
    // Forward `mixed` to VAD processor and recording sender...
}

```

## Summary

- **Dual-stream capture**: Independent `AudioCapture` instances process microphone and system audio with device-specific enhancement chains in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs).
- **Synchronized mixing**: An `AudioMixerRingBuffer` maintains 50ms aligned windows with 400ms maximum latency tolerance, padding missing data with silence to prevent artifacts.
- **Professional audio processing**: The `ProfessionalAudioMixer` applies 70% system audio attenuation and soft-clipping to prevent distortion while maintaining -23 LUFS broadcast loudness on the microphone channel.
- **Dual output routing**: Mixed audio writes to WAV files for recording, while VAD-filtered speech segments route separately to Whisper for transcription.
- **Robust resampling**: Rubato-based resamplers handle arbitrary input sample rates while preserving RMS energy across 512-sample processing blocks.

## Frequently Asked Questions

### How does Meetily prevent audio clipping when mixing two sources?

Meetily implements **soft-clipping normalization** in `ProfessionalAudioMixer::mix_window`. When the sum of microphone and scaled system audio exceeds the ±1.0 floating-point range, the mixer proportionally scales the entire sample back to the unit boundary rather than truncating it. This preserves the relative dynamics between sources while eliminating the harsh digital distortion associated with hard clipping.

### What sample rate does Meetily use for internal audio processing?

The pipeline operates at a native **48 kHz** sample rate in mono format. Both microphone and system audio streams are resampled to this standard using Rubato resamplers when source devices provide alternative rates (such as 44.1 kHz). The resamplers buffer variable-size chunks until accumulating 512-sample blocks to maintain processing efficiency and RMS energy consistency.

### How does the system handle temporary dropouts in one audio stream?

The `AudioMixerRingBuffer` detects missing data during `extract_window` operations and pads the deficient stream with **silence (zero-values)** rather than repeating previous samples or interpolating. This approach prevents audible artifacts like echoes or phasing effects when one device temporarily stops delivering data, ensuring the mixed output remains clean during brief interruptions.

### Where is the audio mixing logic implemented in the codebase?

The professional mixing implementation resides in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs), specifically within the `ProfessionalAudioMixer` struct and its `mix_window` method. The high-level orchestration that starts and stops this pipeline, including device selection and file output management, is handled by `AudioPipelineManager` in [`frontend/src-tauri/src/audio/recording_manager.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/recording_manager.rs). Voice activity detection that feeds the transcription engine operates in [`frontend/src-tauri/src/audio/vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/vad.rs).