# How FFmpegMixer in Meetily Combines Multiple Audio Sources

> Discover how Meetily's FFmpegMixer combines audio sources using adaptive buffering and timestamp sync for clear transcriptions and quality recordings.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: internals
- Published: 2026-07-30

---

**Meetily's FFmpegMixer merges microphone and system audio into a single synchronized stream using per-source adaptive buffering, timestamp-aware synchronization, and RMS-based ducking to ensure clear voice transcription and high-quality recordings.**

Meetily, an open-source meeting assistant from Zackriya-Solutions, handles real-time audio processing by merging multiple input streams into a unified track suitable for both local recording and Whisper transcription. The core of this capability resides in the **FFmpegMixer** ([`frontend/src-tauri/src/audio/ffmpeg_mixer.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/ffmpeg_mixer.rs)), a Rust implementation that adapts FFmpeg's "Cap" buffering strategy for modern device environments. This module processes raw PCM frames from disparate sources—wired microphones, Bluetooth headsets, and system audio—while compensating for network jitter and hardware latency variations.

## Per-Source Adaptive Buffering Architecture

### The SourceBuffer Structure

Each audio source receives its own isolated **SourceBuffer** containing a queue of **TimestampedChunk** objects. This design eliminates the contention issues common in shared ring buffers by maintaining separate `VecDeque` storage for every input stream.

The buffer implementation resides in the `SourceBuffer` struct within [`ffmpeg_mixer.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/ffmpeg_mixer.rs), where the `chunks: VecDeque<TimestampedChunk>` field stores incoming audio frames with precise arrival timestamps.

### Device-Aware Timeout Configuration

Timeout values derive dynamically from **InputDeviceKind** classifications defined in [`device_detection.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/device_detection.rs). Wired devices utilize short 20–50 ms windows for minimal latency, while Bluetooth devices employ extended 80–200 ms windows to accommodate wireless jitter. The `SourceBuffer::new` constructor applies these thresholds via `device_kind.buffer_timeout()`.

## Synchronization and Gap Handling

### Timestamp-Aware Mixing

The mixer timestamps each chunk upon arrival using `TimestampedChunk::new`. Data extraction only occurs once chunk age exceeds the adaptive timeout, ensuring temporal alignment between streams before mixing. The `SourceBuffer::has_data` method checks readiness, while `FFmpegAudioMixer::has_data_ready` confirms both buffers meet synchronization requirements.

### Gap Detection and Silence Insertion

When elapsed time between consecutive chunks exceeds twice the expected duration, `SourceBuffer::push` logs a gap—expected behavior for Bluetooth connections but triggering warnings for wired devices. During buffer underruns, `SourceBuffer::pop_samples` injects calculated silence to maintain output continuity and prevent stream desynchronization.

## Professional Audio Processing Pipeline

### RMS-Based Adaptive Ducking

The `AudioMixer::mix` implementation calculates microphone RMS levels in real-time using `calculate_rms`. When speech detection triggers (`mic_rms > 0.01`), system audio automatically ducks to **60%** volume, preventing background audio from obscuring voice content while preserving ambient presence.

### Clipping Protection and Fixed Latency

Every mixed sample undergoes hard clamping to `[-1.0, 1.0]` via `mixed.clamp(-1.0, 1.0)`, guaranteeing distortion-free WAV output regardless of input gain levels. The system maintains a fixed **50 ms mixing window** (`mixing_window_samples`) across all operations, matching the original Cap implementation for predictable latency characteristics.

## Implementation Workflow

The FFmpegMixer operates through a structured five-stage pipeline:

1. **Initialization** – `FFmpegAudioMixer::new` instantiates two `SourceBuffer` instances (microphone and system) with device-specific timeouts and initializes an `AudioMixer` with adaptive ducking enabled.

2. **Data Ingestion** – Capture threads feed raw PCM frames through `push_mic(samples)` and `push_system(samples)` as 32-bit floating-point vectors.

3. **Readiness Verification** – `has_data_ready()` returns `true` only when both buffers contain chunks older than their respective adaptive timeouts, ensuring synchronous extraction.

4. **Mixing Execution** – `pop_mixed()` extracts 50 ms windows from each buffer via `pop_samples`, processes them through `AudioMixer::mix` with RMS-based ducking, increments the internal window counter, and returns the mixed frame.

5. **Diagnostics Logging** – Every 200 windows (approximately 10 seconds), `FFmpegAudioMixer::log_stats` outputs buffer latency, gap frequencies, and silence insertion metrics for performance tuning.

## Practical Implementation Example

```rust
use meetily_frontend::audio::ffmpeg_mixer::FFmpegAudioMixer;
use meetily_frontend::audio::device_detection::InputDeviceKind;

// Initialize the mixer at 48 kHz sample rate
let mut mixer = FFmpegAudioMixer::new(
    "Built-in Mic".to_string(),
    InputDeviceKind::Wired,
    "BlackHole 2ch".to_string(),
    InputDeviceKind::Bluetooth,
    48_000,
);

// Push raw PCM frames from capture threads
mixer.push_mic(mic_frame);
mixer.push_system(system_frame);

// Retrieve mixed audio when synchronized
if mixer.has_data_ready() {
    if let Some(mixed_frame) = mixer.pop_mixed() {
        // Process for WAV storage or Whisper transcription
        save_to_wav(&mixed_frame);
    }
}

// Access diagnostic statistics
let (mic_stats, sys_stats) = mixer.get_stats();
println!("Mic latency: {:.0} ms, gaps: {}", 
         mic_stats.buffer_latency_ms, 
         mic_stats.gaps_detected);
println!("System latency: {:.0} ms, silence: {:.1} ms", 
         sys_stats.buffer_latency_ms, 
         sys_stats.silence_inserted_ms);

```

## Performance Monitoring and Diagnostics

The mixer logs statistics every 200 windows via `FFmpegAudioMixer::log_stats`, exposing buffer latency, gap counts, and silence insertion durations through `SourceBuffer::stats`. These metrics enable runtime tuning of Bluetooth timeouts and jitter compensation without interrupting the audio pipeline.

## Summary

- **FFmpegMixer** in [`frontend/src-tauri/src/audio/ffmpeg_mixer.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/ffmpeg_mixer.rs) provides device-aware adaptive buffering using separate `SourceBuffer` instances for each audio source.
- **Bluetooth devices** receive 80–200 ms timeout windows compared to 20–50 ms for wired devices, accommodating wireless jitter while maintaining synchronization.
- **RMS-based ducking** automatically reduces system audio to 60% when microphone speech exceeds 0.01 RMS, ensuring voice clarity.
- **Fixed 50 ms mixing windows** with hard clamping to `[-1.0, 1.0]` prevent clipping and provide predictable latency.
- **Automatic silence insertion** and gap detection maintain stream continuity during buffer underruns or Bluetooth interference.

## Frequently Asked Questions

### How does FFmpegMixer handle Bluetooth audio latency differently from wired microphones?

The mixer distinguishes device types through the `InputDeviceKind` enum. Bluetooth sources receive longer timeout windows (80–200 ms) via `device_kind.buffer_timeout()`, while wired devices use shorter 20–50 ms windows. This accommodates the inherent jitter of wireless protocols without introducing unnecessary latency for stable wired connections.

### What triggers the audio ducking feature in Meetily's mixer?

Ducking activates when `calculate_rms` detects microphone levels exceeding **0.01 RMS** within the current 50 ms window. When triggered, `AudioMixer::mix` reduces system audio amplitude to 60% of its original level, restoring full volume when speech ceases. This preserves voice intelligibility without completely muting background audio.

### How does the mixer prevent audio clipping when combining multiple sources?

Every sample undergoes hard limiting via `mixed.clamp(-1.0, 1.0)` within the `AudioMixer::mix` method. This ensures that summed audio signals—particularly when both microphone and system audio peak simultaneously—never exceed the valid range for 32-bit floating-point PCM, guaranteeing clean output WAV files regardless of input gain staging.

### Can FFmpegMixer handle more than two audio sources simultaneously?

The current implementation in [`ffmpeg_mixer.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/ffmpeg_mixer.rs) explicitly instantiates two `SourceBuffer` instances (microphone and system) within `FFmpegAudioMixer::new`. While the underlying `SourceBuffer` architecture supports arbitrary expansion, the existing mixer structure is optimized for dual-source scenarios (voice + system audio) common in meeting recording contexts.