# Key Algorithms Implemented in Meetily: Audio Processing & AI Pipeline Explained

> Explore key audio processing and AI algorithms in Meetily. Discover ring-buffer mixing, Rubato resampling, RNNoise, EBU R128, and Whisper transcription powered by Rust for privacy-first meeting assistance.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: internals
- Published: 2026-07-29

---

**Meetily implements a sophisticated pipeline of audio processing algorithms—including ring-buffer mixing, Rubato resampling, RNNoise suppression, EBU R128 loudness normalization, and Whisper-based transcription—all orchestrated through a Rust core to deliver privacy-first meeting assistance.**

Meetily is a privacy-first AI meeting assistant built as a Tauri desktop application with a high-performance Rust backend. Understanding the key algorithms implemented in Meetily reveals how it captures studio-quality audio, performs on-device speech recognition, and generates intelligent summaries without transmitting sensitive raw audio to external cloud services.

## Audio Capture and Ring-Buffer Management

The foundation of Meetily's audio pipeline relies on **Ring-Buffer Audio Mixing** to handle asynchronous input streams. The system accumulates microphone and system-audio samples in two synchronized `VecDeque` buffers, each sized to hold 400ms of audio—deliberately oversized to absorb jitter from macOS Core Audio and other platform-specific latency variations.

From these buffers, the pipeline extracts fixed-size windows of approximately 50ms for processing. This design ensures continuous audio flow even when the operating system introduces unpredictable delays during capture.

## Professional Audio Mixing and Stream Synchronization

Once aligned, the **Professional Audio Mixer** blends the two streams using soft scaling algorithms to prevent hard clipping. System audio is attenuated to approximately 70% volume while microphone audio remains at full level; when the combined sum exceeds ±1.0, the mixer proportionally scales both channels back to maintain fidelity.

This implementation lives in [`src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/audio/pipeline.rs), where the mixer handles synchronized windows from both capture devices to produce a clean combined track suitable for transcription and recording.

## Resampling and Format Normalization

Meetily handles mismatched device sample rates—common with Bluetooth headsets at 16kHz or legacy hardware at 44.1kHz—through a **Rubato Resampling** algorithm. The pipeline detects rate mismatches and instantiates a persistent `SincFixedIn` resampler from the Rubato crate, configured to process 512-sample chunks.

The resampler adapts its filter length, interpolation type, and oversampling parameters dynamically based on the up- or down-sampling ratio, preserving phase continuity while minimizing CPU overhead. This logic resides in the `needs_resampling` branch of [`src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/audio/pipeline.rs).

## Noise Reduction and Audio Enhancement

Microphone input undergoes three distinct enhancement stages before transcription:

- **RNNoise Suppression**: An optional processor (controlled by the `RNNOISE_APPLY_ENABLED` flag) reduces background noise by 10-15dB, instantiated once per device at the 48kHz target rate
- **High-Pass Filtering**: Removes low-frequency rumble below 80Hz to improve both human listening and downstream transcription accuracy
- **EBU R128 Loudness Normalization**: Implements the professional broadcast standard to drive microphone audio toward -23 LUFS, ensuring consistent volume across different recording environments

All three processors are implemented as fields (`noise_suppressor`, `high_pass_filter`, `normalizer`) within the main pipeline structure in [`src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/audio/pipeline.rs).

## Voice Activity Detection (VAD)

To reduce computational overhead, Meetily employs a **Continuous VAD Processor** that analyzes incoming samples and forwards only speech segments to the transcription engine. This algorithm, implemented in [`src/audio/vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/audio/vad.rs), achieves approximately 70% reduction in transcription workload by filtering out silence and non-speech audio before it reaches the Whisper engine.

## Speech-to-Text with Whisper

The **Whisper Transcription Engine** wraps the `whisper-rs` crate (leveraging Whisper.cpp) to perform on-device speech-to-text. Located in [`src/whisper_engine/whisper_engine.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/whisper_engine/whisper_engine.rs), the engine loads the selected model once at startup, automatically detects GPU capabilities (Metal on macOS, CUDA on NVIDIA, or Vulkan), and gracefully falls back to CPU inference when hardware acceleration is unavailable.

## Parallel Processing Optimization

For improved throughput on multi-core systems, Meetily distributes transcription work across multiple async tasks using the **Parallel Whisper Processor** found in [`src/whisper_engine/parallel_processor.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/whisper_engine/parallel_processor.rs). This algorithm automatically parallelizes chunk processing when the selected model supports it, maximizing utilization of available CPU cores or GPU compute units.

## LLM-Based Meeting Summarization

Once transcription completes, the **LLM Client** handles summarization through an abstraction layer supporting local models (Ollama) and remote providers (OpenRouter, Claude, Groq). Implemented in [`src/summary/llm_client.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/summary/llm_client.rs), this algorithm manages authentication, streaming response handling, and error recovery to deliver concise meeting summaries without storing sensitive data on external servers.

## Real-Time Diagnostics and Monitoring

Throughout the pipeline, **Audio Level Monitoring** and **Metrics Batching** provide visibility into system health. The `SimpleLevelMonitor` in [`src/audio/simple_level_monitor.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/audio/simple_level_monitor.rs) continuously emits buffer sizes, VAD detection rates, and clipping warnings to the frontend UI. Meanwhile, the `AudioMetricsBatcher` in [`src/audio/batch_processor.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/audio/batch_processor.rs) aggregates fine-grained statistics (RMS, peak, loudness) in batches to prevent excessive logging overhead.

## Integration Example

Developers can interact with these algorithms through the Rust API as follows:

```rust
// Initialize microphone capture with full pipeline processing
let mic_capture = AudioCapture::new(
    mic_device.clone(),
    recording_state.clone(),
    mic_device.sample_rate,
    mic_device.channels,
    DeviceType::Microphone,
    Some(audio_sender.clone()),
);

// Start the capture stream (internally applies VAD, normalization, and noise suppression)
mic_capture.start()?;

// Process raw PCM data from the OS callback
fn audio_callback(data: &[f32]) {
    mic_capture.process_audio_data(data);
}

// Send processed audio to Whisper for transcription
whisper_engine.transcribe(mixed_audio).await?;

```

## Summary

- **Ring-buffer architecture** in [`src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/audio/pipeline.rs) uses 400ms oversized `VecDeque` buffers to absorb platform jitter while extracting 50ms processing windows
- **Rubato resampling** handles sample rate mismatches through persistent `SincFixedIn` resamplers working on 512-sample chunks
- **Audio enhancement chain** combines RNNoise suppression, 80Hz high-pass filtering, and EBU R128 loudness normalization targeting -23 LUFS
- **VAD processing** in [`src/audio/vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/audio/vad.rs) reduces transcription workload by 70% through continuous speech detection
- **Whisper engine** with parallel processing support handles on-device transcription across Metal, CUDA, and CPU backends
- **LLM abstraction** in [`src/summary/llm_client.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/summary/llm_client.rs) enables privacy-preserving summarization across multiple provider APIs

## Frequently Asked Questions

### What audio processing algorithms does Meetily use for noise reduction?

Meetily implements RNNoise (controlled by the `RNNOISE_APPLY_ENABLED` flag) for background noise suppression, achieving 10-15dB reduction, alongside an 80Hz high-pass filter to remove low-frequency rumble. These algorithms run in series within the [`pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/pipeline.rs) file before the audio reaches the transcription engine.

### How does Meetily handle different sample rates between audio devices?

The system detects mismatched sample rates using the `needs_resampling` logic in [`src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/src/audio/pipeline.rs) and creates persistent `SincFixedIn` resamplers via the Rubato crate. These resamplers process audio in 512-sample chunks, adapting filter parameters to maintain phase continuity when converting between common rates like 16kHz, 44.1kHz, and the internal 48kHz target.

### Does Meetily process audio in real-time or batch mode?

Meetily processes audio in real-time using a continuous pipeline: the `AudioCapture` streams data through ring buffers, the `ContinuousVadProcessor` filters speech segments immediately, and the `WhisperEngine` transmits partial transcripts via Tauri events. However, summarization operates in batch mode, waiting for the meeting to conclude before sending the complete transcript to the LLM client.

### What makes Meetily's transcription pipeline privacy-preserving?

All audio processing—including noise suppression, resampling, and Whisper transcription—executes locally on the device through the Rust backend, with raw audio never transmitted to external servers. Only the final text transcript leaves the device (if using remote LLM providers), and local model support through Ollama enables fully offline operation.