Key Algorithms Implemented in Meetily: Audio Processing & AI Pipeline Explained
Meetily implements a sophisticated pipeline of audio processing algorithms—including ring-buffer mixing, Rubato resampling, RNNoise suppression, EBU R128 loudness normalization, and Whisper-based transcription—all orchestrated through a Rust core to deliver privacy-first meeting assistance.
Meetily is a privacy-first AI meeting assistant built as a Tauri desktop application with a high-performance Rust backend. Understanding the key algorithms implemented in Meetily reveals how it captures studio-quality audio, performs on-device speech recognition, and generates intelligent summaries without transmitting sensitive raw audio to external cloud services.
Audio Capture and Ring-Buffer Management
The foundation of Meetily's audio pipeline relies on Ring-Buffer Audio Mixing to handle asynchronous input streams. The system accumulates microphone and system-audio samples in two synchronized VecDeque buffers, each sized to hold 400ms of audio—deliberately oversized to absorb jitter from macOS Core Audio and other platform-specific latency variations.
From these buffers, the pipeline extracts fixed-size windows of approximately 50ms for processing. This design ensures continuous audio flow even when the operating system introduces unpredictable delays during capture.
Professional Audio Mixing and Stream Synchronization
Once aligned, the Professional Audio Mixer blends the two streams using soft scaling algorithms to prevent hard clipping. System audio is attenuated to approximately 70% volume while microphone audio remains at full level; when the combined sum exceeds ±1.0, the mixer proportionally scales both channels back to maintain fidelity.
This implementation lives in src/audio/pipeline.rs, where the mixer handles synchronized windows from both capture devices to produce a clean combined track suitable for transcription and recording.
Resampling and Format Normalization
Meetily handles mismatched device sample rates—common with Bluetooth headsets at 16kHz or legacy hardware at 44.1kHz—through a Rubato Resampling algorithm. The pipeline detects rate mismatches and instantiates a persistent SincFixedIn resampler from the Rubato crate, configured to process 512-sample chunks.
The resampler adapts its filter length, interpolation type, and oversampling parameters dynamically based on the up- or down-sampling ratio, preserving phase continuity while minimizing CPU overhead. This logic resides in the needs_resampling branch of src/audio/pipeline.rs.
Noise Reduction and Audio Enhancement
Microphone input undergoes three distinct enhancement stages before transcription:
- RNNoise Suppression: An optional processor (controlled by the
RNNOISE_APPLY_ENABLEDflag) reduces background noise by 10-15dB, instantiated once per device at the 48kHz target rate - High-Pass Filtering: Removes low-frequency rumble below 80Hz to improve both human listening and downstream transcription accuracy
- EBU R128 Loudness Normalization: Implements the professional broadcast standard to drive microphone audio toward -23 LUFS, ensuring consistent volume across different recording environments
All three processors are implemented as fields (noise_suppressor, high_pass_filter, normalizer) within the main pipeline structure in src/audio/pipeline.rs.
Voice Activity Detection (VAD)
To reduce computational overhead, Meetily employs a Continuous VAD Processor that analyzes incoming samples and forwards only speech segments to the transcription engine. This algorithm, implemented in src/audio/vad.rs, achieves approximately 70% reduction in transcription workload by filtering out silence and non-speech audio before it reaches the Whisper engine.
Speech-to-Text with Whisper
The Whisper Transcription Engine wraps the whisper-rs crate (leveraging Whisper.cpp) to perform on-device speech-to-text. Located in src/whisper_engine/whisper_engine.rs, the engine loads the selected model once at startup, automatically detects GPU capabilities (Metal on macOS, CUDA on NVIDIA, or Vulkan), and gracefully falls back to CPU inference when hardware acceleration is unavailable.
Parallel Processing Optimization
For improved throughput on multi-core systems, Meetily distributes transcription work across multiple async tasks using the Parallel Whisper Processor found in src/whisper_engine/parallel_processor.rs. This algorithm automatically parallelizes chunk processing when the selected model supports it, maximizing utilization of available CPU cores or GPU compute units.
LLM-Based Meeting Summarization
Once transcription completes, the LLM Client handles summarization through an abstraction layer supporting local models (Ollama) and remote providers (OpenRouter, Claude, Groq). Implemented in src/summary/llm_client.rs, this algorithm manages authentication, streaming response handling, and error recovery to deliver concise meeting summaries without storing sensitive data on external servers.
Real-Time Diagnostics and Monitoring
Throughout the pipeline, Audio Level Monitoring and Metrics Batching provide visibility into system health. The SimpleLevelMonitor in src/audio/simple_level_monitor.rs continuously emits buffer sizes, VAD detection rates, and clipping warnings to the frontend UI. Meanwhile, the AudioMetricsBatcher in src/audio/batch_processor.rs aggregates fine-grained statistics (RMS, peak, loudness) in batches to prevent excessive logging overhead.
Integration Example
Developers can interact with these algorithms through the Rust API as follows:
// Initialize microphone capture with full pipeline processing
let mic_capture = AudioCapture::new(
mic_device.clone(),
recording_state.clone(),
mic_device.sample_rate,
mic_device.channels,
DeviceType::Microphone,
Some(audio_sender.clone()),
);
// Start the capture stream (internally applies VAD, normalization, and noise suppression)
mic_capture.start()?;
// Process raw PCM data from the OS callback
fn audio_callback(data: &[f32]) {
mic_capture.process_audio_data(data);
}
// Send processed audio to Whisper for transcription
whisper_engine.transcribe(mixed_audio).await?;
Summary
- Ring-buffer architecture in
src/audio/pipeline.rsuses 400ms oversizedVecDequebuffers to absorb platform jitter while extracting 50ms processing windows - Rubato resampling handles sample rate mismatches through persistent
SincFixedInresamplers working on 512-sample chunks - Audio enhancement chain combines RNNoise suppression, 80Hz high-pass filtering, and EBU R128 loudness normalization targeting -23 LUFS
- VAD processing in
src/audio/vad.rsreduces transcription workload by 70% through continuous speech detection - Whisper engine with parallel processing support handles on-device transcription across Metal, CUDA, and CPU backends
- LLM abstraction in
src/summary/llm_client.rsenables privacy-preserving summarization across multiple provider APIs
Frequently Asked Questions
What audio processing algorithms does Meetily use for noise reduction?
Meetily implements RNNoise (controlled by the RNNOISE_APPLY_ENABLED flag) for background noise suppression, achieving 10-15dB reduction, alongside an 80Hz high-pass filter to remove low-frequency rumble. These algorithms run in series within the pipeline.rs file before the audio reaches the transcription engine.
How does Meetily handle different sample rates between audio devices?
The system detects mismatched sample rates using the needs_resampling logic in src/audio/pipeline.rs and creates persistent SincFixedIn resamplers via the Rubato crate. These resamplers process audio in 512-sample chunks, adapting filter parameters to maintain phase continuity when converting between common rates like 16kHz, 44.1kHz, and the internal 48kHz target.
Does Meetily process audio in real-time or batch mode?
Meetily processes audio in real-time using a continuous pipeline: the AudioCapture streams data through ring buffers, the ContinuousVadProcessor filters speech segments immediately, and the WhisperEngine transmits partial transcripts via Tauri events. However, summarization operates in batch mode, waiting for the meeting to conclude before sending the complete transcript to the LLM client.
What makes Meetily's transcription pipeline privacy-preserving?
All audio processing—including noise suppression, resampling, and Whisper transcription—executes locally on the device through the Rust backend, with raw audio never transmitted to external servers. Only the final text transcript leaves the device (if using remote LLM providers), and local model support through Ollama enables fully offline operation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →