Meetily Audio Processing Stages: Resampling, High-Pass Filtering, Noise Suppression, and Loudness Normalization
Meetily processes every microphone stream through four distinct DSP stages—dynamic resampling, 80 Hz high-pass filtering, RNNoise-based noise suppression, and EBU R128 loudness normalization—before forwarding audio to the transcription engine.
Meetily is an open-source meeting transcription application built with Rust and Tauri. According to the source code in Zackriya-Solutions/meetily, the audio_processing module orchestrates a strict signal chain that ensures consistent sample rates, removes unwanted low-frequency content, suppresses background noise, and normalizes loudness to broadcast standards.
The Four-Stage Processing Pipeline
The pipeline is implemented in the process_audio_data method (referenced in frontend/src-tauri/src/audio/audio_processing.rs) and executes in the following order:
1. Dynamic Resampling to 48 kHz
If the capture device’s sample rate differs from the pipeline’s target of 48 kHz, Meetily applies dynamic resampling using the rubato crate’s SincFixedIn resampler. This processor maintains a persistent resampler instance with a 512‑sample internal chunk size, ensuring energy preservation across variable buffer sizes from devices reporting 16 kHz (Bluetooth headsets) or 44.1 kHz (professional interfaces).
2. High-Pass Filtering (80 Hz Cutoff)
Immediately after resampling, the HighPassFilter struct (defined in frontend/src-tauri/src/audio/audio_processing/high_pass_filter.rs) removes low‑frequency rumble below 80 Hz. This prevents subsonic noise from masking speech frequencies and reduces mechanical vibration artifacts common in laptop microphones.
3. RNNoise-Based Noise Suppression
The NoiseSuppressionProcessor (located in frontend/src-tauri/src/audio/audio_processing/noise_suppression.rs) implements RNNoise deep‑learning denoising. When the compile‑time flag RNNOISE_APPLY_ENABLED is active, this stage applies 10–15 dB of noise reduction. The processor is instantiated with NoiseSuppressionProcessor::new(TARGET_SAMPLE_RATE) and processes the mono stream in place.
4. EBU R128 Loudness Normalization
The final stage uses LoudnessNormalizer (defined in frontend/src-tauri/src/audio/audio_processing/loudness_normalizer.rs) to bring microphone audio to ‑23 LUFS, complying with the EBU R128 broadcast standard. This ensures consistent perceived loudness regardless of hardware sensitivity differences across devices.
Implementation Details and Code Flow
The processing chain executes inside AudioCapture::process_audio_data. Before the DSP stages, multichannel audio is converted to mono via audio_to_mono (utility in frontend/src-tauri/src/audio/audio_processing/audio_to_mono.rs), which averages channels mathematically:
// Convert stereo to mono before processing
let mut mono_data = audio_to_mono(input_samples, channels);
The sequential processing then occurs as follows:
// 1. Resample if needed (e.g., 44.1kHz -> 48kHz)
if self.needs_resampling {
mono_data = self.resampler.process(&mono_data);
}
// 2. Apply 80Hz high-pass filter
if let Some(ref mut hpf) = *self.high_pass_filter.lock()? {
mono_data = hpf.process(&mono_data);
}
// 3. RNNoise suppression (optional, compile-time enabled)
if let Some(ref mut ns) = *self.noise_suppressor.lock()? {
mono_data = ns.process(&mono_data);
}
// 4. Normalize to -23 LUFS
if let Some(ref mut norm) = *self.normalizer.lock()? {
mono_data = norm.normalize_loudness(&mono_data);
}
After normalization, the cleaned AudioChunk is forwarded to the AudioPipeline (defined in frontend/src-tauri/src/audio/pipeline.rs) for mixing with system audio and voice‑activity detection.
Source File Architecture
| File | Purpose |
|---|---|
frontend/src-tauri/src/audio/audio_processing/mod.rs |
Declares sub‑modules: audio_to_mono, loudness_normalizer, noise_suppression, high_pass_filter |
frontend/src-tauri/src/audio/audio_processing/audio_to_mono.rs |
Channel averaging utility for mono conversion |
frontend/src-tauri/src/audio/audio_processing/high_pass_filter.rs |
State‑variable filter implementation with configurable cutoff |
frontend/src-tauri/src/audio/audio_processing/noise_suppression.rs |
RNNoise wrapper and processing state management |
frontend/src-tauri/src/audio/audio_processing/loudness_normalizer.rs |
EBU R128 loudness calculation and gain application |
frontend/src-tauri/src/audio/pipeline.rs |
Imports processors and orchestrates the final mixing/VAD pipeline |
Summary
- Dynamic Resampling: Converts any input rate to 48 kHz using Rubato’s high‑quality sinc interpolation.
- High-Pass Filtering: Removes sub‑80 Hz rumble to clean up the speech band.
- Noise Suppression: Applies RNNoise deep‑learning reduction for 10–15 dB cleaner speech.
- Loudness Normalization: Targets ‑23 LUFS per EBU R128 for broadcast‑consistent levels.
These stages execute sequentially in process_audio_data, ensuring that microphone audio is acoustically optimized before transcription or mixing with system capture.
Frequently Asked Questions
What sample rate does Meetily require for transcription?
Meetily internally resamples all microphone input to 48 kHz using a persistent SincFixedIn resampler. This standardizes the pipeline regardless of whether the hardware provides 16 kHz (Bluetooth), 44.1 kHz (consumer audio), or 48 kHz (professional) sources.
Why does the high-pass filter use an 80 Hz cutoff?
The 80 Hz cutoff in HighPassFilter targets electrical hum and mechanical rumble (HVAC, laptop fans, desk vibrations) that occupy the sub‑speech spectrum. Removing this content prevents low‑frequency energy from consuming dynamic range and interfering with voice activity detection algorithms.
Is noise suppression always active in Meetily?
No. RNNoise suppression is controlled by the compile‑time flag RNNOISE_APPLY_ENABLED. When enabled, NoiseSuppressionProcessor applies real‑time denoising; when disabled, the stage is skipped to reduce CPU load on resource‑constrained devices.
What does ‑23 LUFS mean in the loudness normalizer?
LUFS (Loudness Units relative to Full Scale) is a perceptual loudness standard. Meetily’s LoudnessNormalizer targets ‑23 LUFS according to the EBU R128 specification, ensuring that recorded speech matches the loudness of professional broadcast content and preventing volume differences between meeting participants.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →