Real-Time Transcription Workflow in Meetily: From Audio Chunk to Text Explained
Meetily processes live audio transcription entirely on the client through a Rust-based pipeline that captures audio, filters noise, detects speech, and routes segments to local transcription engines—all without sending audio data over the network.
This article breaks down the complete real-time transcription workflow in Zackriya-Solutions/meetily, an open-source meeting assistant built with Tauri and Rust. Understanding this architecture helps developers implement low-latency speech-to-text pipelines for desktop applications.
Stage 1: Audio Capture and Pre-Processing
The pipeline begins in AudioCapture::process_audio_data within frontend/src-tauri/src/audio/pipeline.rs. This stage transforms raw microphone or system audio into a clean, normalized format suitable for downstream processing.
Key transformations applied to every audio chunk:
- Mono conversion — collapses stereo streams to single-channel
- Resampling — uses a persistent resampler to hit 48 kHz when needed
- High-pass filtering — removes low-frequency rumble
- RNNoise integration — applies deep learning-based noise suppression
- EBU R128 loudness normalization — ensures consistent levels
The processed data is wrapped in an AudioChunk struct containing sample data, sample rate, timestamp, chunk ID, and device type. This chunk is then handed to the shared RecordingState.
// Inside AudioCapture::process_audio_data
if self.state.is_recording() {
// ... mono conversion, optional resampling, filtering ...
let chunk = AudioChunk {
data: mono_data,
sample_rate: if self.needs_resampling { 48_000 } else { self.sample_rate },
timestamp: self.state.get_recording_duration().unwrap_or(0.0),
chunk_id: self.chunk_counter.fetch_add(1, Ordering::SeqCst),
device_type: self.device_type.clone(),
};
// Send to pipeline
self.state.send_audio_chunk(chunk)?;
}
Stage 2: Routing Through RecordingState
The RecordingState struct in frontend/src-tauri/src/audio/recording_state.rs serves as a thread-safe hub for audio chunks. Its send_audio_chunk method pushes chunks onto an UnboundedSender created by AudioPipelineManager.
This design decouples capture threads from processing threads, allowing the microphone capture to run independently of the heavier VAD and transcription workloads.
Stage 3: Mixing, VAD, and Speech Segmentation
The AudioPipeline receives chunks through its receiver channel and performs three critical operations:
Audio Mixing from Multiple Sources
When both microphone and system audio are captured, the AudioMixerRingBuffer (add_samples method, lines 49-73) stores streams in overlapping ring buffers. The ProfessionalAudioMixer::mix_window method (lines 54-86) blends these into a single 48 kHz stream with proper latency compensation.
Voice Activity Detection
The mixed audio feeds into ContinuousVadProcessor from frontend/src-tauri/src/audio/vad.rs. This extracts speech-only segments and emits them as new AudioChunk instances at 16 kHz—the optimal sample rate for most transcription engines.
while self.ring_buffer.can_mix() {
if let Some((mic_win, sys_win)) = self.ring_buffer.extract_window() {
let mixed = self.mixer.mix_window(&mic_win, &sys_win);
// VAD produces speech segments at 16 kHz
if let Ok(segments) = self.vad_processor.process_audio(&mixed) {
for seg in segments {
let vt_chunk = AudioChunk {
data: seg.samples,
sample_rate: 16_000,
timestamp: seg.start_timestamp_ms / 1000.0,
chunk_id: self.chunk_id_counter,
device_type: DeviceType::Microphone,
};
self.transcription_sender.send(vt_chunk)?;
self.chunk_id_counter += 1;
}
}
}
}
Stage 4: Transcription Provider Invocation
VAD-produced chunks travel through a dedicated transcription_sender channel to worker tasks. The transcribe_chunk_with_provider function in frontend/src-tauri/src/audio/transcription/worker.rs (lines 6-48) handles provider dispatch.
Three engine types are supported:
- WhisperEngine — OpenAI's Whisper model, local execution
- ParakeetEngine — NVIDIA's Riva Parakeet CTC model
- Custom providers — any type implementing the
TranscriptionProvidertrait
Each provider receives resampled 16 kHz mono audio and returns a TranscriptResult containing recognized text, optional confidence scores, and a partial-flag for streaming display.
let (text, confidence, partial) = match engine {
TranscriptionEngine::Whisper(w) => {
w.transcribe_audio_with_confidence(samples, language).await?
}
TranscriptionEngine::Parakeet(p) => {
(p.transcribe_audio(samples).await?, None, false)
}
TranscriptionEngine::Provider(p) => {
let res = p.transcribe(samples, language).await?;
(res.text, res.confidence, res.is_partial)
}
};
The TranscriptionProvider trait in frontend/src-tauri/src/audio/transcription/provider.rs enables seamless swapping between engines without modifying the worker logic.
Stage 5: Frontend Consumption via Tauri Events
Transcription results flow back to the React/Next.js UI through Tauri's event system. The frontend listens for transcript-update events and updates the meeting display in real time.
await listen<TranscriptUpdate>('transcript-update', (event) => {
setTranscripts(prev => [...prev, event.payload.text]);
});
The UI entry point in frontend/src/app/page.tsx wires these listeners to the transcript state, enabling live captions during meetings.
Architecture Benefits of Meetily's Real-Time Transcription
Privacy-first design — All audio processing happens locally in the Tauri process. No audio data leaves the machine unless explicitly configured for external LLM calls.
Deterministic resource management — Explicit flush signals and channel-based communication ensure clean shutdowns without zombie threads.
Modular provider system — The trait-based TranscriptionProvider interface allows community contributions of new transcription engines without pipeline changes.
Sub-millisecond latency — By avoiding network roundtrips and using lock-free ring buffers, the pipeline achieves processing delays suitable for live captioning.
Summary
- Audio capture in
AudioCapture::process_audio_datanormalizes raw PCM to 48 kHz mono chunks - RecordingState routes chunks through channels to decouple capture from processing
- AudioPipeline mixes multi-source audio and runs
ContinuousVadProcessorto extract 16 kHz speech segments - Transcription workers dispatch segments to Whisper, Parakeet, or custom providers via
transcribe_chunk_with_provider - Tauri events stream results to the React frontend for real-time display
The complete implementation spans frontend/src-tauri/src/audio/pipeline.rs, recording_state.rs, vad.rs, transcription/worker.rs, and transcription/provider.rs.
Frequently Asked Questions
How does Meetily achieve low latency in transcription?
Meetily eliminates network latency by running all transcription models locally through ONNX Runtime or similar inference engines. The Rust-based pipeline uses lock-free ring buffers and dedicated worker threads to keep audio capture, VAD, and transcription in parallel. According to the source code, chunks flow through channels with minimal copying, and VAD segments at 16 kHz match the native sample rate of most speech models, avoiding runtime resampling.
Can I use a custom transcription model with Meetily?
Yes. Implement the TranscriptionProvider trait defined in frontend/src-tauri/src/audio/transcription/provider.rs. Your implementation must provide a transcribe method that takes audio samples and a language code, returning a TranscriptResult. The TranscriptionEngine::Provider variant in the worker will automatically route chunks to your implementation. No changes to the audio pipeline or frontend are required.
Why does Meetily resample to 48 kHz during capture then down to 16 kHz for transcription?
The 48 kHz intermediate format provides headroom for professional audio mixing and aligns with standard system audio interfaces. The ContinuousVadProcessor specifically targets 16 kHz because this is the training sample rate for Whisper, Parakeet, and most other speech recognition models. Downsampling after VAD reduces computational load on the transcription engine while preserving the quality needed for recognition accuracy.
Where does noise suppression happen in the pipeline?
RNNoise filtering and high-pass filtering occur in AudioCapture::process_audio_data before chunks enter the shared state. This placement ensures all downstream components—mixing, VAD, and transcription—receive cleaned audio. The EBU R128 loudness normalization also happens here, preventing clipping and maintaining consistent levels across different microphone hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →