Meetily Real-Time vs Post-Processing Transcription: Performance Tradeoffs and Technical Implementation
Meetily offers two transcription modes: real-time processing during recording minimizes latency to tens of milliseconds but sacrifices some accuracy and increases continuous CPU load, while post-processing batch transcription delivers higher accuracy and cleaner text but requires waiting for the full meeting to process.
Meetily's open-source meeting assistant provides flexible transcription capabilities through dual operating modes implemented in Rust and TypeScript. Understanding the performance tradeoffs between real-time transcription during recording versus post-processing transcription helps developers optimize for latency-sensitive accessibility features or accuracy-critical archival workflows.
How Real-Time Transcription Works in Meetily
The VAD-Driven Pipeline Architecture
The real-time pipeline centers on AudioPipeline::run in frontend/src-tauri/src/audio/pipeline.rs. This implementation uses Voice Activity Detection (VAD) to segment incoming audio streams immediately upon capture. When speech is detected, the pipeline forwards chunks containing at least 800 samples at 16 kHz to transcription engines without waiting for the recording to end.
The main loop processes chunks on the fly at lines 66-68, sending only speech segments (not silence) to either the WhisperEngine or ParakeetEngine. This VAD-gated approach prevents wasted inference cycles on non-speech audio, though the pipeline deliberately drops very short utterances below the 800-sample threshold to maintain responsiveness.
Memory and Latency Optimization
Memory usage remains minimal due to the ring buffer architecture described in frontend/src-tauri/src/audio/pipeline.rs at lines 19-23. The system maintains short-lived 512-sample resampler chunks, discarding old data as the pipeline progresses rather than accumulating the entire meeting history.
This design achieves near-zero latency measured in tens of milliseconds, as each validated VAD segment triggers immediate model inference. The continuous processing demands persistent CPU or GPU resources, making this mode suitable for desktop environments with constant power supply but potentially draining for battery-operated devices.
How Post-Processing Transcription Works in Meetily
Batch Processing Architecture
Post-processing mode activates after recording completion, feeding the full mixed audio file (combining microphone and system audio) to transcription engines in a single operation. Unlike the segmented real-time approach, this method loads the complete WAV file into memory or streams it from disk, invoking the engine once rather than continuously.
The transcription engines—WhisperEngine in frontend/src-tauri/src/whisper_engine/whisper_engine.rs or ParakeetEngine in frontend/src-tauri/src/parakeet_engine/mod.rs—process the entire waveform as a unified input. This batch approach can leverage model-specific optimizations for long-form audio, though it spikes resource utilization for the duration of the meeting length.
Accuracy and Contextual Improvements
Post-processing achieves higher transcription accuracy because models access the full temporal context of the conversation. Without the constraints of 800-sample segmentation boundaries, the engine better handles cross-sentence punctuation, speaker transitions, and contextual disambiguation.
After transcription completes, the PostProcessor class in frontend/src-tauri/src/audio/post_processor.rs applies aggressive cleanup operations. Lines 44-70 implement deduplication, filler-word removal, capitalization fixes, and artifact removal—operations that require lookahead context unavailable during streaming.
Performance Tradeoffs: Real-Time vs Post-Processing
Latency Characteristics: Real-time mode delivers results within tens of milliseconds through immediate VAD segment processing, enabling live captioning and instant feedback. Post-processing delays output until the entire file processes, creating a wait time proportional to meeting duration.
CPU and GPU Load: Real-time transcription creates continuous computational load as every VAD segment triggers inference, though Meetily mitigates this by using lightweight models like Parakeet when available. Post-processing concentrates load into a single burst, potentially more efficient if the model supports batch operations but spiking resources for the meeting's full duration.
Memory Footprint: The real-time pipeline maintains small memory footprints using ring buffers that discard old data after processing 512-sample chunks. Post-processing requires holding several minutes of mixed audio in RAM or maintaining disk streams, significantly increasing memory pressure according to the implementation in audio_capture.rs.
Accuracy and Error Resilience: Real-time processing may miss short utterances under 800 samples or cut speech at unfavorable VAD boundaries. However, errors are handled per-chunk without aborting the meeting—failures log to the console and the pipeline continues. Post-processing provides superior accuracy through full-context analysis, though a single out-of-memory failure can abort the entire transcription requiring a retry.
Power Consumption: Continuous real-time processing maintains active CPU/GPU states throughout the meeting, draining batteries on laptops and mobile devices. Post-processing allows devices to sleep during transcription or run the operation as a background batch job.
Post-Processing Depth: Real-time mode applies only lightweight PostProcessor cleanup for partial results (lines 6-14), while post-processing leverages the full cleanup suite including contextual improvements and aggressive deduplication.
Implementation Examples
Real-Time Transcription Setup
Start a meeting and receive live chunks through Tauri's event system:
// Front-end TypeScript – invoke the Tauri command that starts recording
await invoke('start_recording', {
mic_device_name: 'Built-in Microphone',
system_device_name: 'BlackHole 2ch',
meeting_name: 'Team Standup'
});
// Listen for live transcript updates emitted by the Rust core
await listen<TranscriptUpdate>('transcript-update', (e) => {
console.log('Live:', e.payload.text);
});
Post-Processing Transcription
Run after the recording is saved to disk:
use meetily::whisper_engine::WhisperEngine;
use meetily::audio::post_processor::PostProcessor;
// Load the mixed WAV file that was saved on disk
let wav_path = "/path/to/meeting.wav";
// Initialise Whisper (or Parakeet) – the engine can be reused
let engine = WhisperEngine::new()?; // src-tauri/src/whisper_engine/whisper_engine.rs
let audio = std::fs::read(wav_path)?;
// Transcribe the whole file in one call
let raw_text = engine.transcribe(&audio)?; // high-accuracy batch inference
// Clean up the raw transcript
let post = PostProcessor::new(); // src-tauri/src/audio/post_processor.rs
let request = PostProcessRequest {
sequence_id: 0,
raw_text,
is_partial: false,
timestamp: chrono::Utc::now().to_rfc3339(),
};
post.process_async(request)?;
if let Some(result) = post.try_recv().await {
println!("Final transcript: {}", result.processed_text);
}
Engine Selection Strategy
Choose engines based on your performance requirements:
// Choose the lightweight Parakeet engine for real-time mode
let engine = ParakeetEngine::new()?; // src-tauri/src/parakeet_engine/mod.rs
// Or pick Whisper for the highest quality in post-processing
let engine = WhisperEngine::new()?; // src-tauri/src/whisper_engine/whisper_engine.rs
Summary
- Real-time transcription in Meetily uses a VAD-driven pipeline with 800-sample minimum segments, delivering sub-100ms latency but potentially missing short utterances and requiring continuous CPU/GPU resources.
- Post-processing transcription processes complete mixed audio files after recording, providing higher accuracy through full-context analysis and comprehensive
PostProcessorcleanup at the cost of delayed output. - Memory usage differs significantly: real-time mode uses ring buffers with 512-sample chunks and minimal RAM, while post-processing loads entire meeting audio into memory.
- Error handling favors real-time resilience (per-chunk failures don't crash the meeting) versus post-processing atomicity (single failures abort the entire job).
- Engine selection allows optimization for speed (Parakeet) or accuracy (Whisper) depending on the chosen transcription mode.
Frequently Asked Questions
Which transcription mode should I use for live meetings?
Use real-time transcription when participants require immediate accessibility features like live captions or when you need real-time meeting assistance. According to the Meetily source code, this mode maintains responsiveness by dropping VAD segments below 800 samples, making it ideal for interactive scenarios despite slightly reduced accuracy.
Why is post-processing transcription more accurate than real-time?
Post-processing achieves higher accuracy because the model sees the entire waveform context rather than 800-sample chunks. The PostProcessor in frontend/src-tauri/src/audio/post_processor.rs applies aggressive deduplication, filler-word removal, and contextual punctuation fixes that require lookahead analysis impossible during streaming transcription.
Can I switch between transcription engines during a meeting?
Yes. Meetily supports runtime engine selection through ParakeetEngine::new() for lightweight real-time processing or WhisperEngine::new() for maximum accuracy. However, switching modes (real-time vs post-processing) requires stopping the current pipeline and restarting with the new configuration, as the architectures differ fundamentally in their buffer management and inference timing.
How does Meetily handle transcription errors in real-time mode?
The real-time pipeline implements per-chunk error isolation in AudioPipeline::run at lines 62-66. If a specific VAD segment fails transcription (due to network issues, model errors, or corrupted audio), the system logs the error and continues processing subsequent chunks. This ensures that transient failures don't abort the entire meeting recording, unlike post-processing where a single out-of-memory error can terminate the batch job.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →