How Retranscription Works in Meetily: Re‑processing Stored Audio with Different Settings
Retranscription in Meetily allows users to re‑run saved meeting audio through the Whisper speech‑to‑text engine with new language, model, or provider settings, implemented as a background task in the Tauri core with live progress events and atomic concurrency guards.
Meetily's open-source meeting assistant stores every recorded session locally. When you need higher accuracy, a different language, or alternative AI providers, the retranscription feature reprocesses that stored audio without requiring a new recording. This article explains the complete pipeline implemented in the Zackriya-Solutions/meetily repository, from the frontend command to metadata persistence.
Entry Point: Triggering a Retranscription
The process begins when the frontend invokes the Tauri command start_retranscription_command. This command is registered in src-tauri/src/lib.rs at lines 742–744 and immediately delegates to the internal start_retranscription function.
// Frontend: start a retranscription with different settings
await invoke('start_retranscription', {
meetingId: '12345',
language: 'es', // Switch to Spanish
model: 'medium', // Upgrade from 'small' to 'medium' Whisper model
provider: 'ollama' // Use local Ollama instance instead of cloud
});
The command handler accepts optional parameters for language, model, and provider, allowing any combination of settings changes. If omitted, defaults fall back to system configuration.
Concurrency Control with Atomic Guards
Retranscription is resource-intensive. To prevent GPU/CPU thrashing, Meetily enforces single-instance execution using an atomic flag.
In src-tauri/src/audio/retranscription.rs (lines 19–26), a global RETRANSLATION_IN_PROGRESS AtomicBool guards the entry point. The flag is set inside a guard object that automatically clears on scope exit, ensuring cleanup even if the process panics.
// Simplified guard pattern from the source
let _guard = RetranslationGuard::new(); // Sets atomic flag
// ... retranscription runs ...
// Flag cleared automatically when _guard drops
Attempts to start a second retranscription while one is active return an immediate error: "Retranscription already in progress".
Parameter Validation and Engine Preparation
Before touching audio files, start_retranscription validates that the selected provider supports language-aware Whisper runs. At line 612 in retranscription.rs, unsupported configurations trigger a clear abort with descriptive error messaging.
Once validated, the system instantiates a fresh WhisperEngine via WhisperEngine::new with the requested model and provider. The engine lifecycle follows explicit load/unload semantics:
- Loading: New engine initialization with custom settings
- Processing: Batch transcription of all discovered audio
- Unloading:
unload_engineinsrc-tauri/src/audio/common.rs(line 17) frees memory and GPU context
This explicit cleanup prevents resource leaks when switching between providers with incompatible backends.
Audio Discovery and Decoding
Retranscription targets the meeting's original audio folder. The system enumerates all files matching extensions defined in audio/constants.rs (.wav, .mp3, .m4a, .webm, etc.).
Each file passes through the AudioDecoder in audio/decoder.rs, which:
- Demuxes container formats
- Resamples to Whisper's expected 16kHz mono PCM
- Returns a memory-efficient stream for chunked processing
This re‑decoding step is essential: stored audio is reconverted from its compressed format, not reused from any prior transcription cache.
The Transcription Loop with Progress Events
The core processing loop iterates over decoded audio chunks. For each segment:
// Inside run_retranscription – processing each audio file
for audio_path in audio_files {
let pcm = AudioDecoder::decode(&audio_path).await?;
let transcript = whisper_engine.transcribe(pcm, &language).await?;
// Results accumulated and progress emitted...
}
Progress flows back to the UI through Tauri's event system. At line 510 of retranscription.rs, the "retranscription-progress" event carries percentage completion, enabling live progress bars without blocking the main thread.
Metadata Persistence and History Tracking
Successful completion updates the meeting's metadata.json via write_retranscription_metadata (lines 723–756). The function:
- Preserves all existing fields (original transcription, timestamps, tags)
- Adds
retranscribed_atwith ISO 8601 timestamp - Marks
"source": "retranscription"to distinguish from live recordings
This append-only approach maintains full audit history. Users can compare original and retranscribed outputs without data destruction.
Completion and Error Handling
The pipeline reports final state through two events:
| Event | Trigger | Payload |
|---|---|---|
"retranscription-complete" |
All audio processed successfully | Meeting ID, duration, new transcript path |
"retranscription-error" |
Failure at any stage | Error message, failed file (if applicable) |
These events at lines 116 and 127 enable UI toasts, automatic retry prompts, or silent background retry logic.
Cancellation Support
Long retranscriptions can be interrupted. The cancel_retranscription_command (lines 822–827) flips the atomic progress flag, causing the background task to abort at the next chunk boundary. Resources are released through the normal guard cleanup path.
// Frontend: abort in-progress retranscription
await invoke('cancel_retranscription');
Summary
- Single-entry command:
start_retranscription_commandinlib.rsroutes frontend requests - Atomic concurrency: Guard pattern prevents overlapping runs in
retranscription.rs - Fresh engine per job: Explicit load/unload in
common.rsmanages GPU/CPU resources - Re‑decode everything:
AudioDecoderreprocesses stored files; no cached audio reuse - Event-driven progress: Live percentage updates via
"retranscription-progress" - Non-destructive metadata:
write_retranscription_metadatapreserves history withretranscribed_attimestamp - Cancellable: Atomic flag check enables mid-operation abort
Frequently Asked Questions
What file formats support retranscription in Meetily?
All formats listed in src-tauri/src/audio/constants.rs are eligible: WAV, MP3, M4A, WEBM, OGG, and FLAC. The AudioDecoder handles format detection and resampling automatically, so retranscription works regardless of the original recording format.
Can I retranscribe with a different Whisper model size?
Yes. The model parameter accepts any valid Whisper model name (tiny, base, small, medium, large, large-v3, etc.). Meetily instantiates a fresh engine with the requested model, allowing quality upgrades or faster iterations with smaller variants.
Does retranscription overwrite my original transcript?
No. The original transcript file remains untouched. The new output is written alongside it, and metadata.json tracks both the original and retranscribed versions with timestamps and source markers. You can reference either version in the Meetily UI.
Why does retranscription fail with "provider does not support language selection"?
Not all Whisper providers implement language-aware transcription. The validation at line 612 of retranscription.rs checks provider capabilities before engine initialization. Switch to a compatible provider (OpenAI, local Whisper.cpp, or Ollama with proper configuration) or omit the language parameter to use the provider's auto-detection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →