How Retranscription Works in Meetily: Re‑processing Stored Audio with Different Settings

Retranscription in Meetily allows users to re‑run saved meeting audio through the Whisper speech‑to‑text engine with new language, model, or provider settings, implemented as a background task in the Tauri core with live progress events and atomic concurrency guards.

Meetily's open-source meeting assistant stores every recorded session locally. When you need higher accuracy, a different language, or alternative AI providers, the retranscription feature reprocesses that stored audio without requiring a new recording. This article explains the complete pipeline implemented in the Zackriya-Solutions/meetily repository, from the frontend command to metadata persistence.

Entry Point: Triggering a Retranscription

The process begins when the frontend invokes the Tauri command start_retranscription_command. This command is registered in src-tauri/src/lib.rs at lines 742–744 and immediately delegates to the internal start_retranscription function.

// Frontend: start a retranscription with different settings
await invoke('start_retranscription', {
  meetingId: '12345',
  language: 'es',                 // Switch to Spanish
  model: 'medium',                // Upgrade from 'small' to 'medium' Whisper model
  provider: 'ollama'              // Use local Ollama instance instead of cloud
});

The command handler accepts optional parameters for language, model, and provider, allowing any combination of settings changes. If omitted, defaults fall back to system configuration.

Concurrency Control with Atomic Guards

Retranscription is resource-intensive. To prevent GPU/CPU thrashing, Meetily enforces single-instance execution using an atomic flag.

In src-tauri/src/audio/retranscription.rs (lines 19–26), a global RETRANSLATION_IN_PROGRESS AtomicBool guards the entry point. The flag is set inside a guard object that automatically clears on scope exit, ensuring cleanup even if the process panics.

// Simplified guard pattern from the source
let _guard = RetranslationGuard::new(); // Sets atomic flag
// ... retranscription runs ...
// Flag cleared automatically when _guard drops

Attempts to start a second retranscription while one is active return an immediate error: "Retranscription already in progress".

Parameter Validation and Engine Preparation

Before touching audio files, start_retranscription validates that the selected provider supports language-aware Whisper runs. At line 612 in retranscription.rs, unsupported configurations trigger a clear abort with descriptive error messaging.

Once validated, the system instantiates a fresh WhisperEngine via WhisperEngine::new with the requested model and provider. The engine lifecycle follows explicit load/unload semantics:

  • Loading: New engine initialization with custom settings
  • Processing: Batch transcription of all discovered audio
  • Unloading: unload_engine in src-tauri/src/audio/common.rs (line 17) frees memory and GPU context

This explicit cleanup prevents resource leaks when switching between providers with incompatible backends.

Audio Discovery and Decoding

Retranscription targets the meeting's original audio folder. The system enumerates all files matching extensions defined in audio/constants.rs (.wav, .mp3, .m4a, .webm, etc.).

Each file passes through the AudioDecoder in audio/decoder.rs, which:

  1. Demuxes container formats
  2. Resamples to Whisper's expected 16kHz mono PCM
  3. Returns a memory-efficient stream for chunked processing

This re‑decoding step is essential: stored audio is reconverted from its compressed format, not reused from any prior transcription cache.

The Transcription Loop with Progress Events

The core processing loop iterates over decoded audio chunks. For each segment:

// Inside run_retranscription – processing each audio file
for audio_path in audio_files {
    let pcm = AudioDecoder::decode(&audio_path).await?;
    let transcript = whisper_engine.transcribe(pcm, &language).await?;
    // Results accumulated and progress emitted...
}

Progress flows back to the UI through Tauri's event system. At line 510 of retranscription.rs, the "retranscription-progress" event carries percentage completion, enabling live progress bars without blocking the main thread.

Metadata Persistence and History Tracking

Successful completion updates the meeting's metadata.json via write_retranscription_metadata (lines 723–756). The function:

  • Preserves all existing fields (original transcription, timestamps, tags)
  • Adds retranscribed_at with ISO 8601 timestamp
  • Marks "source": "retranscription" to distinguish from live recordings

This append-only approach maintains full audit history. Users can compare original and retranscribed outputs without data destruction.

Completion and Error Handling

The pipeline reports final state through two events:

Event Trigger Payload
"retranscription-complete" All audio processed successfully Meeting ID, duration, new transcript path
"retranscription-error" Failure at any stage Error message, failed file (if applicable)

These events at lines 116 and 127 enable UI toasts, automatic retry prompts, or silent background retry logic.

Cancellation Support

Long retranscriptions can be interrupted. The cancel_retranscription_command (lines 822–827) flips the atomic progress flag, causing the background task to abort at the next chunk boundary. Resources are released through the normal guard cleanup path.

// Frontend: abort in-progress retranscription
await invoke('cancel_retranscription');

Summary

  • Single-entry command: start_retranscription_command in lib.rs routes frontend requests
  • Atomic concurrency: Guard pattern prevents overlapping runs in retranscription.rs
  • Fresh engine per job: Explicit load/unload in common.rs manages GPU/CPU resources
  • Re‑decode everything: AudioDecoder reprocesses stored files; no cached audio reuse
  • Event-driven progress: Live percentage updates via "retranscription-progress"
  • Non-destructive metadata: write_retranscription_metadata preserves history with retranscribed_at timestamp
  • Cancellable: Atomic flag check enables mid-operation abort

Frequently Asked Questions

What file formats support retranscription in Meetily?

All formats listed in src-tauri/src/audio/constants.rs are eligible: WAV, MP3, M4A, WEBM, OGG, and FLAC. The AudioDecoder handles format detection and resampling automatically, so retranscription works regardless of the original recording format.

Can I retranscribe with a different Whisper model size?

Yes. The model parameter accepts any valid Whisper model name (tiny, base, small, medium, large, large-v3, etc.). Meetily instantiates a fresh engine with the requested model, allowing quality upgrades or faster iterations with smaller variants.

Does retranscription overwrite my original transcript?

No. The original transcript file remains untouched. The new output is written alongside it, and metadata.json tracks both the original and retranscribed versions with timestamps and source markers. You can reference either version in the Meetily UI.

Why does retranscription fail with "provider does not support language selection"?

Not all Whisper providers implement language-aware transcription. The validation at line 612 of retranscription.rs checks provider capabilities before engine initialization. Switch to a compatible provider (OpenAI, local Whisper.cpp, or Ollama with proper configuration) or omit the language parameter to use the provider's auto-detection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →