Meetily Retranscription Workflow: How to Reprocess Recordings with Different Models and Languages
Meetily's retranscription workflow allows you to re-transcribe existing recordings with different languages, models, or providers (Whisper or Parakeet) without re-recording, implemented in frontend/src-tauri/src/audio/retranscription.rs as a cancellable background pipeline.
The retranscription system in the Zackriya-Solutions/meetily repository enables users to reprocess meeting recordings using alternative settings while preserving the original audio. This Rust-based pipeline runs as a background async task within the Tauri application framework, handling everything from audio discovery to database updates.
How Meetily's Retranscription Pipeline Works
The workflow implemented in frontend/src-tauri/src/audio/retranscription.rs processes recordings through several distinct stages, each designed to ensure high-quality transcription results while maintaining system responsiveness.
Guard and Concurrency Control
To prevent resource conflicts, the pipeline begins by acquiring a RetranscriptionGuard. This structure checks the RETRANSCIPTION_IN_PROGRESS atomic flag to guarantee only one retranscription runs simultaneously. If another process is active, the guard returns an error immediately. The system also maintains a RETRANSCIPTION_CANCELLED flag that allows users to abort long-running operations gracefully.
Audio Discovery and Decoding
The system locates audio files using the find_audio_file function, which first searches for common filenames then falls back to scanning for any supported extension defined in constants::AUDIO_EXTENSIONS. Once found, the file undergoes blocking decode operations via decode_audio_file followed by format conversion through to_whisper_format, producing a 16 kHz mono PCM buffer suitable for transcription engines.
Voice Activity Detection (VAD)
The decoded audio passes through get_speech_chunks_with_progress with a 2000 ms redemption time (VAD_REDEMPTION_TIME_MS), bridging natural pauses to create continuous speech segments. For segments exceeding 25 seconds (MAX_SEGMENT_SAMPLES), the pipeline applies split_segment_at_silence to divide them at the lowest-energy silence window, optimizing transcription accuracy for long utterances.
Engine Initialization and Transcription
Based on the provider parameter, the system lazily initializes either Whisper via get_or_init_whisper or Parakeet via get_or_init_parakeet. The main transcription loop iterates over speech segments, skipping any chunks shorter than 100 ms, and calls transcribe_audio_with_confidence for Whisper or transcribe_audio for Parakeet. Results aggregate into a complete transcript with confidence scores and timestamps.
Data Persistence and Progress Reporting
Upon completion, the system executes a database transaction that deletes existing transcripts and inserts new rows in a single atomic operation. It then exports results to transcripts.json and updates metadata.json with retranscribed_at timestamps and status flags via write_transcripts_json and write_retranscription_metadata. Throughout execution, the pipeline emits Tauri events including retranscription-progress, retranscription-complete, and retranscription-error via the emit_progress helper.
Triggering Retranscription from the Frontend
Invoke the retranscription process using Tauri's command system to specify alternative languages, models, or providers:
import { invoke } from '@tauri-apps/api/tauri';
// Start a retranscription (e.g., switch to a different language or model)
await invoke('start_retranscription_command', {
meeting_id: '12345',
meeting_folder_path: '/Users/alice/Meetily/meetings/12345',
language: 'es', // optional ISO language code
model: 'medium', // optional model name
provider: 'localWhisper' // "localWhisper" | "parakeet"
});
Monitoring Progress and Handling Completion
Listen to real-time updates using Tauri's event system:
import { listen } from '@tauri-apps/api/event';
listen('retranscription-progress', event => {
const { meeting_id, stage, progress_percentage, message } = event.payload;
console.log(`[${meeting_id}] ${stage}: ${progress_percentage}% – ${message}`);
});
listen('retranscription-complete', event => {
console.log('Retranscription finished:', event.payload);
});
listen('retranscription-error', event => {
console.error('Retranscription failed:', event.payload);
});
Cancelling Ongoing Retranscriptions
To abort a running retranscription job:
await invoke('cancel_retranscription_command');
This sets the RETRANSCIPTION_CANCELLED atomic flag, causing the pipeline to terminate at the next checkpoint.
Key Implementation Files
The retranscription workflow spans several modules in the Meetily codebase:
frontend/src-tauri/src/audio/retranscription.rs– Main pipeline implementation including guard handling, VAD, engine initialization, and database writes.frontend/src-tauri/src/audio/vad.rs– Speech activity detection with redemption time configuration.frontend/src-tauri/src/whisper_engine/whisper_engine.rs– Whisper implementation withtranscribe_audio_with_confidence.frontend/src-tauri/src/parakeet_engine/parakeet_engine.rs– Parakeet implementation withtranscribe_audio.frontend/src-tauri/src/database/repositories/transcript.rs– Database schema and operations for transcript storage.frontend/src-tauri/src/database/manager.rs– Connection pooling and transaction management.frontend/src-tauri/src/config.rs– Default model constants and configuration defaults.
Summary
- Meetily's retranscription workflow reprocesses existing recordings without requiring new recordings, supporting language changes and model switches.
- Concurrency protection via
RetranscriptionGuardensures only one retranscription runs at a time with cancellable flags. - Audio processing includes 16 kHz mono conversion, VAD with 2000 ms redemption time, and automatic splitting of segments exceeding 25 seconds.
- Dual engine support allows choosing between Whisper (
transcribe_audio_with_confidence) and Parakeet (transcribe_audio) providers. - Atomic updates replace old transcripts in a single database transaction while exporting JSON metadata.
- Real-time communication through Tauri events provides progress updates, completion signals, and error handling.
Frequently Asked Questions
Can I switch transcription models when reprocessing a recording?
Yes, the start_retranscription_command accepts a model parameter (e.g., "medium", "large") and a provider parameter ("localWhisper" or "parakeet"). The system lazily loads the requested engine via get_or_init_whisper or get_or_init_parakeet, allowing you to reprocess the same audio with different models without restarting the application.
How does Meetily handle cancellation during retranscription?
The pipeline checks the RETRANSCIPTION_CANCELLED atomic flag at key checkpoints. Calling cancel_retranscription_command sets this flag, causing the background task to terminate gracefully. The RetranscriptionGuard ensures cleanup of the RETRANSCIPTION_IN_PROGRESS flag regardless of completion status.
What audio formats are supported for retranscription?
The find_audio_file function searches for extensions defined in constants::AUDIO_EXTENSIONS. While specific formats depend on the system's audio decoder capabilities, the pipeline attempts common meeting audio extensions first, then falls back to any supported format found in the meeting folder.
Where are retranscription results stored?
Results persist in three locations: the SQLite database (via atomic transaction in transcript.rs), a JSON export at transcripts.json (via write_transcripts_json), and updated metadata in metadata.json (via write_retranscription_metadata) including the retranscribed_at timestamp and completion status.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →