How Meetily's Retranscription Feature Re-processes Stored Audio with Different Models or Languages
Meetily's retranscription feature re-processes stored meeting audio by decoding the original recording, re-running Voice Activity Detection with batch-optimized settings, and initializing the requested transcription engine with new model and language parameters to generate an updated transcript.
Meetily is an open-source meeting assistant that archives raw audio recordings for every session. The retranscription feature allows users to re-process these stored files using different AI models or language settings without re-recording, providing flexibility to improve accuracy or adapt content for different languages post-meeting.
The Retranscription Pipeline Architecture
When a user initiates a retranscription, the application executes a dedicated batch pipeline that operates independently from real-time transcription. This pipeline handles audio discovery, decoding, segmentation, and transcription engine management as atomic operations.
Locating and Decoding Stored Audio
The pipeline begins by locating the raw audio file within the meeting folder. In frontend/src-tauri/src/audio/retranscription.rs, the find_audio_file function (lines 39-66) scans for common filenames and validates extensions against the AUDIO_EXTENSIONS constant to identify the source file.
Once located, the system decodes the audio using decode_audio_file and normalizes it to 16 kHz mono format—the standard required by both Whisper and Parakeet models. According to the Meetily source code, both decoding and normalization execute inside blocking tasks (lines 99-126) to prevent blocking the async runtime during CPU-intensive operations.
Voice Activity Detection with Batch Optimizations
After decoding, the system runs Voice Activity Detection (VAD) using get_speech_chunks_with_progress. Unlike real-time processing, batch retranscription configures a larger redemption time of 2 seconds (lines 48-52 and 35-38) to reduce fragmentation in pre-recorded audio. Progress updates emit via the retranscription-progress Tauri event (lines 46-58), allowing the frontend to display granular status updates during this phase.
Model and Language Selection Logic
The retranscription pipeline dynamically selects and initializes transcription engines based on user-provided parameters, automatically handling model swaps when the requested configuration differs from cached instances.
Engine Initialization and Model Swapping
In frontend/src-tauri/src/audio/retranscription.rs (lines 84-91), the provider parameter determines the backend: "parakeet" triggers the ParakeetEngine, while any other value defaults to the native WhisperEngine. The system initializes the selected engine only once per job using helper functions get_or_init_whisper or get_or_init_parakeet (lines 20-75 and 77-88).
These helpers check whether the requested model is already loaded in memory. If the user specifies a different model via the model argument, or if no model is specified and the system falls back to the value stored in the transcript_settings table, the engine loads the new weights before processing begins. This guarantees consistent model usage across the entire batch job.
Language Override Handling
When using the Whisper backend, the pipeline supports explicit language overrides. The function call engine.transcribe_audio_with_confidence(..., language.clone()) (lines 79-84) passes the optional language parameter directly to the Whisper inference engine. This allows users to force specific language codes—such as "es" for Spanish—regardless of the language detected during the original recording.
Segment Processing and Quality Optimizations
Long audio segments degrade transcription accuracy, so the pipeline implements intelligent splitting mechanisms. Segments exceeding 25 seconds are automatically divided at silence points using split_segment_at_silence (lines 13-16 and 18-33). This splitting occurs before the main transcription loop.
The transcription loop (lines 42-70) iterates through all speech chunks, checking a cancellation flag before each segment to support abort operations. Results accumulate in the all_transcripts vector, preserving the temporal order of speech segments.
Database Updates and Metadata Management
Upon completion, the system writes results within a database transaction to ensure atomicity. The create_transcript_segments function (lines 21-28 and 30-57) inserts new transcript rows, replacing any previous transcripts for that meeting ID.
Simultaneously, write_retranscription_metadata (lines 23-30 and 33-44) updates the meeting's metadata.json file with a new retranscribed_at timestamp, sets the status to completed, and removes the stale detected_summary_language field to prevent language model mismatches in downstream processing.
Frontend Integration and Progress Events
The Rust core communicates with the frontend through Tauri events throughout the pipeline lifecycle. The system emits retranscription-progress during VAD and transcription, retranscription-complete upon success, and retranscription-error if failures occur (lines 15-24 and 91-98).
To implement retranscription in a client application:
import { invoke } from '@tauri-apps/api/tauri';
import { listen } from '@tauri-apps/api/event';
// Initiate retranscription with specific model and language
async function startRetranscription(meetingId: string, folderPath: string) {
await invoke('start_retranscription_command', {
meeting_id: meetingId,
meeting_folder_path: folderPath,
language: 'es', // Force Spanish language
model: 'medium-v3', // Use larger Whisper model
provider: 'whisper', // or 'parakeet'
});
}
// Monitor progress updates
listen('retranscription-progress', (event) => {
const p = event.payload as {
meeting_id: string;
stage: string;
progress_percentage: number;
message: string;
};
console.log(`[${p.stage}] ${p.progress_percentage}% – ${p.message}`);
});
The Rust command handler spawns the pipeline as a background task to keep the UI responsive:
#[tauri::command]
pub async fn start_retranscription_command<R: Runtime>(
// args omitted for brevity
) -> Result<RetranscriptionStarted, String> {
tauri::async_runtime::spawn(async move {
let _ = start_retranscription(
app, meeting_id, folder, language, model, provider
).await;
});
Ok(RetranscriptionStarted {
meeting_id,
message: "Retranscription started".into()
})
}
Summary
- Audio Discovery:
find_audio_filelocates recordings by scanning standard filenames and supported extensions in the meeting folder. - Format Normalization: The pipeline decodes audio to 16 kHz mono in blocking tasks to satisfy Whisper/Parakeet requirements.
- Batch VAD: Uses extended 2-second redemption time and emits progress events via Tauri for UI feedback.
- Dynamic Engine Loading:
get_or_init_whisperandget_or_init_parakeetswap models automatically when user requests differ from cached instances. - Language Overrides: Supports explicit language codes passed directly to the transcription engine during batch processing.
- Segment Optimization: Automatically splits segments exceeding 25 seconds at silence points to maintain accuracy.
- Atomic Updates: Database transactions ensure transcript replacement is all-or-nothing, with metadata timestamps tracking retranscription history.
Frequently Asked Questions
What audio formats does Meetily support for retranscription?
Meetily supports any audio format whose extension appears in the AUDIO_EXTENSIONS constant defined in the configuration. The find_audio_file function scans the meeting folder for common recording filenames and validates extensions before processing, ensuring compatibility with standard formats like MP4, MP3, WAV, and M4A.
Can I switch between Whisper and Parakeet during retranscription?
Yes. The provider parameter in start_retranscription_command accepts "parakeet" to use the Parakeet engine or defaults to Whisper for any other value. The system initializes the selected engine once at the start of the job and uses it consistently for all segments in that batch.
How does Meetily handle long audio segments during retranscription?
Segments longer than 25 seconds are automatically split at detected silence points using split_segment_at_silence before transcription begins. This prevents accuracy degradation that occurs when models process excessively long audio chunks, ensuring consistent quality across the entire recording.
Will retranscription overwrite my original transcript?
Yes. The system executes create_transcript_segments within a database transaction that replaces previous transcript entries for the meeting ID. The write_retranscription_metadata function updates the metadata.json file with a new retranscribed_at timestamp, effectively versioning the transcript history while removing stale language detection fields.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →