# How Meetily's Retranscription Feature Re-processes Stored Audio with Different Models or Languages

> Learn how Meetily's retranscription feature re-processes stored audio using new models or languages. Discover its advanced decoding, VAD, and transcription engine capabilities.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: internals
- Published: 2026-07-31

---

**Meetily's retranscription feature re-processes stored meeting audio by decoding the original recording, re-running Voice Activity Detection with batch-optimized settings, and initializing the requested transcription engine with new model and language parameters to generate an updated transcript.**

Meetily is an open-source meeting assistant that archives raw audio recordings for every session. The retranscription feature allows users to re-process these stored files using different AI models or language settings without re-recording, providing flexibility to improve accuracy or adapt content for different languages post-meeting.

## The Retranscription Pipeline Architecture

When a user initiates a retranscription, the application executes a dedicated batch pipeline that operates independently from real-time transcription. This pipeline handles audio discovery, decoding, segmentation, and transcription engine management as atomic operations.

### Locating and Decoding Stored Audio

The pipeline begins by locating the raw audio file within the meeting folder. In [`frontend/src-tauri/src/audio/retranscription.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/retranscription.rs), the `find_audio_file` function (lines 39-66) scans for common filenames and validates extensions against the `AUDIO_EXTENSIONS` constant to identify the source file.

Once located, the system decodes the audio using `decode_audio_file` and normalizes it to 16 kHz mono format—the standard required by both Whisper and Parakeet models. According to the Meetily source code, both decoding and normalization execute inside blocking tasks (lines 99-126) to prevent blocking the async runtime during CPU-intensive operations.

### Voice Activity Detection with Batch Optimizations

After decoding, the system runs Voice Activity Detection (VAD) using `get_speech_chunks_with_progress`. Unlike real-time processing, batch retranscription configures a **larger redemption time of 2 seconds** (lines 48-52 and 35-38) to reduce fragmentation in pre-recorded audio. Progress updates emit via the `retranscription-progress` Tauri event (lines 46-58), allowing the frontend to display granular status updates during this phase.

## Model and Language Selection Logic

The retranscription pipeline dynamically selects and initializes transcription engines based on user-provided parameters, automatically handling model swaps when the requested configuration differs from cached instances.

### Engine Initialization and Model Swapping

In [`frontend/src-tauri/src/audio/retranscription.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/retranscription.rs) (lines 84-91), the provider parameter determines the backend: `"parakeet"` triggers the `ParakeetEngine`, while any other value defaults to the native `WhisperEngine`. The system initializes the selected engine **only once per job** using helper functions `get_or_init_whisper` or `get_or_init_parakeet` (lines 20-75 and 77-88).

These helpers check whether the requested model is already loaded in memory. If the user specifies a different model via the `model` argument, or if no model is specified and the system falls back to the value stored in the `transcript_settings` table, the engine loads the new weights before processing begins. This guarantees consistent model usage across the entire batch job.

### Language Override Handling

When using the Whisper backend, the pipeline supports explicit language overrides. The function call `engine.transcribe_audio_with_confidence(..., language.clone())` (lines 79-84) passes the optional `language` parameter directly to the Whisper inference engine. This allows users to force specific language codes—such as `"es"` for Spanish—regardless of the language detected during the original recording.

## Segment Processing and Quality Optimizations

Long audio segments degrade transcription accuracy, so the pipeline implements intelligent splitting mechanisms. Segments exceeding 25 seconds are automatically divided at silence points using `split_segment_at_silence` (lines 13-16 and 18-33). This splitting occurs before the main transcription loop.

The transcription loop (lines 42-70) iterates through all speech chunks, checking a cancellation flag before each segment to support abort operations. Results accumulate in the `all_transcripts` vector, preserving the temporal order of speech segments.

## Database Updates and Metadata Management

Upon completion, the system writes results within a database transaction to ensure atomicity. The `create_transcript_segments` function (lines 21-28 and 30-57) inserts new transcript rows, replacing any previous transcripts for that meeting ID.

Simultaneously, `write_retranscription_metadata` (lines 23-30 and 33-44) updates the meeting's [`metadata.json`](https://github.com/Zackriya-Solutions/meetily/blob/main/metadata.json) file with a new `retranscribed_at` timestamp, sets the status to `completed`, and removes the stale `detected_summary_language` field to prevent language model mismatches in downstream processing.

## Frontend Integration and Progress Events

The Rust core communicates with the frontend through Tauri events throughout the pipeline lifecycle. The system emits `retranscription-progress` during VAD and transcription, `retranscription-complete` upon success, and `retranscription-error` if failures occur (lines 15-24 and 91-98).

To implement retranscription in a client application:

```typescript
import { invoke } from '@tauri-apps/api/tauri';
import { listen } from '@tauri-apps/api/event';

// Initiate retranscription with specific model and language
async function startRetranscription(meetingId: string, folderPath: string) {
  await invoke('start_retranscription_command', {
    meeting_id: meetingId,
    meeting_folder_path: folderPath,
    language: 'es',               // Force Spanish language
    model: 'medium-v3',           // Use larger Whisper model
    provider: 'whisper',          // or 'parakeet'
  });
}

// Monitor progress updates
listen('retranscription-progress', (event) => {
  const p = event.payload as {
    meeting_id: string;
    stage: string;
    progress_percentage: number;
    message: string;
  };
  console.log(`[${p.stage}] ${p.progress_percentage}% – ${p.message}`);
});

```

The Rust command handler spawns the pipeline as a background task to keep the UI responsive:

```rust
#[tauri::command]
pub async fn start_retranscription_command<R: Runtime>(
    // args omitted for brevity
) -> Result<RetranscriptionStarted, String> {
    tauri::async_runtime::spawn(async move {
        let _ = start_retranscription(
            app, meeting_id, folder, language, model, provider
        ).await;
    });
    Ok(RetranscriptionStarted { 
        meeting_id, 
        message: "Retranscription started".into() 
    })
}

```

## Summary

- **Audio Discovery**: `find_audio_file` locates recordings by scanning standard filenames and supported extensions in the meeting folder.
- **Format Normalization**: The pipeline decodes audio to 16 kHz mono in blocking tasks to satisfy Whisper/Parakeet requirements.
- **Batch VAD**: Uses extended 2-second redemption time and emits progress events via Tauri for UI feedback.
- **Dynamic Engine Loading**: `get_or_init_whisper` and `get_or_init_parakeet` swap models automatically when user requests differ from cached instances.
- **Language Overrides**: Supports explicit language codes passed directly to the transcription engine during batch processing.
- **Segment Optimization**: Automatically splits segments exceeding 25 seconds at silence points to maintain accuracy.
- **Atomic Updates**: Database transactions ensure transcript replacement is all-or-nothing, with metadata timestamps tracking retranscription history.

## Frequently Asked Questions

### What audio formats does Meetily support for retranscription?

Meetily supports any audio format whose extension appears in the `AUDIO_EXTENSIONS` constant defined in the configuration. The `find_audio_file` function scans the meeting folder for common recording filenames and validates extensions before processing, ensuring compatibility with standard formats like MP4, MP3, WAV, and M4A.

### Can I switch between Whisper and Parakeet during retranscription?

Yes. The `provider` parameter in `start_retranscription_command` accepts `"parakeet"` to use the Parakeet engine or defaults to Whisper for any other value. The system initializes the selected engine once at the start of the job and uses it consistently for all segments in that batch.

### How does Meetily handle long audio segments during retranscription?

Segments longer than 25 seconds are automatically split at detected silence points using `split_segment_at_silence` before transcription begins. This prevents accuracy degradation that occurs when models process excessively long audio chunks, ensuring consistent quality across the entire recording.

### Will retranscription overwrite my original transcript?

Yes. The system executes `create_transcript_segments` within a database transaction that replaces previous transcript entries for the meeting ID. The `write_retranscription_metadata` function updates the [`metadata.json`](https://github.com/Zackriya-Solutions/meetily/blob/main/metadata.json) file with a new `retranscribed_at` timestamp, effectively versioning the transcript history while removing stale language detection fields.