How to Import and Enhance Existing Audio Files with Different Transcription Models in Meetily

Meetily allows you to import any external audio recording and iteratively improve transcription accuracy by switching between Whisper models or AI providers without re-uploading the original file.

Meetily is an open-source meeting transcription application by Zackriya-Solutions that supports importing legacy audio recordings and enhancing them through flexible transcription model selection. The workflow separates the initial import process from the transcription enhancement phase, storing audio metadata and voice activity data to enable model swapping without file re-processing. This architecture provides flexibility for users who need to balance processing speed against transcription accuracy based on audio quality or content requirements.

Import Workflow Architecture

The import pipeline consists of three coordinated stages that transform raw audio files into searchable meeting transcripts. Each stage emits progress events to keep the frontend synchronized with the background processing task.

File Selection and Validation

The import process begins when the frontend invokes the Tauri command start_import_audio_command registered in frontend/src-tauri/src/lib.rs (lines 746‑750). This command triggers a native file dialog and validates the selected file against the extension whitelist defined in audio/constants.rs. If the file cannot be decoded or uses an unsupported format, the system emits an "import-error" event to the UI; otherwise, it proceeds to the background processing stage.

Audio Processing and Transcription

The heavy lifting occurs in audio/import.rs (lines 254‑327), where the start_import function spawns an asynchronous run_import task. This task performs several critical operations:

  • Loads the selected audio file and resamples it to the pipeline's required 48 kHz sample rate
  • Runs Voice-Activity Detection (VAD) to isolate speech segments and remove silence
  • Invokes the Whisper engine via whisper_engine::load_model followed by transcribe using the user-selected model (e.g., base, small, medium) and provider (e.g., Ollama, OpenRouter)

The system supports multiple provider backends, allowing users to choose between local Whisper models or cloud-based LLM services depending on privacy requirements and computational resources. Throughout this stage, the backend emits "import-progress" and "import-warning" events to update the UI status bar.

Database Persistence

Transcription results and metadata are stored in a local SQLite database through the database/manager.rs API, specifically through functions like import_legacy_database. The import pipeline inserts meeting records, audio file metadata, and timestamped transcript segments into the database, establishing the foundation for subsequent retranscription operations. When processing completes, the system emits an "import-complete" event, making the meeting available for review.

Retranscription and Model Enhancement

After the initial import, users can enhance transcription quality by processing the same audio through different models using the retranscription system implemented in audio/retranscription.rs (line 954).

The retranscribe_meeting function clears the previous Whisper instance and loads a new model configuration without requiring access to the original audio file. It re-processes the stored VAD chunks from the initial import through the updated transcription pipeline and replaces the existing transcript in the database with the new results. Upon completion, the system emits a "retranscription-complete" event, automatically refreshing the meeting view with the enhanced text.

Implementation Examples

Below are practical TypeScript snippets demonstrating how to trigger the import workflow and subsequently request retranscription with a different model.

// Import an external audio file with initial transcription settings
await invoke('start_import_audio_command', {
  source_path: '/path/to/meeting.wav',
  title: 'Quarter‑Q Review',
  language: 'en',
  model: 'small',
  provider: 'openrouter'
});
// Enhance accuracy by retranscribing with a larger model
await invoke('retranscribe_meeting', {
  meeting_id: 42,
  new_model: 'medium',
  new_provider: 'openrouter'
});

Summary

  • Import Command: Use start_import_audio_command in frontend/src-tauri/src/lib.rs (lines 746‑750) to initiate file selection and validation against supported formats in audio/constants.rs.
  • Processing Pipeline: The run_import task in audio/import.rs (lines 254‑327) handles 48 kHz resampling, VAD segmentation, and initial Whisper transcription using the specified model and provider.
  • Persistent Storage: Meeting data and transcripts are stored in SQLite via database/manager.rs, preserving VAD chunks for future reprocessing.
  • Model Flexibility: The retranscribe_meeting function in audio/retranscription.rs (line 954) enables switching transcription models without re-uploading audio files, emitting "retranscription-complete" events to maintain UI synchronization.

Frequently Asked Questions

What audio file formats does Meetily support for import?

Meetily validates imported files against a whitelist of supported extensions defined in audio/constants.rs. The system checks file extensions and attempts to decode audio streams before processing, emitting "import-error" events for unsupported formats.

Can I switch between local and cloud-based transcription models after importing?

Yes. The retranscription system in audio/retranscription.rs allows you to switch between any supported provider (including local Whisper models via Ollama or cloud services via OpenRouter) and model sizes (base, small, medium) without re-uploading the original audio file.

How does Meetily handle audio quality variations during import?

During the import pipeline in audio/import.rs, the system resamples all audio to 48 kHz and applies Voice-Activity Detection (VAD) to isolate speech segments. This preprocessing standardizes input quality and removes non-speech segments before transcription, improving model accuracy across varying source qualities.

Where is the transcription data stored after importing?

All meeting metadata, audio file references, and transcript segments are persisted to a local SQLite database through the database/manager.rs API. This local-first architecture ensures privacy and enables offline retranscription capabilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →