How Meetily Imports External Audio Files for Transcription: A Complete Technical Breakdown

Meetily processes external audio files through a 10-step local pipeline that validates, decodes, resamples, runs voice-activity detection, and transcribes using Whisper or Parakeet—entirely offline with cancellable progress tracking.

Meetily's import feature transforms any supported audio file into a fully timestamped meeting transcript without ever transmitting data to remote servers. This article examines the complete technical implementation in the Zackriya-Solutions/meetily repository, walking through each stage from file selection to database persistence.

Overview of the Import Pipeline

The entire workflow is orchestrated from frontend/src-tauri/src/audio/import.rs. The pipeline is asynchronous, cancellable, and progress-aware, emitting real-time events to the React frontend at every stage.

Key architectural decisions that define the system:

  • Zero network dependency — all processing happens locally via Rust/Tauri
  • Atomic guard pattern — prevents overlapping imports via ImportGuard
  • Speech-aware segmentation — VAD with 2000 ms redemption time removes silence efficiently
  • Dual-engine support — transparently switches between Whisper and Parakeet transcription backends

Step 1: File Selection and Validation

The import begins when the frontend invokes select_and_validate_audio_command. This triggers validate_audio_file (lines 19–90), which enforces strict entry criteria:

  • File must exist and be readable
  • Extension must match AUDIO_EXTENSIONS whitelist
  • Size must not exceed 20 GB
  • Metadata extraction attempts fast path; falls back to full decode if needed

Upon success, the function returns an AudioFileInfo struct containing title, duration, and format—data that populates the UI's import dialog.

import { invoke } from '@tauri-apps/api/tauri';

async function pickAudioFile() {
  const info = await invoke('select_and_validate_audio_command');
  if (info) {
    console.log('Picked:', info.filename, info.duration_seconds);
  }
}

Step 2: Guard Initialization and Cancellation Setup

Before any heavy processing, ImportGuard::acquire atomically sets the global IMPORT_IN_PROGRESS flag. A companion flag, IMPORT_CANCELLED, enables UI-initiated aborts at any stage.

Both flags clear automatically via Drop implementation (lines 31–52), ensuring clean state even on panic. This prevents resource leaks and guarantees that a failed import never blocks future attempts.

Step 3: Meeting Folder Creation and File Copying

The run_import function establishes isolation:

  1. Creates a dedicated folder via create_meeting_folder (lines 41–45)
  2. Copies the original audio using std::fs::copy (lines 46–63)

This preserves the user's original file while giving the pipeline a stable, versioned workspace. The meeting folder becomes the permanent home for transcripts, metadata, and derived assets.

Step 4: Audio Decoding with Progress Streaming

The copied file feeds into decode_audio_file_with_progress, defined in frontend/src-tauri/src/audio/decoder.rs. This leverages the symphonia crate to decode diverse container formats into raw PCM.

The function streams audio while emitting progress events, yielding a DecodedAudio struct containing:

  • Interleaved PCM samples
  • Original sample rate and channel count
  • Precise duration

Step 5: Resampling to Whisper-Compatible Format

Raw decoded audio rarely matches the transcription engine's requirements. The pipeline converts to 16 kHz mono PCM via to_whisper_format_with_progress (lines 104–116 in import.rs).

This standardized format is mandatory for both Whisper and Parakeet backends, eliminating downstream compatibility variance.

Step 6: Voice-Activity Detection (VAD)

The resampled audio flows into get_speech_chunks_with_progress from frontend/src-tauri/src/audio/vad.rs. Key parameters distinguish import processing from live recording:

Aspect Import Pipeline Live Pipeline
Redemption time 2000 ms 400 ms
Purpose Aggressive silence removal Responsive real-time chunking

VAD segments the audio into speech chunks with precise start/end timestamps. This reduces LLM load dramatically by discarding silent periods rather than transcribing them.

Step 7: Optional Segment Splitting

Long speech chunks exceeding 25 seconds undergo additional processing. The split_segment_at_silence function (lines 222–242) fractures these at natural silence boundaries.

This prevents word loss at hard cut points and keeps segments within optimal size ranges for the transcription engine's context window.

Step 8: Engine Selection and Lazy Loading

The provider argument ("parakeet" or default Whisper) determines backend selection. The engine initialization logic (lines 260–331) implements efficient resource management:

  • Engines load lazily via get_or_init_whisper or get_or_init_parakeet
  • If the currently resident model differs from the request, load_model swaps it
  • Model choice respects the database configuration via get_configured_model

This supports runtime switching without application restart.

Step 9: Transcription Loop with Confidence Scoring

Each speech segment enters the transcription loop (lines 254–311):

// Pseudoglimpse of the core loop logic
for chunk in speech_chunks {
    if chunk.samples.len() < 1600 { continue; } // Skip < 0.1s
    
    let (text, confidence) = match provider {
        "parakeet" => engine.transcribe(&chunk), // Fixed confidence
        _ => whisper.transcribe_with_confidence(&chunk), // Scored output
    };
    
    transcripts.push((text, chunk.start_ts, chunk.end_ts));
    emit_progress(&app, stage, percentage, message);
}

The loop accumulates (text, start_ts, end_ts) tuples while tracking aggregate confidence metrics. Progress events stream to the frontend after each segment.

Step 10: Meeting Persistence and Metadata Generation

Transcription completion triggers database persistence (lines 340–418). The pipeline:

  1. Converts raw tuples to TranscriptSegment structs via create_transcript_segments
  2. Inserts the meeting and transcripts into SQLite via create_meeting_with_transcripts from frontend/src-tauri/src/database/manager.rs
  3. Writes transcripts.json and canonical metadata.json to the meeting folder (lines 778–808 in frontend/src-tauri/src/audio/common.rs)

These files enable UI navigation, export functionality, and future reprocessing without retranscribing.

Event-Driven Frontend Integration

The React frontend communicates with the Rust backend through Tauri commands and event listeners:

import { invoke } from '@tauri-apps/api/tauri';
import { listen } from '@tauri-apps/api/event';

// Initiate import with full configuration
async function startImport(filePath: string, title: string) {
  await invoke('start_import_audio_command', {
    source_path: filePath,
    title,
    language: 'en',
    model: undefined,      // Use configured default
    provider: 'whisper',   // or 'parakeet'
  });
}

// Real-time progress updates
listen('import-progress', (e) => {
  const { stage, progress_percentage, message } = e.payload;
  console.log(`[${stage}] ${progress_percentage}% – ${message}`);
});

Cancellation is equally straightforward:

async function cancelImport() {
  await invoke('cancel_import_command');
}

Event types emitted throughout the pipeline:

  • import-progress — stage transitions and percentage completion
  • import-warning — non-fatal issues (unsupported metadata, decode edge cases)
  • import-complete — final success with meeting ID
  • import-error — terminal failure with structured error message

Key Files in the Import Architecture

File Responsibility
frontend/src-tauri/src/audio/import.rs Core pipeline orchestration
frontend/src-tauri/src/audio/decoder.rs Container decoding via symphonia
frontend/src-tauri/src/audio/vad.rs Speech boundary detection
frontend/src-tauri/src/whisper_engine/whisper_engine.rs Whisper model management and inference
frontend/src-tauri/src/parakeet_engine/parakeet_engine.rs Alternative Parakeet backend
frontend/src-tauri/src/database/manager.rs SQLite persistence layer
frontend/src-tauri/src/audio/common.rs Folder creation, transcript formatting, metadata I/O
frontend/src-tauri/src/api/commands.rs Tauri command surface for frontend binding

Performance and Scalability Characteristics

The import pipeline exhibits several optimization patterns worth noting:

  • Memory efficiency — Streaming decode and VAD process audio in chunks rather than loading entire files
  • CPU parallelism — Resampling and V leverage available cores where symphonia and internal algorithms permit
  • Disk isolation — Per-meeting folders prevent I/O contention and simplify cleanup
  • Graceful degradation — Metadata extraction failures fall back to full decode without user intervention

The 20 GB file size limit balances practical use cases (multi-hour recordings) against memory and disk constraints on typical workstations.

Summary

  • Meetily's import feature is implemented entirely in frontend/src-tauri/src/audio/import.rs as a 10-stage asynchronous pipeline
  • Validation enforces format, size, and integrity constraints before any heavy processing begins
  • Decoding and resampling standardize arbitrary audio to 16 kHz mono PCM required by transcription engines
  • VAD with 2000 ms redemption time aggressively removes silence, reducing LLM computation vs. live recording's 400 ms responsiveness
  • Dual-engine support allows runtime selection between Whisper and Parakeet backends with lazy model loading
  • Progress events and cancellation flags enable responsive UI feedback without blocking the main thread
  • Atomic guard patterns prevent overlapping imports and ensure clean state recovery
  • Complete local processing guarantees privacy—no audio data ever leaves the user's machine

Frequently Asked Questions

What audio formats can Meetily import?

Meetily accepts any format supported by the symphonia crate, which includes WAV, MP3, FLAC, OGG, AAC, and common container formats. The AUDIO_EXTENSIONS whitelist in validate_audio_file may restrict this further based on confidence in metadata extraction reliability.

How does Meetily handle very large audio files?

Files up to 20 GB are supported through streaming decode that processes audio in chunks rather than loading everything into memory. The pipeline copies the file to a dedicated meeting folder first, ensuring the original remains untouched while providing stable random access for the decoder.

Can I switch transcription engines mid-import?

No—engine selection occurs at import initiation via the provider parameter. However, the lazy-loading architecture in get_or_init_whisper and get_or_init_parakeet means switching engines between imports requires no application restart, and models are held resident to accelerate subsequent usage of the same backend.

What happens if I cancel an import in progress?

The IMPORT_CANCELLED flag is set atomically, and the pipeline checks this flag between major stages (validation, decode chunks, transcription segments). Partial progress is discarded; the ImportGuard ensures flags clear automatically. The frontend receives a cancellation confirmation event, and the partial meeting folder may be cleaned up or retained for debugging based on build configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →