# How Meetily Imports External Audio Files for Transcription: A Complete Technical Breakdown

> Explore Meetily's 10-step offline pipeline for importing external audio files. Learn how it validates, decodes, resamples, and transcribes using Whisper or Parakeet.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: how-to-guide
- Published: 2026-08-01

---

**Meetily processes external audio files through a 10-step local pipeline that validates, decodes, resamples, runs voice-activity detection, and transcribes using Whisper or Parakeet—entirely offline with cancellable progress tracking.**

Meetily's import feature transforms any supported audio file into a fully timestamped meeting transcript without ever transmitting data to remote servers. This article examines the complete technical implementation in the Zackriya-Solutions/meetily repository, walking through each stage from file selection to database persistence.

## Overview of the Import Pipeline

The entire workflow is orchestrated from **[`frontend/src-tauri/src/audio/import.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/import.rs)**. The pipeline is **asynchronous**, **cancellable**, and **progress-aware**, emitting real-time events to the React frontend at every stage.

Key architectural decisions that define the system:

- **Zero network dependency** — all processing happens locally via Rust/Tauri
- **Atomic guard pattern** — prevents overlapping imports via `ImportGuard`
- **Speech-aware segmentation** — VAD with 2000 ms redemption time removes silence efficiently
- **Dual-engine support** — transparently switches between Whisper and Parakeet transcription backends

## Step 1: File Selection and Validation

The import begins when the frontend invokes `select_and_validate_audio_command`. This triggers `validate_audio_file` (lines 19–90), which enforces strict entry criteria:

- File must exist and be readable
- Extension must match `AUDIO_EXTENSIONS` whitelist
- Size must not exceed **20 GB**
- Metadata extraction attempts fast path; falls back to full decode if needed

Upon success, the function returns an `AudioFileInfo` struct containing title, duration, and format—data that populates the UI's import dialog.

```tsx
import { invoke } from '@tauri-apps/api/tauri';

async function pickAudioFile() {
  const info = await invoke('select_and_validate_audio_command');
  if (info) {
    console.log('Picked:', info.filename, info.duration_seconds);
  }
}

```

## Step 2: Guard Initialization and Cancellation Setup

Before any heavy processing, `ImportGuard::acquire` atomically sets the global `IMPORT_IN_PROGRESS` flag. A companion flag, `IMPORT_CANCELLED`, enables UI-initiated aborts at any stage.

Both flags clear automatically via `Drop` implementation (lines 31–52), ensuring clean state even on panic. This prevents resource leaks and guarantees that a failed import never blocks future attempts.

## Step 3: Meeting Folder Creation and File Copying

The `run_import` function establishes isolation:

1. Creates a dedicated folder via `create_meeting_folder` (lines 41–45)
2. Copies the original audio using `std::fs::copy` (lines 46–63)

This preserves the user's original file while giving the pipeline a stable, versioned workspace. The meeting folder becomes the permanent home for transcripts, metadata, and derived assets.

## Step 4: Audio Decoding with Progress Streaming

The copied file feeds into `decode_audio_file_with_progress`, defined in **[`frontend/src-tauri/src/audio/decoder.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/decoder.rs)**. This leverages the `symphonia` crate to decode diverse container formats into raw PCM.

The function streams audio while emitting progress events, yielding a `DecodedAudio` struct containing:
- Interleaved PCM samples
- Original sample rate and channel count
- Precise duration

## Step 5: Resampling to Whisper-Compatible Format

Raw decoded audio rarely matches the transcription engine's requirements. The pipeline converts to **16 kHz mono PCM** via `to_whisper_format_with_progress` (lines 104–116 in [`import.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/import.rs)).

This standardized format is mandatory for both Whisper and Parakeet backends, eliminating downstream compatibility variance.

## Step 6: Voice-Activity Detection (VAD)

The resampled audio flows into `get_speech_chunks_with_progress` from **[`frontend/src-tauri/src/audio/vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/vad.rs)**. Key parameters distinguish import processing from live recording:

| Aspect | Import Pipeline | Live Pipeline |
|--------|---------------|---------------|
| Redemption time | **2000 ms** | 400 ms |
| Purpose | Aggressive silence removal | Responsive real-time chunking |

VAD segments the audio into speech chunks with precise start/end timestamps. This **reduces LLM load dramatically** by discarding silent periods rather than transcribing them.

## Step 7: Optional Segment Splitting

Long speech chunks exceeding **25 seconds** undergo additional processing. The `split_segment_at_silence` function (lines 222–242) fractures these at natural silence boundaries.

This prevents word loss at hard cut points and keeps segments within optimal size ranges for the transcription engine's context window.

## Step 8: Engine Selection and Lazy Loading

The `provider` argument (`"parakeet"` or default Whisper) determines backend selection. The engine initialization logic (lines 260–331) implements efficient resource management:

- Engines load lazily via `get_or_init_whisper` or `get_or_init_parakeet`
- If the currently resident model differs from the request, `load_model` swaps it
- Model choice respects the database configuration via `get_configured_model`

This supports runtime switching without application restart.

## Step 9: Transcription Loop with Confidence Scoring

Each speech segment enters the transcription loop (lines 254–311):

```rust
// Pseudoglimpse of the core loop logic
for chunk in speech_chunks {
    if chunk.samples.len() < 1600 { continue; } // Skip < 0.1s
    
    let (text, confidence) = match provider {
        "parakeet" => engine.transcribe(&chunk), // Fixed confidence
        _ => whisper.transcribe_with_confidence(&chunk), // Scored output
    };
    
    transcripts.push((text, chunk.start_ts, chunk.end_ts));
    emit_progress(&app, stage, percentage, message);
}

```

The loop accumulates `(text, start_ts, end_ts)` tuples while tracking aggregate confidence metrics. Progress events stream to the frontend after each segment.

## Step 10: Meeting Persistence and Metadata Generation

Transcription completion triggers database persistence (lines 340–418). The pipeline:

1. Converts raw tuples to `TranscriptSegment` structs via `create_transcript_segments`
2. Inserts the meeting and transcripts into SQLite via `create_meeting_with_transcripts` from **[`frontend/src-tauri/src/database/manager.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/database/manager.rs)**
3. Writes [`transcripts.json`](https://github.com/Zackriya-Solutions/meetily/blob/main/transcripts.json) and canonical [`metadata.json`](https://github.com/Zackriya-Solutions/meetily/blob/main/metadata.json) to the meeting folder (lines 778–808 in **[`frontend/src-tauri/src/audio/common.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/common.rs)**)

These files enable UI navigation, export functionality, and future reprocessing without retranscribing.

## Event-Driven Frontend Integration

The React frontend communicates with the Rust backend through Tauri commands and event listeners:

```tsx
import { invoke } from '@tauri-apps/api/tauri';
import { listen } from '@tauri-apps/api/event';

// Initiate import with full configuration
async function startImport(filePath: string, title: string) {
  await invoke('start_import_audio_command', {
    source_path: filePath,
    title,
    language: 'en',
    model: undefined,      // Use configured default
    provider: 'whisper',   // or 'parakeet'
  });
}

// Real-time progress updates
listen('import-progress', (e) => {
  const { stage, progress_percentage, message } = e.payload;
  console.log(`[${stage}] ${progress_percentage}% – ${message}`);
});

```

Cancellation is equally straightforward:

```tsx
async function cancelImport() {
  await invoke('cancel_import_command');
}

```

Event types emitted throughout the pipeline:

- `import-progress` — stage transitions and percentage completion
- `import-warning` — non-fatal issues (unsupported metadata, decode edge cases)
- `import-complete` — final success with meeting ID
- `import-error` — terminal failure with structured error message

## Key Files in the Import Architecture

| File | Responsibility |
|------|--------------|
| [`frontend/src-tauri/src/audio/import.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/import.rs) | Core pipeline orchestration |
| [`frontend/src-tauri/src/audio/decoder.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/decoder.rs) | Container decoding via `symphonia` |
| [`frontend/src-tauri/src/audio/vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/vad.rs) | Speech boundary detection |
| [`frontend/src-tauri/src/whisper_engine/whisper_engine.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/whisper_engine/whisper_engine.rs) | Whisper model management and inference |
| [`frontend/src-tauri/src/parakeet_engine/parakeet_engine.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/parakeet_engine/parakeet_engine.rs) | Alternative Parakeet backend |
| [`frontend/src-tauri/src/database/manager.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/database/manager.rs) | SQLite persistence layer |
| [`frontend/src-tauri/src/audio/common.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/common.rs) | Folder creation, transcript formatting, metadata I/O |
| [`frontend/src-tauri/src/api/commands.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/api/commands.rs) | Tauri command surface for frontend binding |

## Performance and Scalability Characteristics

The import pipeline exhibits several optimization patterns worth noting:

- **Memory efficiency** — Streaming decode and VAD process audio in chunks rather than loading entire files
- **CPU parallelism** — Resampling and V leverage available cores where `symphonia` and internal algorithms permit
- **Disk isolation** — Per-meeting folders prevent I/O contention and simplify cleanup
- **Graceful degradation** — Metadata extraction failures fall back to full decode without user intervention

The 20 GB file size limit balances practical use cases (multi-hour recordings) against memory and disk constraints on typical workstations.

## Summary

- Meetily's import feature is implemented entirely in **[`frontend/src-tauri/src/audio/import.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/import.rs)** as a 10-stage asynchronous pipeline
- **Validation** enforces format, size, and integrity constraints before any heavy processing begins
- **Decoding and resampling** standardize arbitrary audio to 16 kHz mono PCM required by transcription engines
- **VAD with 2000 ms redemption time** aggressively removes silence, reducing LLM computation vs. live recording's 400 ms responsiveness
- **Dual-engine support** allows runtime selection between Whisper and Parakeet backends with lazy model loading
- **Progress events** and **cancellation flags** enable responsive UI feedback without blocking the main thread
- **Atomic guard patterns** prevent overlapping imports and ensure clean state recovery
- **Complete local processing** guarantees privacy—no audio data ever leaves the user's machine

## Frequently Asked Questions

### What audio formats can Meetily import?

Meetily accepts any format supported by the `symphonia` crate, which includes WAV, MP3, FLAC, OGG, AAC, and common container formats. The `AUDIO_EXTENSIONS` whitelist in `validate_audio_file` may restrict this further based on confidence in metadata extraction reliability.

### How does Meetily handle very large audio files?

Files up to 20 GB are supported through streaming decode that processes audio in chunks rather than loading everything into memory. The pipeline copies the file to a dedicated meeting folder first, ensuring the original remains untouched while providing stable random access for the decoder.

### Can I switch transcription engines mid-import?

No—engine selection occurs at import initiation via the `provider` parameter. However, the lazy-loading architecture in `get_or_init_whisper` and `get_or_init_parakeet` means switching engines between imports requires no application restart, and models are held resident to accelerate subsequent usage of the same backend.

### What happens if I cancel an import in progress?

The `IMPORT_CANCELLED` flag is set atomically, and the pipeline checks this flag between major stages (validation, decode chunks, transcription segments). Partial progress is discarded; the `ImportGuard` ensures flags clear automatically. The frontend receives a cancellation confirmation event, and the partial meeting folder may be cleaned up or retained for debugging based on build configuration.