How Meetily Imports External Audio Files for Transcription: A Complete Technical Breakdown
Meetily processes external audio files through a 10-step local pipeline that validates, decodes, resamples, runs voice-activity detection, and transcribes using Whisper or Parakeet—entirely offline with cancellable progress tracking.
Meetily's import feature transforms any supported audio file into a fully timestamped meeting transcript without ever transmitting data to remote servers. This article examines the complete technical implementation in the Zackriya-Solutions/meetily repository, walking through each stage from file selection to database persistence.
Overview of the Import Pipeline
The entire workflow is orchestrated from frontend/src-tauri/src/audio/import.rs. The pipeline is asynchronous, cancellable, and progress-aware, emitting real-time events to the React frontend at every stage.
Key architectural decisions that define the system:
- Zero network dependency — all processing happens locally via Rust/Tauri
- Atomic guard pattern — prevents overlapping imports via
ImportGuard - Speech-aware segmentation — VAD with 2000 ms redemption time removes silence efficiently
- Dual-engine support — transparently switches between Whisper and Parakeet transcription backends
Step 1: File Selection and Validation
The import begins when the frontend invokes select_and_validate_audio_command. This triggers validate_audio_file (lines 19–90), which enforces strict entry criteria:
- File must exist and be readable
- Extension must match
AUDIO_EXTENSIONSwhitelist - Size must not exceed 20 GB
- Metadata extraction attempts fast path; falls back to full decode if needed
Upon success, the function returns an AudioFileInfo struct containing title, duration, and format—data that populates the UI's import dialog.
import { invoke } from '@tauri-apps/api/tauri';
async function pickAudioFile() {
const info = await invoke('select_and_validate_audio_command');
if (info) {
console.log('Picked:', info.filename, info.duration_seconds);
}
}
Step 2: Guard Initialization and Cancellation Setup
Before any heavy processing, ImportGuard::acquire atomically sets the global IMPORT_IN_PROGRESS flag. A companion flag, IMPORT_CANCELLED, enables UI-initiated aborts at any stage.
Both flags clear automatically via Drop implementation (lines 31–52), ensuring clean state even on panic. This prevents resource leaks and guarantees that a failed import never blocks future attempts.
Step 3: Meeting Folder Creation and File Copying
The run_import function establishes isolation:
- Creates a dedicated folder via
create_meeting_folder(lines 41–45) - Copies the original audio using
std::fs::copy(lines 46–63)
This preserves the user's original file while giving the pipeline a stable, versioned workspace. The meeting folder becomes the permanent home for transcripts, metadata, and derived assets.
Step 4: Audio Decoding with Progress Streaming
The copied file feeds into decode_audio_file_with_progress, defined in frontend/src-tauri/src/audio/decoder.rs. This leverages the symphonia crate to decode diverse container formats into raw PCM.
The function streams audio while emitting progress events, yielding a DecodedAudio struct containing:
- Interleaved PCM samples
- Original sample rate and channel count
- Precise duration
Step 5: Resampling to Whisper-Compatible Format
Raw decoded audio rarely matches the transcription engine's requirements. The pipeline converts to 16 kHz mono PCM via to_whisper_format_with_progress (lines 104–116 in import.rs).
This standardized format is mandatory for both Whisper and Parakeet backends, eliminating downstream compatibility variance.
Step 6: Voice-Activity Detection (VAD)
The resampled audio flows into get_speech_chunks_with_progress from frontend/src-tauri/src/audio/vad.rs. Key parameters distinguish import processing from live recording:
| Aspect | Import Pipeline | Live Pipeline |
|---|---|---|
| Redemption time | 2000 ms | 400 ms |
| Purpose | Aggressive silence removal | Responsive real-time chunking |
VAD segments the audio into speech chunks with precise start/end timestamps. This reduces LLM load dramatically by discarding silent periods rather than transcribing them.
Step 7: Optional Segment Splitting
Long speech chunks exceeding 25 seconds undergo additional processing. The split_segment_at_silence function (lines 222–242) fractures these at natural silence boundaries.
This prevents word loss at hard cut points and keeps segments within optimal size ranges for the transcription engine's context window.
Step 8: Engine Selection and Lazy Loading
The provider argument ("parakeet" or default Whisper) determines backend selection. The engine initialization logic (lines 260–331) implements efficient resource management:
- Engines load lazily via
get_or_init_whisperorget_or_init_parakeet - If the currently resident model differs from the request,
load_modelswaps it - Model choice respects the database configuration via
get_configured_model
This supports runtime switching without application restart.
Step 9: Transcription Loop with Confidence Scoring
Each speech segment enters the transcription loop (lines 254–311):
// Pseudoglimpse of the core loop logic
for chunk in speech_chunks {
if chunk.samples.len() < 1600 { continue; } // Skip < 0.1s
let (text, confidence) = match provider {
"parakeet" => engine.transcribe(&chunk), // Fixed confidence
_ => whisper.transcribe_with_confidence(&chunk), // Scored output
};
transcripts.push((text, chunk.start_ts, chunk.end_ts));
emit_progress(&app, stage, percentage, message);
}
The loop accumulates (text, start_ts, end_ts) tuples while tracking aggregate confidence metrics. Progress events stream to the frontend after each segment.
Step 10: Meeting Persistence and Metadata Generation
Transcription completion triggers database persistence (lines 340–418). The pipeline:
- Converts raw tuples to
TranscriptSegmentstructs viacreate_transcript_segments - Inserts the meeting and transcripts into SQLite via
create_meeting_with_transcriptsfromfrontend/src-tauri/src/database/manager.rs - Writes
transcripts.jsonand canonicalmetadata.jsonto the meeting folder (lines 778–808 infrontend/src-tauri/src/audio/common.rs)
These files enable UI navigation, export functionality, and future reprocessing without retranscribing.
Event-Driven Frontend Integration
The React frontend communicates with the Rust backend through Tauri commands and event listeners:
import { invoke } from '@tauri-apps/api/tauri';
import { listen } from '@tauri-apps/api/event';
// Initiate import with full configuration
async function startImport(filePath: string, title: string) {
await invoke('start_import_audio_command', {
source_path: filePath,
title,
language: 'en',
model: undefined, // Use configured default
provider: 'whisper', // or 'parakeet'
});
}
// Real-time progress updates
listen('import-progress', (e) => {
const { stage, progress_percentage, message } = e.payload;
console.log(`[${stage}] ${progress_percentage}% – ${message}`);
});
Cancellation is equally straightforward:
async function cancelImport() {
await invoke('cancel_import_command');
}
Event types emitted throughout the pipeline:
import-progress— stage transitions and percentage completionimport-warning— non-fatal issues (unsupported metadata, decode edge cases)import-complete— final success with meeting IDimport-error— terminal failure with structured error message
Key Files in the Import Architecture
| File | Responsibility |
|---|---|
frontend/src-tauri/src/audio/import.rs |
Core pipeline orchestration |
frontend/src-tauri/src/audio/decoder.rs |
Container decoding via symphonia |
frontend/src-tauri/src/audio/vad.rs |
Speech boundary detection |
frontend/src-tauri/src/whisper_engine/whisper_engine.rs |
Whisper model management and inference |
frontend/src-tauri/src/parakeet_engine/parakeet_engine.rs |
Alternative Parakeet backend |
frontend/src-tauri/src/database/manager.rs |
SQLite persistence layer |
frontend/src-tauri/src/audio/common.rs |
Folder creation, transcript formatting, metadata I/O |
frontend/src-tauri/src/api/commands.rs |
Tauri command surface for frontend binding |
Performance and Scalability Characteristics
The import pipeline exhibits several optimization patterns worth noting:
- Memory efficiency — Streaming decode and VAD process audio in chunks rather than loading entire files
- CPU parallelism — Resampling and V leverage available cores where
symphoniaand internal algorithms permit - Disk isolation — Per-meeting folders prevent I/O contention and simplify cleanup
- Graceful degradation — Metadata extraction failures fall back to full decode without user intervention
The 20 GB file size limit balances practical use cases (multi-hour recordings) against memory and disk constraints on typical workstations.
Summary
- Meetily's import feature is implemented entirely in
frontend/src-tauri/src/audio/import.rsas a 10-stage asynchronous pipeline - Validation enforces format, size, and integrity constraints before any heavy processing begins
- Decoding and resampling standardize arbitrary audio to 16 kHz mono PCM required by transcription engines
- VAD with 2000 ms redemption time aggressively removes silence, reducing LLM computation vs. live recording's 400 ms responsiveness
- Dual-engine support allows runtime selection between Whisper and Parakeet backends with lazy model loading
- Progress events and cancellation flags enable responsive UI feedback without blocking the main thread
- Atomic guard patterns prevent overlapping imports and ensure clean state recovery
- Complete local processing guarantees privacy—no audio data ever leaves the user's machine
Frequently Asked Questions
What audio formats can Meetily import?
Meetily accepts any format supported by the symphonia crate, which includes WAV, MP3, FLAC, OGG, AAC, and common container formats. The AUDIO_EXTENSIONS whitelist in validate_audio_file may restrict this further based on confidence in metadata extraction reliability.
How does Meetily handle very large audio files?
Files up to 20 GB are supported through streaming decode that processes audio in chunks rather than loading everything into memory. The pipeline copies the file to a dedicated meeting folder first, ensuring the original remains untouched while providing stable random access for the decoder.
Can I switch transcription engines mid-import?
No—engine selection occurs at import initiation via the provider parameter. However, the lazy-loading architecture in get_or_init_whisper and get_or_init_parakeet means switching engines between imports requires no application restart, and models are held resident to accelerate subsequent usage of the same backend.
What happens if I cancel an import in progress?
The IMPORT_CANCELLED flag is set atomically, and the pipeline checks this flag between major stages (validation, decode chunks, transcription segments). Partial progress is discarded; the ImportGuard ensures flags clear automatically. The frontend receives a cancellation confirmation event, and the partial meeting folder may be cleaned up or retained for debugging based on build configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →