How Meetily Switches Between Whisper and Parakeet Transcription Engines

Meetily handles the switch between Whisper and Parakeet transcription engines through a unified frontend model selector that encodes the provider as provider:name, then delegates to separate Tauri commands (whisper_transcribe or parakeet_transcribe) implemented in independent Rust modules.

Meetily is an open-source meeting assistant built with Tauri that supports multiple speech-to-text backends. The application allows users to dynamically switch between Whisper and Parakeet transcription engines through a unified interface that abstracts the underlying provider differences while maintaining separate optimized implementations in the Rust core.

Unified Model Discovery and Selection

The frontend treats both transcription engines as a single pool of options while preserving their distinct identities through prefixed keys.

Fetching Available Models from Both Engines

The useTranscriptionModels hook queries the Rust backend independently for each engine's available models. It invokes two distinct Tauri commands to retrieve the model metadata:

// From frontend/src/hooks/useTranscriptionModels.ts
const whisperModels = await invoke<RawModelInfo[]>('whisper_get_available_models');
const parakeetModels = await invoke<RawModelInfo[]>('parakeet_get_available_models');

These calls hit separate Rust functions that scan their respective model directories—Whisper models in src-tauri/src/whisper_engine/ and Parakeet models in src-tauri/src/parakeet_engine/.

The Provider-Prefixed Selection Key

After fetching both model sets, the hook constructs a unified ModelOption[] array where each entry includes a provider field ('whisper' or 'parakeet'). The UI presents these as distinct categories:

  • Whisper options: Displayed as "🏠 Whisper: {model_name}"
  • Parakeet options: Displayed as "⚡ Parakeet: {model_name}"

When a user selects a model, the hook stores the value as a composite string in the format provider:name. For example, selecting the Whisper base model stores whisper:base, while selecting the Parakeet neon model stores parakeet:neon. This encoding happens in frontend/src/hooks/useTranscriptionModels.ts (lines 33-40) and allows the application to retain provider context without complex state objects.

Runtime Engine Dispatch

When recording begins, the frontend parses the provider prefix to determine which Rust command to invoke.

Parsing the Selected Provider

The useRecordingStart hook handles the transcription initialization. It extracts the provider and model name by splitting the stored key on the colon delimiter:

const startTranscription = async () => {
  const [provider, model] = selectedModelKey.split(':');
  const colonIndex = selectedModelKey.indexOf(':');
  const providerName = selectedModelKey.slice(0, colonIndex);
  const modelName = selectedModelKey.slice(colonIndex + 1);
  
  // Provider is now 'whisper' or 'parakeet'
  // Model is the specific model identifier
};

This parsing ensures the frontend knows exactly which transcription backend to activate before making any Rust calls.

Command Invocation Strategy

Based on the parsed provider, the hook invokes the corresponding Tauri command with the model name as a parameter:

  • Whisper: invoke('whisper_transcribe', { model: modelName })
  • Parakeet: invoke('parakeet_transcribe', { model: modelName })

These commands are registered in frontend/src-tauri/src/lib.rs (around line 572), where each function is annotated with #[tauri::command] and mapped to its respective engine implementation. The explicit command separation ensures type safety and allows each engine to handle its own model loading, VAD (Voice Activity Detection) processing, and inference pipelines without interference.

Rust Backend Architecture

The Rust core maintains strict separation between the two transcription engines to prevent resource conflicts and allow engine-specific optimizations.

Independent Engine Modules

The src-tauri/src/ directory contains two distinct subsystems:

  • whisper_engine/: Handles all Whisper-specific logic including GGML/GGUF model loading, tokenization, and inference
  • parakeet_engine/: Manages NVIDIA NeMo-based Parakeet models with CUDA acceleration and specialized audio preprocessing

Each module exports its own command handlers (whisper_transcribe and parakeet_transcribe) which are registered separately in the Tauri application builder. Because the frontend explicitly specifies the provider in the command call, the Rust side does not need to perform runtime detection—it simply routes the request to the pre-registered handler for that engine.

Implementation Example

The following TypeScript React components demonstrate the complete flow from model selection to transcription start:

// Model selection dropdown (simplified from ImportAudioDialog.tsx)
<Select
  value={selectedModelKey}
  onValueChange={setSelectedModelKey}
>
  {/* Whisper options */}
  {whisperModels.map(m => (
    <SelectItem key={m.name} value={`whisper:${m.name}`}>
      🏠 Whisper: {m.name}
    </SelectItem>
  ))}
  
  {/* Parakeet options */}
  {parakeetModels.map(m => (
    <SelectItem key={m.name} value={`parakeet:${m.name}`}>
      ⚡ Parakeet: {m.name}
    </SelectItem>
  ))}
</Select>
// Transcription initiation (from useRecordingStart.ts)
const handleRecordingStart = async () => {
  const colonIndex = selectedModelKey.indexOf(':');
  const provider = selectedModelKey.slice(0, colonIndex);
  const model = selectedModelKey.slice(colonIndex + 1);
  
  if (provider === 'whisper') {
    await invoke('whisper_transcribe', { 
      model, 
      audioPath: recordingPath 
    });
  } else if (provider === 'parakeet') {
    await invoke('parakeet_transcribe', { 
      model, 
      audioPath: recordingPath 
    });
  }
};

This implementation ensures that switching between Whisper and Parakeet transcription engines requires no backend reconfiguration—only a change to the frontend state that propagates through the provider-prefixed key system.

Summary

  • Unified selection: Meetily uses a composite key format (provider:name) to track which transcription engine the user selects in the dropdown
  • Dual command dispatch: The frontend calls whisper_transcribe or parakeet_transcribe based on parsing the provider prefix from the selected model key
  • Isolated Rust modules: Each engine lives in its own directory (whisper_engine/ and parakeet_engine/) with separate command registrations in src-tauri/src/lib.rs
  • Zero backend switching: The Rust backend does not dynamically choose engines; the frontend explicitly invokes the correct command, making the architecture deterministic and type-safe

Frequently Asked Questions

How does Meetily encode the transcription provider selection?

Meetily encodes the provider and model as a single string in the format provider:name (e.g., whisper:base or parakeet:neon). The useTranscriptionModels hook constructs this key when building the dropdown options in frontend/src/hooks/useTranscriptionModels.ts, and the useRecordingStart hook splits this string to determine which Tauri command to invoke.

What happens if an invalid provider string is passed to the transcription command?

The frontend code in useRecordingStart.ts explicitly checks for 'whisper' or 'parakeet' after splitting the key. If the provider does not match either string, the application does not invoke any transcription command, preventing errors in the Rust backend. The type safety of Tauri's command system ensures only registered commands (whisper_transcribe or parakeet_transcribe) can be called.

Are Whisper and Parakeet models interchangeable without changing the frontend code?

No, the models are not interchangeable because each engine requires specific initialization parameters and audio preprocessing. However, the frontend code in ImportAudioDialog.tsx and RetranscribeDialog handles both engines uniformly through the provider-prefixed key system, so users can switch between them seamlessly without restarting the application or modifying the underlying Rust code.

Where are the transcription engines implemented in the Meetily codebase?

The Whisper implementation resides in frontend/src-tauri/src/whisper_engine/ and the Parakeet implementation is located in frontend/src-tauri/src/parakeet_engine/. Both engines are registered as Tauri commands in frontend/src-tauri/src/lib.rs around line 572, where the parakeet_transcribe function is exported alongside its Whisper counterpart.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →