How Meetily Implements GPU Acceleration for Whisper and Parakeet Models in Rust

Meetily implements GPU acceleration for Whisper and Parakeet models using Rust's conditional compilation with Cargo feature flags, supporting CUDA (NVIDIA), Vulkan (AMD/Intel), and Metal (macOS) through a unified acceleration infrastructure in the whisper_engine crate.

Meetily is an open-source AI meeting assistant built with Tauri and Rust. Its speech-to-text and audio generation capabilities rely on GPU-accelerated inference for real-time performance. This article examines how the Meetily codebase enables GPU support for both Whisper transcription and Parakeet audio generation models.

Architecture Overview

Meetily's GPU acceleration follows a feature-gated, runtime-detected architecture. The system decouples compilation targets from runtime availability, allowing a single build to support multiple GPU backends—or fall back to CPU inference when hardware acceleration is unavailable.

The design centers on three principles:

  • Compile-time selection via Cargo features (cuda, vulkan, hipblas, openblas)
  • Runtime detection of installed GPU libraries
  • Unified backend interface shared between Whisper and Parakeet engines

Conditional Compilation with Cargo Features

Meetily uses Rust's cfg attributes to include GPU-specific code only when explicitly requested. This prevents compilation errors on systems without CUDA or Vulkan toolchains.

Feature-Gate Pattern

The pattern appears throughout the codebase, as seen in llama-helper/src/main.rs and whisper_engine/whisper_engine.rs:

#[cfg(feature = "cuda")]
if let Some(vram) = detect_cuda_vram() {
    // CUDA-specific initialization
}
#[cfg(feature = "vulkan")]
else if let Some(vram) = detect_vulkan_vram() {
    // Vulkan-specific initialization
}

Available feature flags in Meetily:

Flag GPU Backend Target Hardware
cuda CUDA NVIDIA GPUs
vulkan Vulkan AMD, Intel GPUs
hipblas HIP/ROCm AMD GPUs (Linux)
openblas OpenBLAS CPU (optimized)
(none) Pure Rust CPU (fallback)

The WhisperContextAcceleration Struct

The core abstraction for GPU acceleration lives in frontend/src-tauri/src/whisper_engine/acceleration.rs. The WhisperContextAcceleration struct encapsulates backend selection:

pub struct WhisperContextAcceleration {
    pub use_cuda: bool,
    pub use_vulkan: bool,
    pub use_metal: bool,
    pub device_id: Option<i32>,
    pub memory_mb: Option<usize>,
}

This struct is constructed after runtime checks in hardware_detector.rs and passed to both Whisper and Parakeet engines. Both engines use identical logic to determine whether GPU inference is available and which backend to prioritize.

Runtime GPU Detection

Meetily performs hardware detection at two stages: build time and runtime.

Build-Time Detection

The frontend/src-tauri/build.rs script emits helpful warnings to guide developers:


# Example output during compilation

💡 For NVIDIA GPU: cargo build --release --features cuda
⚠️  No GPU feature enabled — falling back to CPU inference

Runtime Detection

The audio/hardware_detector.rs module checks for:

  • CUDA: Presence of /usr/local/cuda directory and libcuda.so
  • Vulkan: Availability of libvulkan.so or vulkan-1.dll
  • Metal: macOS platform with compatible hardware

Results populate the WhisperContextAcceleration struct and appear in Tauri application logs.

Whisper Engine GPU Implementation

The Whisper speech-to-text engine in whisper_engine/whisper_engine.rs loads models with automatic backend selection:

// Load model with automatic GPU detection
let engine = WhisperEngine::new(app_handle.clone()).await?;
engine.load_model("large-v3").await?;

// Or force specific backend
let accel = WhisperContextAcceleration {
    use_cuda: false,
    use_vulkan: true,
    use_metal: false,
    ..Default::default()
};
engine.set_acceleration(accel);
engine.load_model("medium").await?;

The engine delegates to whisper_rs (Rust bindings for whisper.cpp) with the appropriate GPU context—gpu::cuda::Context, gpu::vulkan::Context, or gpu::metal::Context—based on enabled features and runtime detection.

Parakeet Engine GPU Implementation

Parakeet—Meetily's audio generation model—shares the same acceleration infrastructure. Located in parakeet_engine/parakeet_engine.rs, it reuses WhisperContextAcceleration:

let parakeet = ParakeetEngine::new(app_handle.clone()).await?;
parakeet.load_model("parakeet-base").await?;

Because both engines reside in the same Tauri workspace and share the whisper_engine crate's acceleration module, any GPU support added for Whisper automatically extends to Parakeet. The underlying inference code (also whisper_rs-compatible) applies the same backend selection logic for audio generation workloads.

Building with GPU Support

NVIDIA (CUDA)

cargo build --release --features cuda

AMD/Intel (Vulkan)

cargo build --release --features vulkan

macOS (Metal)

cargo build --release --features metal

Multiple Backends


# Compile all GPU backends; runtime selects best available

cargo build --release --features "cuda,vulkan,metal"

Key Files in the GPU Acceleration Stack

File Purpose
whisper_engine/acceleration.rs WhisperContextAcceleration struct and backend configuration
whisper_engine/whisper_engine.rs Whisper model loading and inference with GPU backend selection
parakeet_engine/parakeet_engine.rs Parakeet audio generation reusing acceleration infrastructure
audio/hardware_detector.rs Runtime detection of CUDA, Vulkan, and Metal availability
frontend/src-tauri/build.rs Build-time GPU hints and feature validation
llama-helper/src/main.rs Example of #[cfg(feature = "cuda")] conditional compilation

Summary

  • Meetily uses Cargo feature flags (cuda, vulkan, metal) to conditionally compile GPU-specific code
  • The WhisperContextAcceleration struct in whisper_engine/acceleration.rs unifies backend selection for both engines
  • Runtime detection in hardware_detector.rs checks for installed GPU libraries before attempting acceleration
  • Whisper and Parakeet share identical infrastructure—GPU improvements apply to both transcription and audio generation
  • Automatic CPU fallback ensures the application runs on any hardware configuration

Frequently Asked Questions

What GPU vendors does Meetily support?

Meetily supports NVIDIA GPUs via CUDA, AMD and Intel GPUs via Vulkan, and Apple Silicon via Metal. The HIP/ROCm backend (hipblas feature) provides experimental AMD support on Linux. According to the source code in acceleration.rs, the engine prioritizes CUDA when multiple backends are available.

How does Meetily handle systems with no GPU?

When no GPU features are compiled or runtime detection fails, Meetily falls back to CPU-only inference using OpenBLAS (if openblas feature enabled) or pure Rust implementations. The hardware_detector.rs module sets all acceleration flags to false, and both WhisperEngine and ParakeetEngine proceed with CPU contexts.

Can I force a specific GPU backend at runtime?

Yes. Instantiate WhisperContextAcceleration directly with your preferred backend flags and pass it to engine.set_acceleration() before loading models. This overrides automatic detection, as shown in the code example forcing Vulkan over CUDA.

Why do Whisper and Parakeet share GPU code?

Both engines reside in the same Tauri workspace and depend on the whisper_engine crate's acceleration module. Since both use whisper_rs-compatible inference, unifying the GPU abstraction reduces maintenance and ensures consistent behavior across Meetily's AI features.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →