How Meetily Implements GPU Acceleration for Whisper and Parakeet Models in Rust
Meetily implements GPU acceleration for Whisper and Parakeet models using Rust's conditional compilation with Cargo feature flags, supporting CUDA (NVIDIA), Vulkan (AMD/Intel), and Metal (macOS) through a unified acceleration infrastructure in the whisper_engine crate.
Meetily is an open-source AI meeting assistant built with Tauri and Rust. Its speech-to-text and audio generation capabilities rely on GPU-accelerated inference for real-time performance. This article examines how the Meetily codebase enables GPU support for both Whisper transcription and Parakeet audio generation models.
Architecture Overview
Meetily's GPU acceleration follows a feature-gated, runtime-detected architecture. The system decouples compilation targets from runtime availability, allowing a single build to support multiple GPU backends—or fall back to CPU inference when hardware acceleration is unavailable.
The design centers on three principles:
- Compile-time selection via Cargo features (
cuda,vulkan,hipblas,openblas) - Runtime detection of installed GPU libraries
- Unified backend interface shared between Whisper and Parakeet engines
Conditional Compilation with Cargo Features
Meetily uses Rust's cfg attributes to include GPU-specific code only when explicitly requested. This prevents compilation errors on systems without CUDA or Vulkan toolchains.
Feature-Gate Pattern
The pattern appears throughout the codebase, as seen in llama-helper/src/main.rs and whisper_engine/whisper_engine.rs:
#[cfg(feature = "cuda")]
if let Some(vram) = detect_cuda_vram() {
// CUDA-specific initialization
}
#[cfg(feature = "vulkan")]
else if let Some(vram) = detect_vulkan_vram() {
// Vulkan-specific initialization
}
Available feature flags in Meetily:
| Flag | GPU Backend | Target Hardware |
|---|---|---|
cuda |
CUDA | NVIDIA GPUs |
vulkan |
Vulkan | AMD, Intel GPUs |
hipblas |
HIP/ROCm | AMD GPUs (Linux) |
openblas |
OpenBLAS | CPU (optimized) |
| (none) | Pure Rust | CPU (fallback) |
The WhisperContextAcceleration Struct
The core abstraction for GPU acceleration lives in frontend/src-tauri/src/whisper_engine/acceleration.rs. The WhisperContextAcceleration struct encapsulates backend selection:
pub struct WhisperContextAcceleration {
pub use_cuda: bool,
pub use_vulkan: bool,
pub use_metal: bool,
pub device_id: Option<i32>,
pub memory_mb: Option<usize>,
}
This struct is constructed after runtime checks in hardware_detector.rs and passed to both Whisper and Parakeet engines. Both engines use identical logic to determine whether GPU inference is available and which backend to prioritize.
Runtime GPU Detection
Meetily performs hardware detection at two stages: build time and runtime.
Build-Time Detection
The frontend/src-tauri/build.rs script emits helpful warnings to guide developers:
# Example output during compilation
💡 For NVIDIA GPU: cargo build --release --features cuda
⚠️ No GPU feature enabled — falling back to CPU inference
Runtime Detection
The audio/hardware_detector.rs module checks for:
- CUDA: Presence of
/usr/local/cudadirectory andlibcuda.so - Vulkan: Availability of
libvulkan.soorvulkan-1.dll - Metal: macOS platform with compatible hardware
Results populate the WhisperContextAcceleration struct and appear in Tauri application logs.
Whisper Engine GPU Implementation
The Whisper speech-to-text engine in whisper_engine/whisper_engine.rs loads models with automatic backend selection:
// Load model with automatic GPU detection
let engine = WhisperEngine::new(app_handle.clone()).await?;
engine.load_model("large-v3").await?;
// Or force specific backend
let accel = WhisperContextAcceleration {
use_cuda: false,
use_vulkan: true,
use_metal: false,
..Default::default()
};
engine.set_acceleration(accel);
engine.load_model("medium").await?;
The engine delegates to whisper_rs (Rust bindings for whisper.cpp) with the appropriate GPU context—gpu::cuda::Context, gpu::vulkan::Context, or gpu::metal::Context—based on enabled features and runtime detection.
Parakeet Engine GPU Implementation
Parakeet—Meetily's audio generation model—shares the same acceleration infrastructure. Located in parakeet_engine/parakeet_engine.rs, it reuses WhisperContextAcceleration:
let parakeet = ParakeetEngine::new(app_handle.clone()).await?;
parakeet.load_model("parakeet-base").await?;
Because both engines reside in the same Tauri workspace and share the whisper_engine crate's acceleration module, any GPU support added for Whisper automatically extends to Parakeet. The underlying inference code (also whisper_rs-compatible) applies the same backend selection logic for audio generation workloads.
Building with GPU Support
NVIDIA (CUDA)
cargo build --release --features cuda
AMD/Intel (Vulkan)
cargo build --release --features vulkan
macOS (Metal)
cargo build --release --features metal
Multiple Backends
# Compile all GPU backends; runtime selects best available
cargo build --release --features "cuda,vulkan,metal"
Key Files in the GPU Acceleration Stack
| File | Purpose |
|---|---|
whisper_engine/acceleration.rs |
WhisperContextAcceleration struct and backend configuration |
whisper_engine/whisper_engine.rs |
Whisper model loading and inference with GPU backend selection |
parakeet_engine/parakeet_engine.rs |
Parakeet audio generation reusing acceleration infrastructure |
audio/hardware_detector.rs |
Runtime detection of CUDA, Vulkan, and Metal availability |
frontend/src-tauri/build.rs |
Build-time GPU hints and feature validation |
llama-helper/src/main.rs |
Example of #[cfg(feature = "cuda")] conditional compilation |
Summary
- Meetily uses Cargo feature flags (
cuda,vulkan,metal) to conditionally compile GPU-specific code - The
WhisperContextAccelerationstruct inwhisper_engine/acceleration.rsunifies backend selection for both engines - Runtime detection in
hardware_detector.rschecks for installed GPU libraries before attempting acceleration - Whisper and Parakeet share identical infrastructure—GPU improvements apply to both transcription and audio generation
- Automatic CPU fallback ensures the application runs on any hardware configuration
Frequently Asked Questions
What GPU vendors does Meetily support?
Meetily supports NVIDIA GPUs via CUDA, AMD and Intel GPUs via Vulkan, and Apple Silicon via Metal. The HIP/ROCm backend (hipblas feature) provides experimental AMD support on Linux. According to the source code in acceleration.rs, the engine prioritizes CUDA when multiple backends are available.
How does Meetily handle systems with no GPU?
When no GPU features are compiled or runtime detection fails, Meetily falls back to CPU-only inference using OpenBLAS (if openblas feature enabled) or pure Rust implementations. The hardware_detector.rs module sets all acceleration flags to false, and both WhisperEngine and ParakeetEngine proceed with CPU contexts.
Can I force a specific GPU backend at runtime?
Yes. Instantiate WhisperContextAcceleration directly with your preferred backend flags and pass it to engine.set_acceleration() before loading models. This overrides automatic detection, as shown in the code example forcing Vulkan over CUDA.
Why do Whisper and Parakeet share GPU code?
Both engines reside in the same Tauri workspace and depend on the whisper_engine crate's acceleration module. Since both use whisper_rs-compatible inference, unifying the GPU abstraction reduces maintenance and ensures consistent behavior across Meetily's AI features.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →