Meetily GPU Acceleration Backends: Metal, CUDA, Vulkan, and HIP BLAS Support

Meetily supports five GPU acceleration backends—Metal, CUDA, Vulkan, HIP BLAS, and CPU—selected at compile time via Cargo feature flags and validated at runtime through system environment checks.

Meetily, the open-source meeting transcription tool from Zackriya-Solutions, leverages hardware-accelerated Whisper transcription through multiple GPU backends. Understanding which GPU acceleration backends Meetily supports and how they are detected helps developers optimize build configurations and runtime performance across diverse hardware environments.

Supported GPU Acceleration Backends

Meetily targets five distinct compute backends, each controlled by Cargo feature flags and detected through specific system probes:

  • Metal – Targets Apple Silicon (aarch64) devices using the metal feature or default macOS builds
  • CUDA – Accelerates transcription on NVIDIA GPUs when the cuda feature is enabled
  • Vulkan – Provides GPU acceleration for AMD, Intel, and other Vulkan-compatible hardware via the vulkan feature
  • HIP BLAS – Supports AMD GPUs through the ROCm stack using the hipblas feature
  • CPU – Fallback compute mode when no GPU features are compiled or hardware is unavailable

Compile-Time Backend Selection

The WhisperCompiledBackend::current() function in frontend/src-tauri/src/whisper_engine/acceleration.rs determines which backend the binary was built with by evaluating Cargo feature flags at compile time:

if cfg!(feature = "cuda") { Self::Cuda }
else if cfg!(feature = "vulkan") { Self::Vulkan }
else if cfg!(feature = "hipblas") { Self::HipBlas }
else if cfg!(target_os = "macos") || cfg!(feature = "metal") { Self::Metal }
else { Self::Cpu }

This static selection establishes the maximum capability set available to the runtime detector. If the binary is compiled without GPU features, the system defaults to CPU-only inference regardless of available hardware.

Runtime GPU Detection

When the application initializes, HardwareProfile::detect() in frontend/src-tauri/src/audio/hardware_detector.rs executes detect_gpu() to probe the host system and validate GPU availability:

  1. Metal detection – Confirms Apple Silicon presence through architecture validation (ARCH == "aarch64")
  2. CUDA detection – Inspects environment variables (CUDA_PATH or CUDA_HOME) or checks for the existence of /usr/local/cuda
  3. Vulkan detection – Validates the VULKAN_SDK environment variable on Unix systems or inspects for vulkan-1.dll on Windows platforms

If detection fails for all backends, the system reports has_gpu_acceleration = false with gpu_type = GpuType::None, forcing CPU fallback even if the binary was compiled with GPU support.

Acceleration Decision Logic

The whisper_context_acceleration_for() function combines the compiled backend, runtime-detected GPU type, and hardware performance tier to finalize the acceleration configuration:

let use_gpu = !matches!(compiled_backend, WhisperCompiledBackend::Cpu);
let fast_tier = matches!(performance_tier, PerformanceTier::High | PerformanceTier::Ultra);
let flash_attn = matches!(compiled_backend, WhisperCompiledBackend::Metal | WhisperCompiledBackend::Cuda) && fast_tier;

GPU usage remains disabled only when the compiled backend is Cpu. Flash Attention—a memory-efficient attention mechanism—activates exclusively for Metal and CUDA backends when the performance tier is High or Ultra. Vulkan and HIP BLAS never enable flash attention, even when GPU hardware is present and detected.

Implementation Examples

Check the detected hardware profile before initializing the Whisper engine:

use crate::audio::HardwareProfile;

fn log_hardware() {
    let profile = HardwareProfile::detect();
    log::info!(
        "CPU cores: {}, GPU: {:?} (acceleration: {}), Memory: {} GB, Tier: {:?}",
        profile.cpu_cores,
        profile.gpu_type,
        profile.has_gpu_acceleration,
        profile.memory_gb,
        profile.performance_tier,
    );
}

Configure the appropriate Whisper context based on compiled backend and detected hardware:

use crate::whisper_engine::{whisper_context_acceleration_for, WhisperCompiledBackend};
use crate::audio::{HardwareProfile, PerformanceTier};

fn create_whisper_context() {
    let hw = HardwareProfile::detect();
    let compiled = WhisperCompiledBackend::current();

    let accel = whisper_context_acceleration_for(
        compiled,
        hw.gpu_type,
        hw.performance_tier,
    );

    log::info!("Whisper acceleration: {}", accel.status_label());
    // Pass `accel` to the Whisper engine initializer…
}

Summary

  • Meetily supports Metal, CUDA, Vulkan, HIP BLAS, and CPU backends for Whisper transcription acceleration
  • Compile-time selection occurs through Cargo feature flags in WhisperCompiledBackend::current() located in frontend/src-tauri/src/whisper_engine/acceleration.rs
  • Runtime detection probes system environments via HardwareProfile::detect() in frontend/src-tauri/src/audio/hardware_detector.rs, checking for Metal (aarch64), CUDA (environment paths), or Vulkan (SDK variables)
  • Flash Attention is restricted to Metal and CUDA backends operating at High or Ultra performance tiers
  • GPU acceleration requires both correct feature flags at build time and successful hardware detection at runtime

Frequently Asked Questions

How do I enable CUDA support when building Meetily?

Enable the cuda feature flag in your Cargo build configuration. The WhisperCompiledBackend::current() function checks for cfg!(feature = "cuda") first in its conditional chain, compiling the binary with NVIDIA CUDA support. Ensure the target system has the CUDA toolkit installed, as runtime detection looks for CUDA_PATH, CUDA_HOME, or /usr/local/cuda to confirm driver availability.

Why doesn't Meetily detect my AMD GPU for acceleration?

Meetily detects AMD GPUs through HIP BLAS (for ROCm) or Vulkan (for general compute), not through CUDA. Verify you compiled with the hipblas or vulkan feature enabled. Runtime detection for Vulkan checks the VULKAN_SDK environment variable or vulkan-1.dll on Windows, while HIP BLAS detection relies on the compiled backend matching the WhisperCompiledBackend::HipBlas variant.

What is Flash Attention and when is it enabled in Meetily?

Flash Attention is a memory-efficient attention algorithm that reduces memory bandwidth bottlenecks during transcription. In Meetily, flash attention activates only when the compiled backend is Metal or CUDA and the detected performance tier is High or Ultra. Even if Vulkan or HIP BLAS hardware is detected, flash attention remains disabled for these backends according to the logic in whisper_context_acceleration_for().

Can Meetily use GPU acceleration on macOS Intel machines?

No. Meetily's Metal support specifically targets Apple Silicon (aarch64) architecture through the metal feature. Intel-based Macs fall back to CPU-only inference unless experimental Vulkan drivers are installed and the binary is compiled with the vulkan feature, though this configuration is not the primary development target for the Meetily project.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →