How Meetily's Acceleration Module Detects GPU Capabilities with WhisperCompiledBackend

Meetily detects GPU capabilities through a dual-layer approach: compile-time feature flags in the WhisperCompiledBackend enum determine which GPU backends (Metal, CUDA, Vulkan) are built into the binary, while runtime hardware detection via HardwareProfile::detect() verifies actual GPU presence before enabling acceleration.

Meetily is an open-source meeting assistant that leverages the Whisper‑rs library for local speech transcription. To maximize performance across diverse hardware—from Apple Silicon Macs to NVIDIA workstations—the Acceleration module implements sophisticated GPU detection logic. This code resides in the Rust-based Tauri backend, specifically within the whisper_engine and audio subsystems.

Compile-Time Backend Detection with WhisperCompiledBackend

The foundation of GPU support lies in the WhisperCompiledBackend enum defined in frontend/src-tauri/src/whisper_engine/acceleration.rs. This enum encodes which GPU API the binary was compiled to support using Rust's conditional compilation features.

The Rust Feature-Gated Enum

pub enum WhisperCompiledBackend {
    Metal,   // macOS – enabled with `--features metal`
    Cuda,    // NVIDIA – enabled with `--features cuda`
    Vulkan,  // AMD/Intel – enabled with `--features vulkan`
    HipBlas, // AMD – enabled with `--features hipblas`
    Cpu,     // No GPU feature compiled (fallback)
}

The current() method inspects Cargo feature flags at runtime to return the active variant. According to the source code in acceleration.rs, this method uses the cfg! macro to check which features were enabled during compilation:

impl WhisperCompiledBackend {
    pub fn current() -> Self {
        if cfg!(feature = "cuda") { 
            Self::Cuda 
        } else if cfg!(feature = "vulkan") { 
            Self::Vulkan 
        } else if cfg!(feature = "hipblas") { 
            Self::HipBlas 
        } else if cfg!(target_os = "macos") || cfg!(feature = "metal") { 
            Self::Metal 
        } else { 
            Self::Cpu 
        }
    }
}

This compile-time detection ensures that the binary only attempts to use GPU APIs it was actually built to interface with, preventing runtime linking errors.

Runtime GPU Detection via HardwareProfile

Even when compiled with GPU support, the host machine might lack the appropriate hardware. Meetily addresses this through runtime probing in the audio subsystem.

Probing System Hardware

The HardwareProfile::detect() method (located in frontend/src-tauri/src/audio/hardware_detector.rs) performs system introspection to identify available compute resources. This detection ultimately calls audio::hardware_detector::detect_gpu(), which returns a GpuType enum indicating whether the system has CUDA, Vulkan, Metal, or no GPU available.

let hardware_profile = crate::audio::HardwareProfile::detect();   // runtime probe
let runtime_gpu = hardware_profile.gpu_type;                     // GpuType enum

The detection logic also calculates a PerformanceTier (Low, Medium, High, or Ultra) based on CPU core count and available RAM, which influences whether advanced optimizations like Flash-Attention should be enabled.

Combining Compile-Time and Runtime Data

The core decision-making occurs in the whisper_context_acceleration_for function, which reconciles what the binary can do with what the hardware actually provides.

The Decision Engine Logic

This function receives three critical inputs:

  • compiled_backend: The WhisperCompiledBackend variant from current()
  • runtime_detected_gpu: The GpuType from HardwareProfile
  • performance_tier: The system's PerformanceTier classification

According to the implementation in acceleration.rs, the function constructs a WhisperContextAcceleration struct that determines both GPU usage and Flash-Attention eligibility:

pub fn whisper_context_acceleration_for(
    compiled_backend: WhisperCompiledBackend,
    runtime_detected_gpu: GpuType,
    performance_tier: PerformanceTier,
) -> WhisperContextAcceleration {
    // Use GPU if a GPU backend was compiled in (i.e. not Cpu)
    let use_gpu = !matches!(compiled_backend, WhisperCompiledBackend::Cpu);
    
    // Flash‑Attention only on Metal or CUDA when tier is High/Ultra
    let fast_tier = matches!(performance_tier, PerformanceTier::High | PerformanceTier::Ultra);
    let flash_attn = match compiled_backend {
        WhisperCompiledBackend::Metal | WhisperCompiledBackend::Cuda => fast_tier,
        _ => false,
    };

    WhisperContextAcceleration {
        compiled_backend,
        runtime_detected_gpu,
        use_gpu,
        flash_attn: use_gpu && flash_attn,
        gpu_device: 0,
    }
}

Key behavioral rules:

  • Flash-Attention is only enabled for Metal or CUDA backends on High or Ultra tier machines
  • Vulkan and HIP-BLAS never enable Flash-Attention (as implemented in the current codebase)
  • If compiled as Cpu, all GPU features are disabled regardless of hardware detection

Configuration Flow in WhisperEngine

The WhisperEngine struct orchestrates the detection sequence when initializing a transcription session.

Logging and Context Initialization

When WhisperEngine::new() executes, it calls detect_gpu_acceleration() (which wraps the backend detection) and logs the compiled configuration. After creating the acceleration profile, it outputs a diagnostic summary:

log::info!(
    "Whisper acceleration decision: compiled_backend={} runtime_detected_gpu={:?} use_gpu={} flash_attn={} gpu_device={}",
    acceleration.compiled_backend.as_str(),
    acceleration.runtime_detected_gpu,
    acceleration.use_gpu,
    acceleration.flash_attn,
    acceleration.gpu_device,
);

The complete integration flow from detection to model loading follows this pattern:

// 1. Determine compile-time backend at startup
let compiled = WhisperCompiledBackend::current();

// 2. Detect runtime GPU capabilities
let hardware = crate::audio::HardwareProfile::detect();

// 3. Build acceleration profile
let accel = whisper_context_acceleration_for(
    compiled,
    hardware.gpu_type,
    hardware.performance_tier,
);

// 4. Configure Whisper-rs context parameters
let ctx_params = WhisperContextParameters {
    use_gpu: accel.use_gpu,
    gpu_device: accel.gpu_device,
    flash_attn: accel.flash_attn,
    ..Default::default()
};

// 5. Initialize the transcription context
let ctx = WhisperContext::new_with_params(&model_path, ctx_params)?;

Summary

  • Compile-time detection via WhisperCompiledBackend::current() uses Rust cfg! macros to identify which GPU features (Metal, CUDA, Vulkan, HIP-BLAS) were enabled during the build process.
  • Runtime detection through HardwareProfile::detect() and detect_gpu() verifies actual GPU presence and classifies hardware into performance tiers.
  • Decision logic in whisper_context_acceleration_for() (located in frontend/src-tauri/src/whisper_engine/acceleration.rs) enables GPU acceleration only when both the binary supports it and hardware is detected.
  • Flash-Attention is automatically enabled for Metal and CUDA backends on High or Ultra tier machines to maximize transcription throughput.
  • Safety mechanisms ensure CPU fallback occurs gracefully when GPU features are not compiled or when compatible hardware is absent.

Frequently Asked Questions

How does Meetily choose between Metal, CUDA, and Vulkan?

Meetily does not dynamically choose between GPU backends at runtime. The backend is determined at compile time by the Cargo features passed during build. If you compile with --features cuda, the binary uses CUDA; with --features metal, it uses Metal on macOS; and with --features vulkan, it uses the Vulkan API. The WhisperCompiledBackend::current() method simply reports which feature was baked into the binary.

What happens if I compile with GPU support but no GPU is present?

If you compile with a GPU feature (e.g., --features cuda) but run on a machine without an NVIDIA GPU, the use_gpu flag will still be set to true in the acceleration profile. However, the runtime detection will report GpuType::None. While the code attempts to initialize GPU context, Whisper‑rs will typically fail gracefully or fall back, though Meetily's current implementation in whisper_engine.rs logs the discrepancy so you can diagnose the configuration mismatch.

When does Meetily enable Flash-Attention?

Flash-Attention is enabled only when three conditions are met: the compiled backend is either Metal or CUDA, the performance tier is classified as High or Ultra, and use_gpu is true. This optimization is specifically disabled for Vulkan and HIP-BLAS backends regardless of hardware capabilities, as implemented in the whisper_context_acceleration_for function.

Where is the GPU detection logic located in the codebase?

The primary GPU detection logic resides in three key locations: frontend/src-tauri/src/whisper_engine/acceleration.rs contains the WhisperCompiledBackend enum and decision functions; frontend/src-tauri/src/audio/hardware_detector.rs implements the runtime hardware probing via detect_gpu(); and frontend/src-tauri/src/whisper_engine/whisper_engine.rs orchestrates the initialization and logging of the acceleration profile when creating a new transcription engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →