How Meetily Manages Cross-Platform GPU Acceleration for Metal, CUDA, and Vulkan

Meetily uses compile-time feature flags to embed specific GPU backends (Metal, CUDA, Vulkan) into the whisper_rs transcription engine, then applies runtime hardware detection to automatically configure the optimal acceleration profile for macOS, Windows, or Linux.

Meetily implements a hybrid compile-time and runtime strategy to deliver GPU acceleration across diverse operating systems and hardware configurations. The open-source transcription pipeline in Zackriya-Solutions/meetily leverages feature-gated builds combined with dynamic hardware probing to seamlessly switch between Metal on macOS, CUDA on NVIDIA systems, and Vulkan for cross-platform GPU compute. This architecture ensures that the WhisperEngine automatically utilizes the best available hardware while maintaining fallback support for CPU-only operation.

Compile-Time Backend Selection via Cargo Features

Meetily’s GPU acceleration strategy begins at build time in frontend/src-tauri/Cargo.toml. The project defines optional features that map directly to specific GPU APIs, allowing the compiler to include only the relevant backend code for the target platform.

The available feature flags include:

  • metal – Enables Apple Metal acceleration for macOS
  • cuda – Enables NVIDIA CUDA support for Windows and Linux
  • vulkan – Enables Vulkan compute for cross-platform GPU acceleration
  • hipblas – Enables AMD GPU support via HIP
  • coreml – Enables Apple Core ML for Apple Silicon devices

# frontend/src-tauri/Cargo.toml (excerpt)

[features]
metal = ["whisper-rs/metal"]
cuda  = ["whisper-rs/cuda"]
vulkan = ["whisper-rs/vulkan"]
hipblas = ["whisper-rs/hipblas"]
coreml = ["whisper-rs/coreml"]

Platform-specific sections in Cargo.toml automatically enable sensible defaults. On macOS, the build system enables metal and coreml for Apple Silicon. On Windows and Linux, the default is openblas for CPU-only operation, though users can explicitly enable cuda, vulkan, or hipblas during compilation.

Runtime GPU Detection and Hardware Profiling

Once compiled, Meetily determines the actual hardware capabilities at runtime through the hardware_detector.rs module. This component inspects the operating system and environment variables to identify the available GPU type.

The detection logic in frontend/src-tauri/src/audio/hardware_detector.rs returns a HardwareProfile struct containing:

  • has_gpu_acceleration – Boolean indicating GPU availability
  • gpu_type – Enum variant (Metal, Cuda, Vulkan, OpenCL, or None)
  • performance_tier – Classification from Low to Ultra based on CPU cores and memory

The detection mechanism checks for macOS Metal support via system APIs, searches for CUDA_PATH or CUDA_HOME environment variables for NVIDIA GPUs, and verifies the presence of VULKAN_SDK or Vulkan library files (/usr/lib/libvulkan.so, vulkan-1.dll, etc.) to determine Vulkan availability.

use crate::audio::hardware_detector::HardwareProfile;

// Returns a cached profile describing the host’s CPU/GPU capabilities
let profile = HardwareProfile::detect();
println!(
    "GPU: {:?}, cores: {}, tier: {:?}",
    profile.gpu_type, profile.cpu_cores, profile.performance_tier
);

Mapping Compiled Backends to Runtime Configuration

The bridge between compile-time features and runtime detection occurs in frontend/src-tauri/src/whisper_engine/acceleration.rs. This module defines the WhisperCompiledBackend enum (Metal, Cuda, Vulkan, HipBlas, Cpu) and the critical whisper_context_acceleration_for function.

This function combines three inputs to produce a WhisperContextAcceleration configuration:

  1. Compiled backend – Which feature was enabled at build time
  2. Runtime-detected GPU – The actual GpuType discovered on the host
  3. Performance tier – Derived from system CPU cores and memory capacity

The resulting struct determines:

  • use_gpu – Whether to enable GPU acceleration
  • flash_attn – Automatically enabled for Metal and CUDA when a High or Ultra tier GPU is present
  • gpu_device – Currently defaults to 0 (first available device)
use meetily::frontend::src_tauri::whisper_engine::{
    acceleration::{whisper_context_acceleration_for, WhisperCompiledBackend},
    GpuType, PerformanceTier,
};

let compiled = WhisperCompiledBackend::Cuda; // Compiled with --features cuda
let runtime = GpuType::Cuda;
let tier = PerformanceTier::High;

let accel = whisper_context_acceleration_for(compiled, runtime, tier);
println!(
    "Using GPU: {}, FlashAttention: {}",
    accel.use_gpu, accel.flash_attn
);

The logic explicitly enables Flash-Attention only for Metal and CUDA backends when running on high-tier hardware, optimizing memory bandwidth and compute efficiency during transcription.

Engine Initialization and Backend Activation

When instantiating the transcription pipeline, WhisperEngine::new() in frontend/src-tauri/src/whisper_engine/whisper_engine.rs orchestrates the entire detection flow. The constructor calls detect_gpu_acceleration(), which logs the active backend and acceleration status.

The initialization sequence follows this order:

  1. Build phase – Cargo compiles with the appropriate feature flag (metal, cuda, or vulkan)
  2. Startup – WhisperEngine::new() invokes hardware detection
  3. Runtime probe – hardware_detector::detect_gpu() identifies the actual GPU type
  4. Configuration – whisper_context_acceleration_for() creates the acceleration struct
  5. Execution – The whisper-rs context is configured with the detected parameters

The log output indicates which backend is active, allowing developers to verify whether the system is utilizing Metal on macOS, CUDA on Windows, or falling back to CPU-only operation when GPU libraries are absent.

use meetily::frontend::src_tauri::whisper_engine::WhisperEngine;

let engine = WhisperEngine::new()?; // Logs active backend (Metal, CUDA, etc.)

Summary

  • Compile-time features in Cargo.toml control which GPU backends (Metal, CUDA, Vulkan) are included in the binary, preventing unnecessary dependencies.
  • Runtime detection in hardware_detector.rs probes the OS and environment variables to identify the actual GPU hardware and performance capabilities.
  • Configuration mapping in acceleration.rs bridges compiled features with detected hardware, automatically enabling GPU acceleration and Flash-Attention for high-tier devices.
  • Platform defaults automatically select Metal for macOS builds while requiring explicit opt-in for CUDA or Vulkan on Windows and Linux.
  • CPU fallback occurs automatically when no compatible GPU is detected or when compiled without GPU features.

Frequently Asked Questions

How does Meetily choose between Metal, CUDA, and Vulkan at runtime?

Meetily relies on the HardwareProfile::detect() method in hardware_detector.rs to inspect the host system. On macOS, it checks for Metal support via system APIs. On Windows and Linux, it looks for CUDA_PATH environment variables to detect NVIDIA GPUs, or searches for Vulkan SDK installations and library files (vulkan-1.dll, libvulkan.so). The detected GpuType is then matched against the WhisperCompiledBackend to ensure the runtime GPU matches the compiled backend capabilities.

Can I force CPU-only mode even if a GPU is detected?

Yes. CPU-only operation occurs automatically if Meetily is compiled without GPU feature flags (defaulting to the cpu backend). Additionally, the whisper_context_acceleration_for function in acceleration.rs determines GPU usage based on the compiled backend; if the backend is WhisperCompiledBackend::Cpu, the resulting configuration sets use_gpu to false regardless of what hardware is detected at runtime.

What triggers Flash-Attention activation in Meetily?

Flash-Attention is automatically enabled in acceleration.rs when two conditions are met: the compiled backend is either Metal or CUDA, and the detected PerformanceTier is High or Ultra. The logic uses pattern matching to set flash_attn to true only for these specific combinations, optimizing memory efficiency for capable Apple Silicon and NVIDIA hardware while avoiding compatibility issues on lower-tier or different backend systems.

Which file handles the detection of the CUDA environment on Windows?

The detection logic resides in frontend/src-tauri/src/audio/hardware_detector.rs. This module specifically checks for the presence of CUDA_PATH or CUDA_HOME environment variables, which are standard indicators of a CUDA installation on Windows systems. It also validates the existence of necessary library files to confirm that the CUDA runtime is actually functional before reporting GpuType::Cuda to the acceleration configuration system.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →