How Meetily Monitors GPU Usage During Transcription When Acceleration Is Enabled

Meetily uses a three-stage pipeline for GPU monitoring: detection at startup via HardwareDetector, acceleration configuration in WhisperContextAcceleration, and indirect runtime safeguards through CPU/memory/temperature polling in SystemMonitor—it does not query GPU load directly during transcription.

Meetily is an open-source meeting intelligence application built with Tauri and Rust that leverages OpenAI's Whisper for speech-to-text transcription. When GPU acceleration is enabled, the application implements a sophisticated resource management strategy to ensure stable performance without overwhelming the host system. This article examines exactly how Meetily monitors and safeguards GPU usage during transcription, based on the source code in the Zackriya-Solutions/meetily repository.

Three-Stage GPU Monitoring Architecture

Meetily does not poll GPU metrics every millisecond during transcription. Instead, it follows a deliberate three-stage approach that establishes GPU capability, configures acceleration parameters, and enforces system-wide resource constraints.

Stage 1: GPU Detection at Startup

The HardwareDetector scans the operating system once during application initialization to identify available GPU backends. Located in src-tauri/src/audio/hardware_detector.rs, this module probes for:

  • Metal (macOS)
  • CUDA (NVIDIA)
  • Vulkan (AMD/Intel)
  • OpenCL (fallback)

The detector records the result as a GpuType enum variant (Metal, Cuda, Vulkan, OpenCL, or None) and assigns a PerformanceTier based on detected capabilities. This single detection event provides the foundation for all subsequent GPU decisions.

let hardware_profile = crate::audio::HardwareProfile::detect();   // <- gives GpuType

Stage 2: Acceleration Decision

When WhisperEngine::load_model executes, it constructs a WhisperContextAcceleration struct that merges two critical inputs:

  1. Compiled backend: WhisperCompiledBackend::current() — the GPU feature flags enabled at build time
  2. Runtime-detected GPU: hardware_profile.gpu_type from the startup detection

This logic resides primarily in src-tauri/src/whisper_engine/acceleration.rs with the context creation in src-tauri/src/whisper_engine/whisper_engine.rs (lines 52-78).

let acceleration = whisper_context_acceleration_for(
    WhisperCompiledBackend::current(),
    hardware_profile.gpu_type,
    hardware_profile.performance_tier,
);
let ctx = WhisperContext::new_with_params(&model_path, WhisperContextParameters {
    use_gpu: acceleration.use_gpu,
    gpu_device: acceleration.gpu_device,
    flash_attn: acceleration.flash_attn,
    ..Default::default()
})?;

The resulting decision is logged explicitly (lines 298-304 of whisper_engine.rs):


Whisper acceleration decision: compiled_backend=CUDA runtime_detected_gpu=Cuda use_gpu=true flash_attn=true gpu_device=0

Stage 3: Runtime Resource Guard

During actual transcription, Meetily employs a system monitor rather than direct GPU polling. The SystemMonitor in src-tauri/src/whisper_engine/system_monitor.rs checks CPU usage, memory pressure, and system temperature every resource_check_interval_ms (default: 10 seconds).

If constraints are violated, the monitor emits a ProcessingEvent::ResourceConstraint and pauses workers; when metrics recover, processing resumes automatically.

Because whisper-rs bindings do not expose GPU utilization metrics, Meetily relies on indirect safeguards: GPU work consumes memory and CPU cycles for data transfers, so memory-percent and CPU-percent thresholds act as proxy guards for GPU health.

// Inside parallel_processor.rs (lines 72-84)
while let Some(chunk) = receiver.recv().await {
    if monitor.check_resource_constraints().await?.is_healthy() {
        self.process_chunk(chunk).await?;
    } else {
        self.pause_with_backoff().await;
    }
}

How Resource Constraints Protect GPU Transcription

The SystemMonitor enforces default thresholds that protect both CPU and GPU paths:

Metric Default Threshold Action When Exceeded
Memory usage 70% Pause processing, queue chunks
CPU usage 80% Pause processing, queue chunks
Temperature Platform-specific Emergency pause with extended backoff

These constraints are configurable but designed to prevent the cascading failures that occur when GPU memory is exhausted or thermal limits are approached.

Querying Current GPU Status

You can inspect the runtime GPU configuration through Meetily's internal API:

use crate::audio::HardwareProfile;
use crate::whisper_engine::acceleration::{whisper_context_acceleration_for, WhisperCompiledBackend};

pub async fn get_gpu_status() -> Result<String> {
    // Detect hardware once
    let profile = HardwareProfile::detect();

    // Build the acceleration struct
    let accel = whisper_context_acceleration_for(
        WhisperCompiledBackend::current(),
        profile.gpu_type,
        profile.performance_tier,
    );

    // Human-readable description
    Ok(format!(
        "GPU: {:?}, enabled: {}, flash_attn: {}",
        accel.runtime_detected_gpu,
        accel.use_gpu,
        accel.flash_attn
    ))
}

Manual Resource Checking During Transcription

For debugging or custom integrations, you can instantiate the monitor directly:

use crate::whisper_engine::system_monitor::create_system_monitor;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let monitor = create_system_monitor();

    // Refresh the cache (required before the first read)
    monitor.refresh_system_info().await?;

    // Show CPU & memory usage
    let resources = monitor.get_current_resources().await?;
    println!(
        "CPU {:.1}% | Mem {:.1}% (avail {} MiB)",
        resources.cpu_usage_percent,
        resources.memory_used_percent,
        resources.available_memory_mb
    );

    // Verify constraints
    let status = monitor.check_resource_constraints().await?;
    if !status.is_healthy() {
        println!("⚠️ Resource limits exceeded: {:?}", status.warnings);
    }

    Ok(())
}

Pausing Transcription on Resource Pressure

The ParallelProcessor integrates with SystemMonitor to enable automatic circuit-breaking:

// Inside a Tauri command that holds a ParallelProcessor instance `processor`
async fn maybe_pause(processor: Arc<ParallelProcessor>) {
    let monitor = processor.system_monitor.clone();
    let constraints = monitor.check_resource_constraints().await;
    if let Ok(status) = constraints {
        if !status.is_healthy() {
            processor.pause_processing().await;
            log::warn!("Transcription paused – {}", status.warnings.join(", "));
        }
    }
}

Key Source Files

File Responsibility
src-tauri/src/audio/hardware_detector.rs Detects Metal, CUDA, Vulkan, OpenCL; provides GpuType and PerformanceTier
src-tauri/src/whisper_engine/acceleration.rs Combiles compiled backend and runtime GPU info into WhisperContextAcceleration
src-tauri/src/whisper_engine/whisper_engine.rs Creates Whisper context with GPU flags; logs acceleration decision
src-tauri/src/whisper_engine/system_monitor.rs Polls CPU, memory, temperature; decides if processing may continue
src-tauri/src/whisper_engine/parallel_processor.rs Runs transcription workers; integrates SystemMonitor for pause/resume
llama-helper/src/main.rs Demonstrates VRAM querying logic (nvidia-smi, sysctl) used during build configuration

Summary

  • GPU presence is detected once at startup via HardwareProfile::detect() in hardware_detector.rs
  • GPU activation is decided by combining compile-time flags with runtime detection in acceleration.rs, then logged during context creation
  • Runtime safety is enforced through SystemMonitor polling CPU, memory, and temperature—indirectly protecting GPU workloads since whisper-rs does not expose GPU metrics
  • Automatic fallback occurs if resource constraints are violated, with the parallel transcription loop respecting pause/resume signals to maintain system health

Frequently Asked Questions

Does Meetily monitor GPU utilization in real-time during transcription?

No. Meetily does not query GPU load percentages or VRAM usage directly during transcription because the whisper-rs bindings do not expose these metrics. Instead, it monitors CPU usage, system memory, and temperature every 10 seconds as indirect indicators. If these metrics exceed thresholds, transcription pauses automatically, which protects GPU-backed processes from resource exhaustion.

What happens if a GPU driver crashes or becomes unavailable during transcription?

If GPU acceleration was enabled at context creation but the GPU subsequently fails, the compiled backend detection would have already established the GPU type. The SystemMonitor would detect elevated CPU usage (from fallback processing) and memory pressure, potentially triggering a pause. For a cleaner fallback, the application must be restarted to re-run HardwareProfile::detect() and disable use_gpu in the WhisperContextParameters.

How can I verify that GPU acceleration is actually active?

Check the application logs for the line emitted by whisper_engine.rs around line 298-304. It explicitly states: Whisper acceleration decision: compiled_backend=X runtime_detected_gpu=Y use_gpu=Z flash_attn=W gpu_device=V. If use_gpu=true and your expected backend matches runtime_detected_gpu, acceleration is active. You can also programmatically call get_gpu_status() as shown in the code examples above.

Can I adjust the resource monitoring thresholds?

Yes. The SystemMonitor accepts configuration for resource_check_interval_ms (default 10000), memory threshold percentage, and CPU threshold percentage. These values are typically passed during ParallelProcessor initialization. Modify them in src-tauri/src/whisper_engine/system_monitor.rs or through your application's configuration layer if exposed via Tauri commands.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →