How Meetily Monitors GPU Usage During Transcription When Acceleration Is Enabled
Meetily uses a three-stage pipeline for GPU monitoring: detection at startup via HardwareDetector, acceleration configuration in WhisperContextAcceleration, and indirect runtime safeguards through CPU/memory/temperature polling in SystemMonitor—it does not query GPU load directly during transcription.
Meetily is an open-source meeting intelligence application built with Tauri and Rust that leverages OpenAI's Whisper for speech-to-text transcription. When GPU acceleration is enabled, the application implements a sophisticated resource management strategy to ensure stable performance without overwhelming the host system. This article examines exactly how Meetily monitors and safeguards GPU usage during transcription, based on the source code in the Zackriya-Solutions/meetily repository.
Three-Stage GPU Monitoring Architecture
Meetily does not poll GPU metrics every millisecond during transcription. Instead, it follows a deliberate three-stage approach that establishes GPU capability, configures acceleration parameters, and enforces system-wide resource constraints.
Stage 1: GPU Detection at Startup
The HardwareDetector scans the operating system once during application initialization to identify available GPU backends. Located in src-tauri/src/audio/hardware_detector.rs, this module probes for:
- Metal (macOS)
- CUDA (NVIDIA)
- Vulkan (AMD/Intel)
- OpenCL (fallback)
The detector records the result as a GpuType enum variant (Metal, Cuda, Vulkan, OpenCL, or None) and assigns a PerformanceTier based on detected capabilities. This single detection event provides the foundation for all subsequent GPU decisions.
let hardware_profile = crate::audio::HardwareProfile::detect(); // <- gives GpuType
Stage 2: Acceleration Decision
When WhisperEngine::load_model executes, it constructs a WhisperContextAcceleration struct that merges two critical inputs:
- Compiled backend:
WhisperCompiledBackend::current()— the GPU feature flags enabled at build time - Runtime-detected GPU:
hardware_profile.gpu_typefrom the startup detection
This logic resides primarily in src-tauri/src/whisper_engine/acceleration.rs with the context creation in src-tauri/src/whisper_engine/whisper_engine.rs (lines 52-78).
let acceleration = whisper_context_acceleration_for(
WhisperCompiledBackend::current(),
hardware_profile.gpu_type,
hardware_profile.performance_tier,
);
let ctx = WhisperContext::new_with_params(&model_path, WhisperContextParameters {
use_gpu: acceleration.use_gpu,
gpu_device: acceleration.gpu_device,
flash_attn: acceleration.flash_attn,
..Default::default()
})?;
The resulting decision is logged explicitly (lines 298-304 of whisper_engine.rs):
Whisper acceleration decision: compiled_backend=CUDA runtime_detected_gpu=Cuda use_gpu=true flash_attn=true gpu_device=0
Stage 3: Runtime Resource Guard
During actual transcription, Meetily employs a system monitor rather than direct GPU polling. The SystemMonitor in src-tauri/src/whisper_engine/system_monitor.rs checks CPU usage, memory pressure, and system temperature every resource_check_interval_ms (default: 10 seconds).
If constraints are violated, the monitor emits a ProcessingEvent::ResourceConstraint and pauses workers; when metrics recover, processing resumes automatically.
Because whisper-rs bindings do not expose GPU utilization metrics, Meetily relies on indirect safeguards: GPU work consumes memory and CPU cycles for data transfers, so memory-percent and CPU-percent thresholds act as proxy guards for GPU health.
// Inside parallel_processor.rs (lines 72-84)
while let Some(chunk) = receiver.recv().await {
if monitor.check_resource_constraints().await?.is_healthy() {
self.process_chunk(chunk).await?;
} else {
self.pause_with_backoff().await;
}
}
How Resource Constraints Protect GPU Transcription
The SystemMonitor enforces default thresholds that protect both CPU and GPU paths:
| Metric | Default Threshold | Action When Exceeded |
|---|---|---|
| Memory usage | 70% | Pause processing, queue chunks |
| CPU usage | 80% | Pause processing, queue chunks |
| Temperature | Platform-specific | Emergency pause with extended backoff |
These constraints are configurable but designed to prevent the cascading failures that occur when GPU memory is exhausted or thermal limits are approached.
Querying Current GPU Status
You can inspect the runtime GPU configuration through Meetily's internal API:
use crate::audio::HardwareProfile;
use crate::whisper_engine::acceleration::{whisper_context_acceleration_for, WhisperCompiledBackend};
pub async fn get_gpu_status() -> Result<String> {
// Detect hardware once
let profile = HardwareProfile::detect();
// Build the acceleration struct
let accel = whisper_context_acceleration_for(
WhisperCompiledBackend::current(),
profile.gpu_type,
profile.performance_tier,
);
// Human-readable description
Ok(format!(
"GPU: {:?}, enabled: {}, flash_attn: {}",
accel.runtime_detected_gpu,
accel.use_gpu,
accel.flash_attn
))
}
Manual Resource Checking During Transcription
For debugging or custom integrations, you can instantiate the monitor directly:
use crate::whisper_engine::system_monitor::create_system_monitor;
#[tokio::main]
async fn main() -> anyhow::Result<()> {
let monitor = create_system_monitor();
// Refresh the cache (required before the first read)
monitor.refresh_system_info().await?;
// Show CPU & memory usage
let resources = monitor.get_current_resources().await?;
println!(
"CPU {:.1}% | Mem {:.1}% (avail {} MiB)",
resources.cpu_usage_percent,
resources.memory_used_percent,
resources.available_memory_mb
);
// Verify constraints
let status = monitor.check_resource_constraints().await?;
if !status.is_healthy() {
println!("⚠️ Resource limits exceeded: {:?}", status.warnings);
}
Ok(())
}
Pausing Transcription on Resource Pressure
The ParallelProcessor integrates with SystemMonitor to enable automatic circuit-breaking:
// Inside a Tauri command that holds a ParallelProcessor instance `processor`
async fn maybe_pause(processor: Arc<ParallelProcessor>) {
let monitor = processor.system_monitor.clone();
let constraints = monitor.check_resource_constraints().await;
if let Ok(status) = constraints {
if !status.is_healthy() {
processor.pause_processing().await;
log::warn!("Transcription paused – {}", status.warnings.join(", "));
}
}
}
Key Source Files
| File | Responsibility |
|---|---|
src-tauri/src/audio/hardware_detector.rs |
Detects Metal, CUDA, Vulkan, OpenCL; provides GpuType and PerformanceTier |
src-tauri/src/whisper_engine/acceleration.rs |
Combiles compiled backend and runtime GPU info into WhisperContextAcceleration |
src-tauri/src/whisper_engine/whisper_engine.rs |
Creates Whisper context with GPU flags; logs acceleration decision |
src-tauri/src/whisper_engine/system_monitor.rs |
Polls CPU, memory, temperature; decides if processing may continue |
src-tauri/src/whisper_engine/parallel_processor.rs |
Runs transcription workers; integrates SystemMonitor for pause/resume |
llama-helper/src/main.rs |
Demonstrates VRAM querying logic (nvidia-smi, sysctl) used during build configuration |
Summary
- GPU presence is detected once at startup via
HardwareProfile::detect()inhardware_detector.rs - GPU activation is decided by combining compile-time flags with runtime detection in
acceleration.rs, then logged during context creation - Runtime safety is enforced through
SystemMonitorpolling CPU, memory, and temperature—indirectly protecting GPU workloads sincewhisper-rsdoes not expose GPU metrics - Automatic fallback occurs if resource constraints are violated, with the parallel transcription loop respecting pause/resume signals to maintain system health
Frequently Asked Questions
Does Meetily monitor GPU utilization in real-time during transcription?
No. Meetily does not query GPU load percentages or VRAM usage directly during transcription because the whisper-rs bindings do not expose these metrics. Instead, it monitors CPU usage, system memory, and temperature every 10 seconds as indirect indicators. If these metrics exceed thresholds, transcription pauses automatically, which protects GPU-backed processes from resource exhaustion.
What happens if a GPU driver crashes or becomes unavailable during transcription?
If GPU acceleration was enabled at context creation but the GPU subsequently fails, the compiled backend detection would have already established the GPU type. The SystemMonitor would detect elevated CPU usage (from fallback processing) and memory pressure, potentially triggering a pause. For a cleaner fallback, the application must be restarted to re-run HardwareProfile::detect() and disable use_gpu in the WhisperContextParameters.
How can I verify that GPU acceleration is actually active?
Check the application logs for the line emitted by whisper_engine.rs around line 298-304. It explicitly states: Whisper acceleration decision: compiled_backend=X runtime_detected_gpu=Y use_gpu=Z flash_attn=W gpu_device=V. If use_gpu=true and your expected backend matches runtime_detected_gpu, acceleration is active. You can also programmatically call get_gpu_status() as shown in the code examples above.
Can I adjust the resource monitoring thresholds?
Yes. The SystemMonitor accepts configuration for resource_check_interval_ms (default 10000), memory threshold percentage, and CPU threshold percentage. These values are typically passed during ParallelProcessor initialization. Modify them in src-tauri/src/whisper_engine/system_monitor.rs or through your application's configuration layer if exposed via Tauri commands.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →