How Does Meetily's System Monitor Track GPU Utilization During Transcription Workloads?
Contrary to typical resource monitoring, Meetily's SystemMonitor does not poll real-time GPU utilization metrics; instead, it relies on compile-time detection and CPU/memory-based throttling to manage transcription workloads safely.
The Meetily transcription pipeline, housed in the Zackriya-Solutions/meetily repository, employs a SystemMonitor component to prevent system overload during intensive audio processing. While the monitor actively tracks CPU and memory health using the sysinfo crate, it manages GPU acceleration through one-time hardware detection rather than continuous utilization tracking.
How SystemMonitor Gathers CPU and Memory Metrics
Located in frontend/src-tauri/src/whisper_engine/system_monitor.rs, the SystemMonitor struct serves as the primary interface for resource polling. The component initializes a sysinfo::System instance and refreshes it periodically throughout the transcription lifecycle.
The monitor tracks three key host metrics:
- Memory utilization – Total and used system memory expressed as a percentage.
- CPU usage – Average load across all logical cores alongside the total core count.
- CPU temperature – Platform-specific thermal readings (currently implemented as a stub).
The struct enforces ResourceLimits with conservative defaults: 70% memory usage, 80% CPU usage, and 85°C temperature. Developers can instantiate custom thresholds via SystemMonitor::with_limits before launching transcription jobs.
Why GPU Utilization Is Not Monitored in Real-Time
Despite supporting GPU-accelerated transcription, the SystemMonitor explicitly excludes GPU metrics from its polling cycle. The get_current_resources method returns a SystemResources struct containing only memory and CPU statistics, omitting GPU load percentages, VRAM usage, or GPU temperature entirely.
This architectural separation ensures that resource throttling decisions remain independent of graphics subsystem variability. The monitor's calculate_safe_worker_count method derives parallel worker limits based solely on available CPU cycles and free memory, capping the count at four workers as a hard safety ceiling. The logic assumes that if a GPU is present and the Whisper backend is compiled with CUDA, Vulkan, or Metal support, the hardware will handle throughput optimization without requiring runtime utilization feedback.
GPU Detection and Acceleration Layer
GPU capability is determined elsewhere in the codebase through static analysis and hardware probing. In frontend/src-tauri/src/whisper_engine/acceleration.rs, the compiled Whisper backend is inspected and stored as a GpuType enum variant (CUDA, Vulkan, Metal, or CPU).
The HardwareDetector in frontend/src-tauri/src/audio/hardware_detector.rs performs a one-time system scan via has_gpu_acceleration, recording whether discrete or integrated graphics are available. This boolean flag and GPU type inform the WhisperEngine (whisper_engine.rs) whether to initialize GPU buffers, but the engine never queries current GPU load during transcription. Note that llama-helper/src/main.rs contains example code for external GPU querying (--query-gpu=memory.free), but this utility is not integrated into the core monitoring logic.
Calculating Safe Worker Counts
The calculate_safe_worker_count method implements the core throttling logic for transcription workloads. When invoked during a job, it executes the following sequence:
- Refresh system info – Calls
refresh_system_infoto update thesysinfosnapshot. - Read current metrics – Invokes
get_current_resourcesto obtain current CPU and memory percentages. - Check constraints – Uses
check_resource_constraintsto generate aResourceStatuscontaining violation flags and warning messages. - Compute workers – Derives the maximum safe parallel workers based on remaining CPU and memory headroom, ensuring at least one worker runs while never exceeding four.
If check_resource_constraints returns can_proceed: false, the engine pauses spawning new workers and logs the specific resource violations until system headroom improves.
// Create a monitor with default limits
let monitor = SystemMonitor::new();
// Refresh the snapshot (called periodically during transcription)
monitor.refresh_system_info().await?;
// Obtain current resource usage
let resources = monitor.get_current_resources().await?;
println!(
"Mem {:.1}% | CPU {:.1}% | Cores {}",
resources.memory_used_percent,
resources.cpu_usage_percent,
resources.cpu_cores
);
// Verify we are within safe limits before starting a new worker
let status = monitor.check_resource_constraints().await?;
if !status.can_proceed {
eprintln!("Resource constraints violated: {:?}", status.warnings);
}
// Compute how many parallel workers we may spawn safely
let safe_workers = monitor.calculate_safe_worker_count().await?;
println!("Launching {} transcription workers", safe_workers);
Summary
- Meetily's SystemMonitor focuses exclusively on CPU and memory health during transcription workloads, utilizing the
sysinfocrate for periodic polling. - GPU utilization is not tracked in real-time; the monitor relies on compile-time backend detection in
acceleration.rsand one-time hardware probing inhardware_detector.rs. - The
calculate_safe_worker_countmethod caps parallel transcription workers at four based on CPU and memory constraints, assuming GPU presence improves throughput without requiring load monitoring. - ResourceLimits (default: 70% memory, 80% CPU, 85°C) can be customized via
SystemMonitor::with_limitsto suit different hardware configurations. - GPU-enabled transcription is automatic when the Whisper backend is compiled with CUDA, Vulkan, or Metal support, but the monitor itself never queries GPU metrics during runtime.
Frequently Asked Questions
Does Meetily track GPU temperature during transcription?
No, the SystemMonitor only tracks CPU temperature through the sysinfo crate, and even that feature is currently a platform-specific stub. GPU thermal data is not polled during transcription workloads, nor is it used to throttle worker counts.
How does Meetily determine if a GPU is available for transcription?
The system checks GPU availability once at startup through the HardwareDetector in audio/hardware_detector.rs, which sets a has_gpu_acceleration flag. Additionally, whisper_engine/acceleration.rs inspects the compiled Whisper backend to determine if CUDA, Vulkan, or Metal support is available, storing the result as a GpuType enum.
What happens if CPU and memory limits are exceeded during a transcription job?
When check_resource_constraints detects violations against the configured ResourceLimits, it returns a ResourceStatus with can_proceed: false and specific warning messages. The Whisper engine uses this signal to pause spawning new workers until resources free up, preventing system overload while allowing existing workers to complete.
Can the resource limits for SystemMonitor be customized?
Yes, developers can instantiate a monitor with custom thresholds using SystemMonitor::with_limits, passing a ResourceLimits struct with desired memory, CPU, and temperature ceilings instead of relying on the default 70%/80%/85°C values. This allows tuning for high-performance workstations or resource-constrained environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →