# Performance Considerations for Meetily: Optimizing Audio, Transcription, and Summarization

> Explore performance considerations for Meetily. Optimize audio, transcription, and summarization using GPU-accelerated Whisper, ring-buffer mixing, and chunked LLM streaming in Rust/Tauri.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: performance
- Published: 2026-07-27

---

**Meetily achieves low-latency local processing by combining 50 ms ring-buffer audio mixing, GPU-accelerated Whisper inference with automatic backend selection, zero-cost logging macros, and chunked LLM streaming in its Rust/Tauri core.**

Meetily is a privacy-first AI meeting assistant built on Tauri with a Rust backend and a Next.js/TypeScript frontend. Understanding the performance considerations for Meetily is critical when tuning real-time audio capture, on-device transcription, and local summarization workloads. The architecture targets deterministic latency through lock-free state management, platform-specific audio backends, and optional GPU acceleration.

## Audio Pipeline Performance Considerations for Meetily

### Ring-Buffer Mixing and VAD Filtering

In [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs), the `AudioMixerRingBuffer` aligns microphone and system audio streams in **50 ms windows**. This design avoids jitter and ensures deterministic latency during capture. The pipeline also applies **RMS-based ducking** to balance levels without expensive per-sample calculations, and runs a **Voice Activity Detection (VAD)** filter that discards non-speech chunks. According to the Meetily source code, VAD filtering reduces Whisper load by approximately **70 percent** by preventing silent audio from reaching the transcription engine.

```rust
// frontend/src-tauri/src/audio/pipeline.rs
let mixer = AudioMixerRingBuffer::new()
    .with_window_duration(Duration::from_millis(50)) // 50 ms windows
    .with_max_buffer(Duration::from_secs(2));       // Prevents unbounded growth

```

Smaller windows reduce latency while the two-second cap safeguards against memory overrun during burst capture.

```rust
// frontend/src-tauri/src/audio/vad.rs
let vad = VoiceActivityDetector::new()
    .threshold(0.35)          // Raise for noisy environments
    .silence_padding(300);    // Keep 300 ms of context before/after speech

```

Fine-tuning the threshold balances false positives against Whisper workload.

### Platform-Specific Audio Capture

Meetily isolates heavy platform dependencies into modular backends found under `frontend/src-tauri/src/audio/devices/platform/*.rs`. Each operating system uses its own low-level implementation: **ScreenCaptureKit** on macOS, **WASAPI** on Windows, and **ALSA/PulseAudio** on Linux. Because only the relevant platform code compiles and runs, binary size and startup time stay minimal.

### SIMD Resampling to 48 kHz

The [`frontend/src-tauri/src/audio_v2/resampler.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio_v2/resampler.rs) module pre-allocates buffers and performs **SIMD-based resampling** to Meetily’s required 48 kHz sample rate. Pre-allocation prevents runtime allocation spikes during active recording, keeping the audio thread stable.

## Whisper Transcription Performance Considerations for Meetily

### Automatic GPU Backend Selection

The [`frontend/src-tauri/src/whisper_engine/whisper_engine.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/whisper_engine/whisper_engine.rs) file implements a load-time detection routine that automatically picks the fastest Whisper backend. The priority order is **Metal → CUDA → Vulkan → CPU**. Hardware acceleration can be toggled at compile time with Cargo features such as `--features cuda` or `--features vulkan`, as implemented in [`frontend/src-tauri/src/whisper_engine/acceleration.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/whisper_engine/acceleration.rs). When a GPU is present, inference time drops from several seconds per minute of audio to sub-second latency.

### Model Caching and Parallel Batching

Meetily loads a Whisper model **once per session**; changing models requires an application restart to avoid repeated heavy I/O. The [`parallel_processor.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/parallel_processor.rs) batches incoming audio chunks to keep the GPU busy and maximize throughput.

```rust
// frontend/src-tauri/src/whisper_engine/whisper_engine.rs
let engine = WhisperEngine::new()
    .with_gpu_backend(GpuBackend::Auto)   // Auto-detects Metal, CUDA or Vulkan
    .load_model("medium")?;               // Load a medium-size model once

```

The `with_gpu_backend` call chooses the fastest available GPU backend; fallback to CPU is automatic if no GPU is detected.

## Summarization Engine Performance Considerations for Meetily

### Streaming JSON-RPC and Back-Pressure

In [`frontend/src-tauri/src/summary/summary_engine/model_manager.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/summary/summary_engine/model_manager.rs), Meetily communicates with local LLMs via **streaming JSON-RPC** with back-pressure. This approach prevents memory spikes when processing long transcripts. The engine also splits transcripts into **chunked prompts** that fit within model context limits, enabling progressive summarization.

### Local Model Storage

Local LLM storage removes network latency entirely. Models are cached in the OS-specific application data directory:

- macOS: `~/Library/Application Support/Meetily/models/`
- Windows: `%APPDATA%\Meetily\models\`

```typescript
// frontend/src/app/page.tsx
const summary = await invoke<string>('summarize_meeting', {
  meeting_id: currentMeeting.id,
  chunk_size: 5000,            // Send 5 k character chunks
});

```

Chunked streaming keeps LLM payloads within model limits and reduces memory pressure.

## Runtime Optimizations in the Rust Core

### Zero-Cost Debug Logging

The `perf_debug!()` macro, defined in [`frontend/src-tauri/src/lib.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/lib.rs), expands to a no-op in release builds. Developers can leave verbose instrumentation in hot paths without paying runtime costs in production.

```rust
// Any hot path, e.g., src-tauri/src/audio/level_monitor.rs
perf_debug!("Audio level: {:.2} dB", rms);

```

In a release build this line compiles out, incurring zero runtime cost.

### Thread-Safe Shared State

[`frontend/src-tauri/src/audio/recording_state.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/recording_state.rs) uses `Arc<AtomicBool>` for simple flags and `Arc<RwLock<Option<mpsc::UnboundedSender<AudioChunk>>>>` for complex cross-task coordination. This pattern enables lock-free reads on the fast path while preserving safe mutable access to the audio channel when needed.

## Summary

- **Ring-buffer mixing** in 50 ms windows keeps audio latency deterministic, while VAD filtering cuts Whisper workload by roughly 70 percent.
- **GPU auto-detection** (Metal → CUDA → Vulkan → CPU) and per-session model caching minimize transcription latency.
- **Zero-cost logging** via `perf_debug!()` and pre-allocated SIMD resampling remove unnecessary runtime overhead.
- **Platform-specific backends** compile only the code required for the host OS, shrinking binaries and improving startup time.
- **Chunked streaming and local model storage** in the summarization engine balance memory usage and throughput for LLM inference.

## Frequently Asked Questions

### How does Meetily reduce CPU usage during meetings?

Meetily reduces CPU usage by running a Voice Activity Detection (VAD) filter in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs) that discards non-speech audio before it reaches the Whisper engine, cutting processing load by approximately 70 percent. The Rust core also uses zero-cost `perf_debug!()` macros that compile away entirely in release builds.

### What GPU backends does Meetily support for transcription?

According to the Meetily source code, Whisper inference supports **Metal**, **CUDA**, and **Vulkan**, selected automatically at load time in [`frontend/src-tauri/src/whisper_engine/whisper_engine.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/whisper_engine/whisper_engine.rs). CPU fallback is automatic when no compatible GPU is detected, and hardware features are enabled at compile time via Cargo flags in [`frontend/src-tauri/src/whisper_engine/acceleration.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/whisper_engine/acceleration.rs).

### How does Meetily keep audio latency low?

The application maintains low latency through **50 ms ring-buffer mixing** in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs), platform-specific native capture APIs under `frontend/src-tauri/src/audio/devices/platform/*.rs`, and SIMD resampling in [`frontend/src-tauri/src/audio_v2/resampler.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio_v2/resampler.rs). The buffer is also capped at two seconds to prevent unbounded memory growth during burst capture.

### Where does Meetily store local LLM models?

Local models are cached in the OS-specific application support directory: `~/Library/Application Support/Meetily/models/` on macOS and `%APPDATA%\Meetily\models\` on Windows. This local storage strategy eliminates network round-trips and is managed by [`frontend/src-tauri/src/summary/summary_engine/model_manager.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/summary/summary_engine/model_manager.rs).