Performance Considerations for Meetily: Optimizing Audio, Transcription, and Summarization

Meetily achieves low-latency local processing by combining 50 ms ring-buffer audio mixing, GPU-accelerated Whisper inference with automatic backend selection, zero-cost logging macros, and chunked LLM streaming in its Rust/Tauri core.

Meetily is a privacy-first AI meeting assistant built on Tauri with a Rust backend and a Next.js/TypeScript frontend. Understanding the performance considerations for Meetily is critical when tuning real-time audio capture, on-device transcription, and local summarization workloads. The architecture targets deterministic latency through lock-free state management, platform-specific audio backends, and optional GPU acceleration.

Audio Pipeline Performance Considerations for Meetily

Ring-Buffer Mixing and VAD Filtering

In frontend/src-tauri/src/audio/pipeline.rs, the AudioMixerRingBuffer aligns microphone and system audio streams in 50 ms windows. This design avoids jitter and ensures deterministic latency during capture. The pipeline also applies RMS-based ducking to balance levels without expensive per-sample calculations, and runs a Voice Activity Detection (VAD) filter that discards non-speech chunks. According to the Meetily source code, VAD filtering reduces Whisper load by approximately 70 percent by preventing silent audio from reaching the transcription engine.

// frontend/src-tauri/src/audio/pipeline.rs
let mixer = AudioMixerRingBuffer::new()
    .with_window_duration(Duration::from_millis(50)) // 50 ms windows
    .with_max_buffer(Duration::from_secs(2));       // Prevents unbounded growth

Smaller windows reduce latency while the two-second cap safeguards against memory overrun during burst capture.

// frontend/src-tauri/src/audio/vad.rs
let vad = VoiceActivityDetector::new()
    .threshold(0.35)          // Raise for noisy environments
    .silence_padding(300);    // Keep 300 ms of context before/after speech

Fine-tuning the threshold balances false positives against Whisper workload.

Platform-Specific Audio Capture

Meetily isolates heavy platform dependencies into modular backends found under frontend/src-tauri/src/audio/devices/platform/*.rs. Each operating system uses its own low-level implementation: ScreenCaptureKit on macOS, WASAPI on Windows, and ALSA/PulseAudio on Linux. Because only the relevant platform code compiles and runs, binary size and startup time stay minimal.

SIMD Resampling to 48 kHz

The frontend/src-tauri/src/audio_v2/resampler.rs module pre-allocates buffers and performs SIMD-based resampling to Meetily’s required 48 kHz sample rate. Pre-allocation prevents runtime allocation spikes during active recording, keeping the audio thread stable.

Whisper Transcription Performance Considerations for Meetily

Automatic GPU Backend Selection

The frontend/src-tauri/src/whisper_engine/whisper_engine.rs file implements a load-time detection routine that automatically picks the fastest Whisper backend. The priority order is Metal → CUDA → Vulkan → CPU. Hardware acceleration can be toggled at compile time with Cargo features such as --features cuda or --features vulkan, as implemented in frontend/src-tauri/src/whisper_engine/acceleration.rs. When a GPU is present, inference time drops from several seconds per minute of audio to sub-second latency.

Model Caching and Parallel Batching

Meetily loads a Whisper model once per session; changing models requires an application restart to avoid repeated heavy I/O. The parallel_processor.rs batches incoming audio chunks to keep the GPU busy and maximize throughput.

// frontend/src-tauri/src/whisper_engine/whisper_engine.rs
let engine = WhisperEngine::new()
    .with_gpu_backend(GpuBackend::Auto)   // Auto-detects Metal, CUDA or Vulkan
    .load_model("medium")?;               // Load a medium-size model once

The with_gpu_backend call chooses the fastest available GPU backend; fallback to CPU is automatic if no GPU is detected.

Summarization Engine Performance Considerations for Meetily

Streaming JSON-RPC and Back-Pressure

In frontend/src-tauri/src/summary/summary_engine/model_manager.rs, Meetily communicates with local LLMs via streaming JSON-RPC with back-pressure. This approach prevents memory spikes when processing long transcripts. The engine also splits transcripts into chunked prompts that fit within model context limits, enabling progressive summarization.

Local Model Storage

Local LLM storage removes network latency entirely. Models are cached in the OS-specific application data directory:

  • macOS: ~/Library/Application Support/Meetily/models/
  • Windows: %APPDATA%\Meetily\models\
// frontend/src/app/page.tsx
const summary = await invoke<string>('summarize_meeting', {
  meeting_id: currentMeeting.id,
  chunk_size: 5000,            // Send 5 k character chunks
});

Chunked streaming keeps LLM payloads within model limits and reduces memory pressure.

Runtime Optimizations in the Rust Core

Zero-Cost Debug Logging

The perf_debug!() macro, defined in frontend/src-tauri/src/lib.rs, expands to a no-op in release builds. Developers can leave verbose instrumentation in hot paths without paying runtime costs in production.

// Any hot path, e.g., src-tauri/src/audio/level_monitor.rs
perf_debug!("Audio level: {:.2} dB", rms);

In a release build this line compiles out, incurring zero runtime cost.

Thread-Safe Shared State

frontend/src-tauri/src/audio/recording_state.rs uses Arc<AtomicBool> for simple flags and Arc<RwLock<Option<mpsc::UnboundedSender<AudioChunk>>>> for complex cross-task coordination. This pattern enables lock-free reads on the fast path while preserving safe mutable access to the audio channel when needed.

Summary

  • Ring-buffer mixing in 50 ms windows keeps audio latency deterministic, while VAD filtering cuts Whisper workload by roughly 70 percent.
  • GPU auto-detection (Metal → CUDA → Vulkan → CPU) and per-session model caching minimize transcription latency.
  • Zero-cost logging via perf_debug!() and pre-allocated SIMD resampling remove unnecessary runtime overhead.
  • Platform-specific backends compile only the code required for the host OS, shrinking binaries and improving startup time.
  • Chunked streaming and local model storage in the summarization engine balance memory usage and throughput for LLM inference.

Frequently Asked Questions

How does Meetily reduce CPU usage during meetings?

Meetily reduces CPU usage by running a Voice Activity Detection (VAD) filter in frontend/src-tauri/src/audio/pipeline.rs that discards non-speech audio before it reaches the Whisper engine, cutting processing load by approximately 70 percent. The Rust core also uses zero-cost perf_debug!() macros that compile away entirely in release builds.

What GPU backends does Meetily support for transcription?

According to the Meetily source code, Whisper inference supports Metal, CUDA, and Vulkan, selected automatically at load time in frontend/src-tauri/src/whisper_engine/whisper_engine.rs. CPU fallback is automatic when no compatible GPU is detected, and hardware features are enabled at compile time via Cargo flags in frontend/src-tauri/src/whisper_engine/acceleration.rs.

How does Meetily keep audio latency low?

The application maintains low latency through 50 ms ring-buffer mixing in frontend/src-tauri/src/audio/pipeline.rs, platform-specific native capture APIs under frontend/src-tauri/src/audio/devices/platform/*.rs, and SIMD resampling in frontend/src-tauri/src/audio_v2/resampler.rs. The buffer is also capped at two seconds to prevent unbounded memory growth during burst capture.

Where does Meetily store local LLM models?

Local models are cached in the OS-specific application support directory: ~/Library/Application Support/Meetily/models/ on macOS and %APPDATA%\Meetily\models\ on Windows. This local storage strategy eliminates network round-trips and is managed by frontend/src-tauri/src/summary/summary_engine/model_manager.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →