Meetily Technology Stack: A Deep Dive Into the Privacy-First AI Meeting Assistant Architecture
Meetily is built with Tauri 2.x, Next.js 14 with React 18, Rust, whisper-cpp for local transcription, and SQLite for offline storage.
Meetily is a privacy-first AI meeting assistant developed by Zackriya-Solutions that runs entirely on the user's machine. Its technology stack combines modern web technologies with systems programming to deliver cross-platform, GPU-accelerated audio capture and transcription without sending data to external servers.
Desktop Shell: Tauri 2.x
Tauri 2.x serves as the native desktop framework, providing the application window, system integration, and the bridge between JavaScript/TypeScript and Rust. This choice replaces heavier alternatives like Electron with a Rust-based runtime that minimizes memory footprint while maintaining full OS-level access.
The Tauri commands are registered in frontend/src-tauri/src/lib.rs:
#[tauri::command]
async fn start_recording<R: Runtime>(
app: AppHandle<R>,
mic_device_name: Option<String>,
system_device_name: Option<String>,
meeting_name: Option<String>,
) -> Result<(), String> {
audio::recording_commands::start(app, mic_device_name, system_device_name, meeting_name)
.await
}
Tauri's event system enables real-time communication between the Rust backend and React frontend through typed payloads.
Frontend: Next.js 14 and React 18
The user interface layer uses Next.js 14 with React 18 and TypeScript, rendered inside the Tauri webview. This provides:
- Server-side rendering capabilities from Next.js for faster initial loads
- Concurrent React features for responsive UI updates during intensive transcription
- Type safety across the API boundary between UI and Rust commands
The React frontend invokes Tauri commands using the @tauri-apps/api/tauri package:
import { invoke } from '@tauri-apps/api/tauri';
async function startMeeting() {
await invoke('start_recording', {
mic_device_name: 'Built-in Microphone',
system_device_name: 'BlackHole 2ch',
meeting_name: 'Team Sync'
});
}
State management handles meeting lists, recording status, and transcript streams received from backend events.
Backend Core: Rust with Async Architecture
Rust powers all performance-critical operations in Meetily. The backend is structured around async Tauri commands that manage:
- Audio capture and real-time mixing
- Voice activity detection (VAD)
- Whisper transcription pipeline
- Local SQLite persistence
- LLM orchestration
The entry point at frontend/src-tauri/src/lib.rs registers all commands and sets up the event emission system for pushing transcript updates to the UI.
Audio Engine: Cross-Platform Capture and Mixing
Meetily's audio pipeline handles microphone and system audio simultaneously through platform-specific backends:
- cpal — Cross-platform audio I/O abstraction
- ScreenCaptureKit — macOS system audio capture
- WASAPI — Windows system audio capture
- ALSA/PulseAudio — Linux audio capture
The core mixing logic resides in frontend/src-tauri/src/audio/pipeline.rs, which implements a ring-buffer pipeline that:
- Captures simultaneous microphone and system audio streams
- Mixes them with professional-grade synchronization
- Applies voice activity detection to filter silent segments
- Feeds speech-only chunks to the transcription engine
This architecture ensures gapless recording even during system load spikes.
Speech-to-Text: Local Whisper with GPU Acceleration
whisper-cpp (exposed through whisper-rs) provides on-device transcription without network dependencies. The integration in frontend/src-tauri/src/whisper_engine/whisper_engine.rs supports:
- CPU inference for universal compatibility
- Metal acceleration on Apple Silicon
- CUDA for NVIDIA GPUs
- Vulkan for cross-platform GPU compute
The engine loads models dynamically and transcribes VAD-filtered chunks:
let transcript = whisper_engine
.transcribe(&audio_chunk)
.await
.map_err(|e| format!("Transcription failed: {}", e))?;
Results are emitted back to the React UI via Tauri's event system:
app.emit_all("transcript-update", TranscriptUpdate {
text: transcript,
timestamp: chrono::Utc::now(),
})?;
LLM Integration: Flexible AI Generation
Meetily integrates multiple large language model providers for meeting summarization and action item extraction:
- Ollama — Fully local, privacy-maximal option
- Claude — Anthropic's API for high-quality summaries
- Groq — Low-latency inference provider
- OpenRouter — Unified access to multiple models
All LLM interactions are user-configurable, defaulting to local Ollama instances to maintain the privacy-first guarantee.
Local Persistence: SQLite with sqlx/rusqlite
SQLite stores all meetings, transcripts, and application configuration via Rust's sqlx and rusqlite crates. This eliminates:
- External database dependencies
- Network sync requirements
- Cloud subscription costs
The schema handles relational data between meetings, transcript segments, and generated summaries while supporting full-text search for historical retrieval.
Dependency Management
Key dependencies are declared in Cargo.toml at the repository root:
[dependencies]
tauri = { version = "2.0", features = [] }
cpal = "0.15"
whisper-rs = { version = "0.8", features = ["cuda", "metal"] }
sqlx = { version = "0.7", features = ["sqlite", "runtime-tokio"] }
serde = { version = "1.0", features = ["derive"] }
tokio = { version = "1.35", features = ["full"] }
TypeScript configuration in frontend/tsconfig.json ensures strict type checking across the Next.js application.
Summary
- Tauri 2.x provides the lightweight native desktop shell bridging Rust and React
- Next.js 14 + React 18 deliver the modern, type-safe user interface
- Rust handles all performance-critical audio, transcription, and persistence operations
- whisper-cpp/whisper-rs enable local, GPU-accelerated speech-to-text
- Cross-platform audio backends (cpal, ScreenCaptureKit, WASAPI, ALSA) capture high-fidelity mixed audio
- SQLite ensures complete data remains on the user's device
- Multiple LLM providers offer flexible AI generation while preserving local-first defaults
Frequently Asked Questions
What makes Meetily different from cloud-based meeting assistants?
Meetily processes all audio and transcription locally on your device. Unlike services that upload recordings to remote servers, Meetily's Rust backend with whisper-cpp runs entirely offline. Your meeting data never leaves your machine, eliminating privacy risks from data breaches or unauthorized access.
Does Meetily require an internet connection?
No. Core functionality—recording, audio mixing, transcription, and local storage—works completely offline. Internet connectivity is only needed if you choose to use remote LLM providers (Claude, Groq, OpenRouter) rather than the default local Ollama integration.
What hardware accelerates Meetily's transcription?
Meetily automatically detects and uses GPU acceleration when available: Apple Silicon via Metal, NVIDIA cards via CUDA, and Vulkan-compatible hardware on other platforms. The whisper-rs engine in frontend/src-tauri/src/whisper_engine/whisper_engine.rs falls back to optimized CPU inference if no GPU is present.
How does the React frontend communicate with Rust backend?
Communication occurs through Tauri's invoke API and event system. The frontend calls Rust functions using invoke() with typed arguments, while the backend pushes real-time updates (transcript segments, recording status) via emit_all() events that React components subscribe to through Tauri's JavaScript event listeners.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →