What Are the Main Modules and Services in Meetily? A Technical Architecture Guide
Meetily is built as a hybrid Rust-and-Next.js application where a Tauri-based backend manages audio capture, transcription, and AI summarization through specialized modules, while a React frontend handles the user interface.
Meetily is a privacy-first AI meeting assistant developed by Zackriya-Solutions that runs as a cross-platform desktop application. Understanding the main modules and services in Meetily requires examining its dual-layer architecture: a high-performance Rust core that processes audio streams and executes machine learning workloads, and a Next.js frontend that provides the user interface. The codebase is organized into distinct domains spanning audio processing, speech recognition, natural language generation, and state persistence.
Core Architectural Layers
The application's backend resides in frontend/src-tauri/src/ and is organized into functional layers that communicate through Tauri's command pattern and event system.
| Layer | Responsibility | Key Source Files |
|---|---|---|
| Tauri Command Bus | Exposes Rust functions to the UI and manages application lifecycle | frontend/src-tauri/src/lib.rs |
| Audio Pipeline | Captures microphone and system audio, applies mixing and voice detection | frontend/src-tauri/src/audio/pipeline.rs, vad.rs |
| Transcription | Converts speech to text using Whisper or Parakeet | frontend/src-tauri/src/audio/transcription/whisper_provider.rs |
| Summarization | Processes transcripts through LLMs to generate meeting summaries | frontend/src-tauri/src/summary/processor.rs |
| Provider Integrations | Connects to external AI services (OpenAI, Groq, Ollama) | frontend/src-tauri/src/ollama/ollama.rs, openai/openai.rs |
| Persistence | Stores audio files and metadata locally | frontend/src-tauri/src/database/mod.rs, recording_saver.rs |
Audio Capture and Processing Pipeline
Device Detection and Permission Management
The audio subsystem begins with platform-specific device enumeration handled in frontend/src-tauri/src/audio/device_detection.rs. This module detects available microphones and system audio endpoints, while audio/permissions.rs manages OS-level permission requests required for screen audio capture.
Recording Orchestration
The frontend/src-tauri/src/audio/recording_manager.rs file serves as the primary controller for recording sessions. It instantiates the RecordingState struct defined in recording_state.rs and coordinates the lifecycle of audio streams. When a user initiates a capture, the manager creates platform-specific stream handlers and delegates processing to the pipeline module.
Audio Mixing and Voice Activity Detection
At the heart of Meetily's audio processing lies frontend/src-tauri/src/audio/pipeline.rs. This module synchronizes incoming microphone and system audio streams, applies professional RMS-based ducking to balance levels, and performs voice-activity detection (VAD) using the logic in frontend/src-tauri/src/audio/vad.rs. Validated speech chunks are forwarded to the transcription queue, while silent segments are discarded to optimize processing.
Speech-to-Text Transcription Services
Provider Architecture
Meetily abstracts transcription behind the AudioTranscriptionProvider trait, enabling pluggable backends. The primary implementation resides in frontend/src-tauri/src/audio/transcription/whisper_provider.rs, which wraps Whisper-cpp for local inference. An alternative provider in audio/transcription/parakeet_provider.rs supports NVIDIA's Parakeet model for GPU-accelerated transcription.
Both providers expose a consistent async interface:
async fn transcribe(chunk: AudioChunk) -> String
Model Loading and Hardware Acceleration
The frontend/src-tauri/src/whisper_engine/mod.rs module handles model initialization and hardware detection. At startup, it probes for Metal (macOS), CUDA (NVIDIA), or Vulkan support and selects the optimal computation backend for the loaded Whisper model, ensuring real-time transcription performance.
AI Summarization Engine
Summary Processing Pipeline
Once a recording concludes, the summarization workflow activates through frontend/src-tauri/src/summary/commands.rs, which delegates to frontend/src-tauri/src/summary/processor.rs. This processor handles transcript chunking, automatic language detection via summary/language_detection.rs, and template-based formatting using Mustache-style templates stored in summary/templates/.
LLM Provider Integrations
The summarization service supports multiple LLM backends through dedicated client modules:
- Ollama: Local model inference via
frontend/src-tauri/src/ollama/ollama.rs - OpenAI: Cloud API integration in
frontend/src-tauri/src/openai/openai.rs - Groq: High-performance inference via
frontend/src-tauri/src/groq/groq.rs - OpenRouter/Claude: Additional providers in
openrouter/openrouter.rs
Each module implements provider-specific authentication, request formatting, and response parsing, allowing users to select their preferred AI backend through the generate_summary command.
State Management and Data Persistence
Global application state is maintained in frontend/src-tauri/src/state.rs, which keeps the Rust backend synchronized with the React frontend's expectations. For long-term storage, frontend/src-tauri/src/database/mod.rs manages a local SQLite database containing meeting metadata, transcript file paths, and timestamps. Raw audio recordings are persisted as WAV files through frontend/src-tauri/src/audio/recording_saver.rs, which organizes files by meeting ID in the application's data directory.
System Integration Layer
Tauri Command Registry
The frontend/src-tauri/src/lib.rs file functions as the API gateway, registering all exposed commands including start_recording, stop_recording, and generate_summary. It also initializes the event emitter used to push live transcript updates to the frontend via the transcript-update event channel.
System Notifications
User feedback occurs through frontend/src-tauri/src/notifications/manager.rs for desktop alerts and frontend/src-tauri/src/tray.rs for system tray icon management, ensuring users receive status updates even when the application window is minimized.
Next.js Frontend Interface
The user interface resides in frontend/src/ as a Next.js application. The entry point at frontend/src/app/page.tsx provides controls for initiating recordings, displays live transcripts streamed from the Rust backend, and renders generated summaries. Global UI state management, including sidebar navigation state, is handled by frontend/src/components/Sidebar/SidebarProvider.tsx.
Practical Integration Examples
Initiating a Recording Session
To start capturing audio from the frontend, invoke the registered Tauri command with device specifications:
import { invoke } from '@tauri-apps/api/tauri';
await invoke('start_recording', {
mic_device_name: 'Built-in Microphone',
system_device_name: 'BlackHole 2ch',
meeting_name: 'Team Sync 2024-07-29'
});
This command is processed by the command handler in frontend/src-tauri/src/lib.rs and passed to recording_manager.rs to begin the audio pipeline.
Consuming Live Transcript Updates
The frontend listens for real-time transcription results via Tauri's event system:
import { listen } from '@tauri-apps/api/event';
listen('transcript-update', event => {
console.log('New transcript segment:', event.payload);
});
Events are emitted from frontend/src-tauri/src/audio/pipeline.rs whenever the VAD-filtered audio chunks complete processing through the transcription provider.
Generating Meeting Summaries
After concluding a recording, trigger the summarization workflow:
import { invoke } from '@tauri-apps/api/tauri';
const summary = await invoke<string>('generate_summary', {
meeting_id: currentMeetingId,
model: 'ollama:llama2' // or 'openai:gpt-4', 'groq:mixtral'
});
This invokes the logic chain in summary/commands.rs → summary/processor.rs → specific LLM client (e.g., ollama/ollama.rs).
Accessing Recorded Audio Files
Retrieve the local path to saved recordings:
import { path } from '@tauri-apps/api';
const audioPath = await path.appDataDir();
const wavFile = `${audioPath}/meetings/${meetingId}.wav`;
The recording_saver.rs module handles the actual file system operations, ensuring audio data persists between application sessions.
Summary
- Meetily combines a Rust-based Tauri backend with a Next.js frontend to deliver a privacy-focused AI meeting assistant.
- The audio pipeline (
pipeline.rs,vad.rs) handles complex tasks including audio mixing, RMS ducking, and voice activity detection before transcription. - Transcription providers (
whisper_provider.rs,parakeet_provider.rs) implement a common trait for interchangeable speech-to-text engines with GPU acceleration support. - The summarization engine (
summary/processor.rs) integrates with multiple LLM backends including Ollama, OpenAI, and Groq for flexible meeting analysis. - State and persistence modules ensure meeting metadata resides in SQLite while audio files are stored locally as WAV.
- Tauri commands (
lib.rs) and events bridge the Rust core and React UI, enabling real-time transcript streaming and system notifications.
Frequently Asked Questions
How does Meetily handle audio capture from both microphone and system audio simultaneously?
Meetily's frontend/src-tauri/src/audio/pipeline.rs creates distinct input streams for the microphone and system audio (using platform-specific virtual audio drivers on macOS like BlackHole). It synchronizes these streams, applies RMS-based ducking to prevent audio clipping, and mixes them into a single buffer before processing. The recording_manager.rs orchestrates this while vad.rs filters out silent segments to reduce transcription overhead.
What transcription engines does Meetily support and how are they configured?
Meetily supports both Whisper-cpp via frontend/src-tauri/src/audio/transcription/whisper_provider.rs and NVIDIA Parakeet via parakeet_provider.rs. Both implement the AudioTranscriptionProvider trait with an async transcribe method. The whisper_engine/mod.rs module automatically detects available hardware acceleration (Metal, CUDA, or Vulkan) and loads the appropriate model variant. Selection between providers occurs at runtime based on configuration passed through the Tauri command layer.
How does the summarization module process meeting transcripts into structured summaries?
The frontend/src-tauri/src/summary/processor.rs handles the summarization workflow by first chunking the complete transcript and detecting the language via language_detection.rs. It then selects the appropriate LLM client based on the user's provider preference (Ollama, OpenAI, Groq, or OpenRouter), sends the chunked text with a Mustache-style template, and aggregates the response into a formatted meeting summary. The summary/commands.rs module exposes this functionality to the frontend through the generate_summary command.
Where does Meetily store recording data and how is it organized?
Audio recordings are saved as WAV files in the application's data directory (e.g., ~/Library/Application Support/meetily/meetings/ on macOS) via frontend/src-tauri/src/audio/recording_saver.rs. Metadata including meeting titles, timestamps, and file paths are stored in a local SQLite database managed by frontend/src-tauri/src/database/mod.rs. This hybrid approach ensures that sensitive audio data remains local while maintaining a queryable index of meeting history accessible through the state.rs synchronization layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →