# Real-Time Transcription Workflow in Meetily: From Audio Chunk to Text Explained

> Discover Meetily's real-time transcription workflow. Learn how it processes audio chunks to text locally using a Rust pipeline, ensuring privacy and speed without network uploads.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: internals
- Published: 2026-08-01

---

**Meetily processes live audio transcription entirely on the client through a Rust-based pipeline that captures audio, filters noise, detects speech, and routes segments to local transcription engines—all without sending audio data over the network.**

This article breaks down the complete real-time transcription workflow in [Zackriya-Solutions/meetily](https://github.com/Zackriya-Solutions/meetily), an open-source meeting assistant built with Tauri and Rust. Understanding this architecture helps developers implement low-latency speech-to-text pipelines for desktop applications.

---

## Stage 1: Audio Capture and Pre-Processing

The pipeline begins in `AudioCapture::process_audio_data` within [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs). This stage transforms raw microphone or system audio into a clean, normalized format suitable for downstream processing.

**Key transformations applied to every audio chunk:**

- **Mono conversion** — collapses stereo streams to single-channel
- **Resampling** — uses a persistent resampler to hit 48 kHz when needed
- **High-pass filtering** — removes low-frequency rumble
- **RNNoise integration** — applies deep learning-based noise suppression
- **EBU R128 loudness normalization** — ensures consistent levels

The processed data is wrapped in an `AudioChunk` struct containing sample data, sample rate, timestamp, chunk ID, and device type. This chunk is then handed to the shared `RecordingState`.

```rust
// Inside AudioCapture::process_audio_data
if self.state.is_recording() {
    // ... mono conversion, optional resampling, filtering ...

    let chunk = AudioChunk {
        data: mono_data,
        sample_rate: if self.needs_resampling { 48_000 } else { self.sample_rate },
        timestamp: self.state.get_recording_duration().unwrap_or(0.0),
        chunk_id: self.chunk_counter.fetch_add(1, Ordering::SeqCst),
        device_type: self.device_type.clone(),
    };

    // Send to pipeline
    self.state.send_audio_chunk(chunk)?;
}

```

---

## Stage 2: Routing Through RecordingState

The `RecordingState` struct in [`frontend/src-tauri/src/audio/recording_state.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/recording_state.rs) serves as a thread-safe hub for audio chunks. Its `send_audio_chunk` method pushes chunks onto an `UnboundedSender` created by `AudioPipelineManager`.

This design decouples capture threads from processing threads, allowing the microphone capture to run independently of the heavier VAD and transcription workloads.

---

## Stage 3: Mixing, VAD, and Speech Segmentation

The `AudioPipeline` receives chunks through its `receiver` channel and performs three critical operations:

### Audio Mixing from Multiple Sources

When both microphone and system audio are captured, the `AudioMixerRingBuffer` (`add_samples` method, lines 49-73) stores streams in overlapping ring buffers. The `ProfessionalAudioMixer::mix_window` method (lines 54-86) blends these into a single 48 kHz stream with proper latency compensation.

### Voice Activity Detection

The mixed audio feeds into `ContinuousVadProcessor` from [`frontend/src-tauri/src/audio/vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/vad.rs). This extracts speech-only segments and emits them as new `AudioChunk` instances at 16 kHz—the optimal sample rate for most transcription engines.

```rust
while self.ring_buffer.can_mix() {
    if let Some((mic_win, sys_win)) = self.ring_buffer.extract_window() {
        let mixed = self.mixer.mix_window(&mic_win, &sys_win);
        // VAD produces speech segments at 16 kHz
        if let Ok(segments) = self.vad_processor.process_audio(&mixed) {
            for seg in segments {
                let vt_chunk = AudioChunk {
                    data: seg.samples,
                    sample_rate: 16_000,
                    timestamp: seg.start_timestamp_ms / 1000.0,
                    chunk_id: self.chunk_id_counter,
                    device_type: DeviceType::Microphone,
                };
                self.transcription_sender.send(vt_chunk)?;
                self.chunk_id_counter += 1;
            }
        }
    }
}

```

---

## Stage 4: Transcription Provider Invocation

VAD-produced chunks travel through a dedicated `transcription_sender` channel to worker tasks. The `transcribe_chunk_with_provider` function in [`frontend/src-tauri/src/audio/transcription/worker.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/transcription/worker.rs) (lines 6-48) handles provider dispatch.

**Three engine types are supported:**

- **WhisperEngine** — OpenAI's Whisper model, local execution
- **ParakeetEngine** — [NVIDIA's Riva Parakeet CTC model](https://github.com/nvidia/riva-asr-api-spec)
- **Custom providers** — any type implementing the `TranscriptionProvider` trait

Each provider receives resampled 16 kHz mono audio and returns a `TranscriptResult` containing recognized text, optional confidence scores, and a partial-flag for streaming display.

```rust
let (text, confidence, partial) = match engine {
    TranscriptionEngine::Whisper(w) => {
        w.transcribe_audio_with_confidence(samples, language).await?
    }
    TranscriptionEngine::Parakeet(p) => {
        (p.transcribe_audio(samples).await?, None, false)
    }
    TranscriptionEngine::Provider(p) => {
        let res = p.transcribe(samples, language).await?;
        (res.text, res.confidence, res.is_partial)
    }
};

```

The `TranscriptionProvider` trait in [`frontend/src-tauri/src/audio/transcription/provider.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/transcription/provider.rs) enables seamless swapping between engines without modifying the worker logic.

---

## Stage 5: Frontend Consumption via Tauri Events

Transcription results flow back to the React/Next.js UI through Tauri's event system. The frontend listens for `transcript-update` events and updates the meeting display in real time.

```typescript
await listen<TranscriptUpdate>('transcript-update', (event) => {
  setTranscripts(prev => [...prev, event.payload.text]);
});

```

The UI entry point in [`frontend/src/app/page.tsx`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src/app/page.tsx) wires these listeners to the transcript state, enabling live captions during meetings.

---

## Architecture Benefits of Meetily's Real-Time Transcription

**Privacy-first design** — All audio processing happens locally in the Tauri process. No audio data leaves the machine unless explicitly configured for external LLM calls.

**Deterministic resource management** — Explicit flush signals and channel-based communication ensure clean shutdowns without zombie threads.

**Modular provider system** — The trait-based `TranscriptionProvider` interface allows community contributions of new transcription engines without pipeline changes.

**Sub-millisecond latency** — By avoiding network roundtrips and using lock-free ring buffers, the pipeline achieves processing delays suitable for live captioning.

---

## Summary

- **Audio capture** in `AudioCapture::process_audio_data` normalizes raw PCM to 48 kHz mono chunks
- **RecordingState** routes chunks through channels to decouple capture from processing
- **AudioPipeline** mixes multi-source audio and runs `ContinuousVadProcessor` to extract 16 kHz speech segments
- **Transcription workers** dispatch segments to Whisper, Parakeet, or custom providers via `transcribe_chunk_with_provider`
- **Tauri events** stream results to the React frontend for real-time display

The complete implementation spans [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs), [`recording_state.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/recording_state.rs), [`vad.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/vad.rs), [`transcription/worker.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/transcription/worker.rs), and [`transcription/provider.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/transcription/provider.rs).

---

## Frequently Asked Questions

### How does Meetily achieve low latency in transcription?

Meetily eliminates network latency by running all transcription models locally through ONNX Runtime or similar inference engines. The Rust-based pipeline uses lock-free ring buffers and dedicated worker threads to keep audio capture, VAD, and transcription in parallel. According to the source code, chunks flow through channels with minimal copying, and VAD segments at 16 kHz match the native sample rate of most speech models, avoiding runtime resampling.

### Can I use a custom transcription model with Meetily?

Yes. Implement the `TranscriptionProvider` trait defined in [`frontend/src-tauri/src/audio/transcription/provider.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/transcription/provider.rs). Your implementation must provide a `transcribe` method that takes audio samples and a language code, returning a `TranscriptResult`. The `TranscriptionEngine::Provider` variant in the worker will automatically route chunks to your implementation. No changes to the audio pipeline or frontend are required.

### Why does Meetily resample to 48 kHz during capture then down to 16 kHz for transcription?

The 48 kHz intermediate format provides headroom for professional audio mixing and aligns with standard system audio interfaces. The `ContinuousVadProcessor` specifically targets 16 kHz because this is the training sample rate for Whisper, Parakeet, and most other speech recognition models. Downsampling after VAD reduces computational load on the transcription engine while preserving the quality needed for recognition accuracy.

### Where does noise suppression happen in the pipeline?

RNNoise filtering and high-pass filtering occur in `AudioCapture::process_audio_data` before chunks enter the shared state. This placement ensures all downstream components—mixing, VAD, and transcription—receive cleaned audio. The EBU R128 loudness normalization also happens here, preventing clipping and maintaining consistent levels across different microphone hardware.