# How to Implement Speaker Diarization in Meetily's Transcription Pipeline

> Learn how to implement speaker diarization in Meetily's transcription pipeline. Cluster speaker embeddings, assign IDs, and display labels on the frontend for enhanced audio analysis.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: how-to-guide
- Published: 2026-08-04

---

**Speaker diarization in Meetily is implemented by clustering the 128-dimensional speaker embeddings already generated during transcription, assigning cluster labels to each `TranscriptSegment` via a new `speaker_id` field, and emitting those labels to the React frontend for display.**

Meetily's transcription pipeline already extracts speaker embeddings from every speech chunk, but it does not yet cluster them into distinct speaker identities. This article walks through the complete implementation—extending the core data structures, creating a Rust-based clustering module, and wiring it into the existing worker flow—so you can add speaker labels to your meeting transcripts without disrupting real-time performance.

## Meetily's Existing Transcription Architecture

Before adding diarization, you need to understand how Meetily processes audio. The pipeline in [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs) captures mixed microphone and system audio, applies **Voice Activity Detection (VAD)** to extract speech chunks, and pushes those chunks through a ring buffer to transcription workers.

The transcription worker at [`audio/transcription/worker.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/audio/transcription/worker.rs) delegates each chunk to a provider—either [`whisper_provider.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/whisper_provider.rs) or [`parakeet_provider.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/parakeet_provider.rs). These providers return an `SttResult` containing the transcribed text and a pre-computed **speaker embedding**. This embedding is a 128-dimensional vector that captures the acoustic characteristics of the speaker's voice.

**Why this matters:** Because embeddings are already extracted, speaker diarization becomes a post-processing clustering problem rather than a model training problem. You do not need to modify the VAD or transcription stages.

## Step 1: Extend the Transcript Data Structure

The first code change is in [`frontend/src-tauri/src/audio/stt.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/stt.rs). The existing `TranscriptSegment` struct already stores the embedding; you only need to add a field for the assigned speaker label.

```rust
pub struct TranscriptSegment {
    pub text: String,
    pub confidence: f32,
    pub speaker_embedding: Vec<f32>,   // existing: 128-dim vector from Whisper/Parakeet
    pub speaker_id: Option<String>,   // NEW: cluster-assigned label like "Speaker 1"
}

```

This minimal change preserves backward compatibility—`speaker_id` is `Option<String>` so existing code continues to compile—while providing the hook for your diarization logic to attach labels.

## Step 2: Create the Diarization Module

Create a new file at [`frontend/src-tauri/src/audio/diarization.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/diarization.rs). This module implements clustering using the `linfa` machine learning crate, which provides a well-tested K-means implementation in pure Rust.

```rust
use linfa::prelude::*;
use linfa_clustering::KMeans;
use ndarray::Array2;

/// Assigns speaker labels to transcript segments by clustering embeddings.
/// 
/// # Arguments

/// * `segments` - Mutable slice of transcript segments with populated embeddings
/// * `max_speakers` - Upper bound for distinct speakers (e.g., 4 for a small meeting)
pub fn diarize(segments: &mut [TranscriptSegment], max_speakers: usize) -> anyhow::Result<()> {
    // Guard against empty input
    let embed_dim = segments
        .first()
        .ok_or_else(|| anyhow::anyhow!("no segments to diarize"))?
        .speaker_embedding
        .len();

    // Flatten embeddings into an n×d matrix for linfa
    let flattened: Vec<f32> = segments
        .iter()
        .flat_map(|s| s.speaker_embedding.clone())
        .collect();

    let data: Array2<f32> = Array2::from_shape_vec(
        (segments.len(), embed_dim),
        flattened,
    )?;

    // Run K-means with configurable iteration limit
    let model = KMeans::params_with_n_clusters(max_speakers)
        .max_n_iterations(100)
        .tolerance(1e-4)
        .fit(&data)
        .map_err(|e| anyhow::anyhow!("clustering failed: {}", e))?;

    // Map cluster indices back to segment labels
    let predictions = model.predict(&data);
    for (segment, &cluster_idx) in segments.iter_mut().zip(predictions.iter()) {
        segment.speaker_id = Some(format!("Speaker {}", cluster_idx + 1));
    }

    Ok(())
}

```

Add to [`Cargo.toml`](https://github.com/Zackriya-Solutions/meetily/blob/main/Cargo.toml):

```toml
[dependencies]
linfa = "0.7"
linfa-clustering = "0.7"
ndarray = "0.15"
anyhow = "1.0"

```

**Alternative approaches:** If you want lighter dependencies, replace the K-means implementation with **cosine similarity thresholding** or **agglomerative clustering** using only `ndarray`. The interface remains identical—accept a mutable slice, populate `speaker_id`.

## Step 3: Integrate Diarization into the Transcription Worker

In [`frontend/src-tauri/src/audio/transcription/worker.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/transcription/worker.rs), import the new module and invoke it after accumulating a batch of segments:

```rust
use crate::audio::diarization::diarize;

// Inside the worker's processing loop...
let mut batch_segments: Vec<TranscriptSegment> = Vec::new();

// ... populate batch_segments from provider results ...

// Run diarization when flag is enabled (configured via UI)
if state.diarization_enabled && !batch_segments.is_empty() {
    if let Err(e) = diarize(&mut batch_segments, state.max_speakers) {
        log::error!("Diarization failed: {}", e);
        // Continue emitting segments without speaker labels rather than failing hard
    }
}

// Emit to frontend; speaker_id is now present when diarization succeeded
for segment in batch_segments {
    app.emit("transcript-update", segment.clone())?;
}

```

**Performance consideration:** Clustering runs on the CPU with a matrix size of `(batch_size, 128)`. For typical batches of 50-200 segments, this completes in single-digit milliseconds. Run it in `tokio::task::spawn_blocking` if you want to guarantee no disruption to the async audio pipeline, though in practice the overhead is negligible.

## Step 4: Update the Frontend TypeScript Types

The React frontend already listens for `transcript-update` events. Extend the type definition to include the new field:

```typescript
interface TranscriptUpdate {
  text: string;
  confidence: number;
  speaker_id?: string;  // undefined when diarization is disabled or failed
}

listen<TranscriptUpdate>('transcript-update', (event) => {
  const { text, confidence, speaker_id } = event.payload;
  const label = speaker_id ?? 'Unknown Speaker';
  
  // Render with distinct color per speaker
  addTranscriptLine({
    text,
    confidence,
    speakerLabel: label,
    speakerColor: hashColor(label),  // deterministic color from label string
  });
});

```

The existing event emission in [`lib.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/lib.rs) requires no changes—it already serializes and emits the full `TranscriptSegment` struct.

## Step 5: Add User-Facing Configuration

Expose a toggle in the React UI that controls `diarization_enabled` and `max_speakers` in the Rust `RecordingState`. The worker checks these flags before invoking the diarization function, allowing users to disable the feature for privacy or performance reasons.

```rust
// In your state management
pub struct RecordingState {
    pub diarization_enabled: bool,
    pub max_speakers: usize,  // default 4, configurable up to 10
}

```

## Key Source Files Reference

| File | Purpose |
|------|---------|
| [`frontend/src-tauri/src/audio/stt.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/stt.rs) | Core transcript data structures |
| [`frontend/src-tauri/src/audio/diarization.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/diarization.rs) | **New:** Clustering implementation |
| [`frontend/src-tauri/src/audio/transcription/worker.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/transcription/worker.rs) | Orchestrates transcription and diarization |
| [`frontend/src-tauri/src/audio/transcription/whisper_provider.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/transcription/whisper_provider.rs) | Generates speaker embeddings |
| [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs) | Audio capture and VAD |
| [`frontend/src-tauri/src/lib.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/lib.rs) | Event emission to frontend |

## Summary

- **Meetily already generates speaker embeddings** during transcription—no additional ML models needed for embedding extraction.
- **Add a `speaker_id` field** to `TranscriptSegment` in [`stt.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/stt.rs) to store cluster labels.
- **Create [`audio/diarization.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/audio/diarization.rs)** implementing K-means or alternative clustering with `linfa`.
- **Wire into [`worker.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/worker.rs)** to run clustering after each transcription batch, gated by a user toggle.
- **Frontend receives labels automatically** through the existing `transcript-update` event channel.

## Frequently Asked Questions

### Does implementing speaker diarization add latency to live transcription?

No. The diarization step runs **after** transcription completes on a batch of segments, not on the critical path of audio capture and VAD. For typical meeting sizes (under 10 speakers, batches under 200 segments), clustering completes in under 10 milliseconds. You can optionally offload it to a blocking worker thread with `tokio::task::spawn_blocking` if absolute audio pipeline isolation is required.

### Why use K-means instead of a neural diarization model?

K-means clustering on pre-extracted embeddings is **sufficient for Meetily's architecture** because Whisper and Parakeet already provide high-quality 128-dimensional speaker vectors. Neural diarization models (like pyannote.audio) excel when you must directly process raw waveforms, but they add heavy Python dependencies and GPU requirements. The Rust-native `linfa` approach keeps Meetily's binary portable and offline-capable.

### Can I adjust the number of speakers after a meeting starts?

The `max_speakers` parameter sets an upper bound, not an exact count. K-means will use fewer clusters if the data supports it. For dynamic adjustment, you could implement **online clustering** (e.g., incremental DBSCAN) that creates new speaker labels when encountering embeddings dissimilar to existing clusters. The modular structure of [`diarization.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/diarization.rs) makes this swap straightforward—replace the function body while keeping the same interface.

### What if two speakers have similar embeddings?

Speaker embedding similarity is a fundamental limitation of all diarization systems. In Meetily, this manifests as occasional cluster merging. Mitigate by:
- Increasing embedding quality via larger Whisper models (if supported by your provider)
- Post-processing with duration constraints (short segments inherit neighbor labels)
- Using **agglomerative clustering** with a strict cophenetic distance threshold instead of K-means