How to Implement Speaker Diarization in Meetily's Transcription Pipeline
Speaker diarization in Meetily is implemented by clustering the 128-dimensional speaker embeddings already generated during transcription, assigning cluster labels to each TranscriptSegment via a new speaker_id field, and emitting those labels to the React frontend for display.
Meetily's transcription pipeline already extracts speaker embeddings from every speech chunk, but it does not yet cluster them into distinct speaker identities. This article walks through the complete implementation—extending the core data structures, creating a Rust-based clustering module, and wiring it into the existing worker flow—so you can add speaker labels to your meeting transcripts without disrupting real-time performance.
Meetily's Existing Transcription Architecture
Before adding diarization, you need to understand how Meetily processes audio. The pipeline in frontend/src-tauri/src/audio/pipeline.rs captures mixed microphone and system audio, applies Voice Activity Detection (VAD) to extract speech chunks, and pushes those chunks through a ring buffer to transcription workers.
The transcription worker at audio/transcription/worker.rs delegates each chunk to a provider—either whisper_provider.rs or parakeet_provider.rs. These providers return an SttResult containing the transcribed text and a pre-computed speaker embedding. This embedding is a 128-dimensional vector that captures the acoustic characteristics of the speaker's voice.
Why this matters: Because embeddings are already extracted, speaker diarization becomes a post-processing clustering problem rather than a model training problem. You do not need to modify the VAD or transcription stages.
Step 1: Extend the Transcript Data Structure
The first code change is in frontend/src-tauri/src/audio/stt.rs. The existing TranscriptSegment struct already stores the embedding; you only need to add a field for the assigned speaker label.
pub struct TranscriptSegment {
pub text: String,
pub confidence: f32,
pub speaker_embedding: Vec<f32>, // existing: 128-dim vector from Whisper/Parakeet
pub speaker_id: Option<String>, // NEW: cluster-assigned label like "Speaker 1"
}
This minimal change preserves backward compatibility—speaker_id is Option<String> so existing code continues to compile—while providing the hook for your diarization logic to attach labels.
Step 2: Create the Diarization Module
Create a new file at frontend/src-tauri/src/audio/diarization.rs. This module implements clustering using the linfa machine learning crate, which provides a well-tested K-means implementation in pure Rust.
use linfa::prelude::*;
use linfa_clustering::KMeans;
use ndarray::Array2;
/// Assigns speaker labels to transcript segments by clustering embeddings.
///
/// # Arguments
/// * `segments` - Mutable slice of transcript segments with populated embeddings
/// * `max_speakers` - Upper bound for distinct speakers (e.g., 4 for a small meeting)
pub fn diarize(segments: &mut [TranscriptSegment], max_speakers: usize) -> anyhow::Result<()> {
// Guard against empty input
let embed_dim = segments
.first()
.ok_or_else(|| anyhow::anyhow!("no segments to diarize"))?
.speaker_embedding
.len();
// Flatten embeddings into an n×d matrix for linfa
let flattened: Vec<f32> = segments
.iter()
.flat_map(|s| s.speaker_embedding.clone())
.collect();
let data: Array2<f32> = Array2::from_shape_vec(
(segments.len(), embed_dim),
flattened,
)?;
// Run K-means with configurable iteration limit
let model = KMeans::params_with_n_clusters(max_speakers)
.max_n_iterations(100)
.tolerance(1e-4)
.fit(&data)
.map_err(|e| anyhow::anyhow!("clustering failed: {}", e))?;
// Map cluster indices back to segment labels
let predictions = model.predict(&data);
for (segment, &cluster_idx) in segments.iter_mut().zip(predictions.iter()) {
segment.speaker_id = Some(format!("Speaker {}", cluster_idx + 1));
}
Ok(())
}
Add to Cargo.toml:
[dependencies]
linfa = "0.7"
linfa-clustering = "0.7"
ndarray = "0.15"
anyhow = "1.0"
Alternative approaches: If you want lighter dependencies, replace the K-means implementation with cosine similarity thresholding or agglomerative clustering using only ndarray. The interface remains identical—accept a mutable slice, populate speaker_id.
Step 3: Integrate Diarization into the Transcription Worker
In frontend/src-tauri/src/audio/transcription/worker.rs, import the new module and invoke it after accumulating a batch of segments:
use crate::audio::diarization::diarize;
// Inside the worker's processing loop...
let mut batch_segments: Vec<TranscriptSegment> = Vec::new();
// ... populate batch_segments from provider results ...
// Run diarization when flag is enabled (configured via UI)
if state.diarization_enabled && !batch_segments.is_empty() {
if let Err(e) = diarize(&mut batch_segments, state.max_speakers) {
log::error!("Diarization failed: {}", e);
// Continue emitting segments without speaker labels rather than failing hard
}
}
// Emit to frontend; speaker_id is now present when diarization succeeded
for segment in batch_segments {
app.emit("transcript-update", segment.clone())?;
}
Performance consideration: Clustering runs on the CPU with a matrix size of (batch_size, 128). For typical batches of 50-200 segments, this completes in single-digit milliseconds. Run it in tokio::task::spawn_blocking if you want to guarantee no disruption to the async audio pipeline, though in practice the overhead is negligible.
Step 4: Update the Frontend TypeScript Types
The React frontend already listens for transcript-update events. Extend the type definition to include the new field:
interface TranscriptUpdate {
text: string;
confidence: number;
speaker_id?: string; // undefined when diarization is disabled or failed
}
listen<TranscriptUpdate>('transcript-update', (event) => {
const { text, confidence, speaker_id } = event.payload;
const label = speaker_id ?? 'Unknown Speaker';
// Render with distinct color per speaker
addTranscriptLine({
text,
confidence,
speakerLabel: label,
speakerColor: hashColor(label), // deterministic color from label string
});
});
The existing event emission in lib.rs requires no changes—it already serializes and emits the full TranscriptSegment struct.
Step 5: Add User-Facing Configuration
Expose a toggle in the React UI that controls diarization_enabled and max_speakers in the Rust RecordingState. The worker checks these flags before invoking the diarization function, allowing users to disable the feature for privacy or performance reasons.
// In your state management
pub struct RecordingState {
pub diarization_enabled: bool,
pub max_speakers: usize, // default 4, configurable up to 10
}
Key Source Files Reference
| File | Purpose |
|---|---|
frontend/src-tauri/src/audio/stt.rs |
Core transcript data structures |
frontend/src-tauri/src/audio/diarization.rs |
New: Clustering implementation |
frontend/src-tauri/src/audio/transcription/worker.rs |
Orchestrates transcription and diarization |
frontend/src-tauri/src/audio/transcription/whisper_provider.rs |
Generates speaker embeddings |
frontend/src-tauri/src/audio/pipeline.rs |
Audio capture and VAD |
frontend/src-tauri/src/lib.rs |
Event emission to frontend |
Summary
- Meetily already generates speaker embeddings during transcription—no additional ML models needed for embedding extraction.
- Add a
speaker_idfield toTranscriptSegmentinstt.rsto store cluster labels. - Create
audio/diarization.rsimplementing K-means or alternative clustering withlinfa. - Wire into
worker.rsto run clustering after each transcription batch, gated by a user toggle. - Frontend receives labels automatically through the existing
transcript-updateevent channel.
Frequently Asked Questions
Does implementing speaker diarization add latency to live transcription?
No. The diarization step runs after transcription completes on a batch of segments, not on the critical path of audio capture and VAD. For typical meeting sizes (under 10 speakers, batches under 200 segments), clustering completes in under 10 milliseconds. You can optionally offload it to a blocking worker thread with tokio::task::spawn_blocking if absolute audio pipeline isolation is required.
Why use K-means instead of a neural diarization model?
K-means clustering on pre-extracted embeddings is sufficient for Meetily's architecture because Whisper and Parakeet already provide high-quality 128-dimensional speaker vectors. Neural diarization models (like pyannote.audio) excel when you must directly process raw waveforms, but they add heavy Python dependencies and GPU requirements. The Rust-native linfa approach keeps Meetily's binary portable and offline-capable.
Can I adjust the number of speakers after a meeting starts?
The max_speakers parameter sets an upper bound, not an exact count. K-means will use fewer clusters if the data supports it. For dynamic adjustment, you could implement online clustering (e.g., incremental DBSCAN) that creates new speaker labels when encountering embeddings dissimilar to existing clusters. The modular structure of diarization.rs makes this swap straightforward—replace the function body while keeping the same interface.
What if two speakers have similar embeddings?
Speaker embedding similarity is a fundamental limitation of all diarization systems. In Meetily, this manifests as occasional cluster merging. Mitigate by:
- Increasing embedding quality via larger Whisper models (if supported by your provider)
- Post-processing with duration constraints (short segments inherit neighbor labels)
- Using agglomerative clustering with a strict cophenetic distance threshold instead of K-means
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →