# Implementing Voice Activity Detection and Speaker Diarization in Palmier Pro

> Learn to implement voice activity detection and speaker diarization in Palmier Pro. Discover how Palmier Pro processes audio with SpeechVAD and MLX for accurate results.

- Repository: [Palmier/palmier-pro](https://github.com/palmier-io/palmier-pro)
- Tags: how-to-guide
- Published: 2026-07-27

---

**Palmier Pro performs both voice activity detection (VAD) and speaker diarization using the open-source SpeechVAD package and a custom MLX-powered analysis pipeline that processes audio in 32 ms chunks while enforcing strict concurrency limits.**

The `palmier-io/palmier-pro` repository implements a production-ready audio analysis stack that combines on-device machine learning with robust error handling. This article examines the Swift implementation responsible for identifying speech segments and assigning speaker identities, referencing the actual source files and concurrency architecture used in the application.

## Core VAD Implementation in VoiceActivity.swift

The primary entry point for speech detection is `VoiceActivity.analysis(for:mediaRef:)` in [`Sources/PalmierPro/Audio/Analysis/VoiceActivity.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Audio/Analysis/VoiceActivity.swift). This method orchestrates a multi-stage pipeline that transforms raw audio files into structured speech spans.

**Audio Decoding with AudioTrackReader**

The process begins with `AudioTrackReader.readMonoFloats(from:sampleRate:)`, which loads the source file off the main actor into a mono float buffer sampled at 16 kHz. This decouples the heavy I/O operation from the UI thread and standardizes the input format for the neural network.

**MLX-Powered Silero Model**

The VAD engine uses a pre-trained Silero model accelerated via MLX. Because the underlying `SileroVADModel` is **not thread-safe**, the code wraps model access in a `ModelBox` actor. This actor serializes all inference calls, preventing race conditions when multiple audio files are processed concurrently.

The model loads lazily via `SileroVADModel.fromPretrained(engine: .mlx)` and processes audio in fixed 512-sample chunks—approximately 32 ms of audio per frame.

**Chunk Processing and Probability Extraction**

The `detectSpeech(vad:samples:)` method iterates through the float buffer using `stride(from:to:by:)` and feeds each chunk to `SileroVADModel.processChunk`. The method respects task cancellation via `Task.checkCancellation()` and aggregates per-frame speech probabilities into an array for downstream processing.

## The VAD Pipeline and Configuration

Raw probabilities require transformation into boolean speech masks. The `VADPipeline` class, configured via `VADConfig`, handles this binarization.

The configuration dynamically sets `windowDuration` based on the number of probability frames multiplied by the chunk duration (`Float(probabilities.count) * chunkDuration`). Calling `pipeline.binarize(probs:)` produces an array of `VoiceActivity.Span` structs, each containing start and end times for contiguous speech segments.

This abstraction separates the probabilistic model output from the deterministic logic that merges adjacent frames into usable time ranges.

## Speaker Diarization with SpeakerIdentity.swift

Building upon the VAD output, [`Sources/PalmierPro/Audio/Analysis/SpeakerIdentity.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Audio/Analysis/SpeakerIdentity.swift) implements speaker diarization by importing the `SpeechVAD` package (available when `BUNDLED_SPEECH` is enabled during compilation).

The module follows the same architectural pattern as the VAD implementation: it reads the audio asset, runs the speech detection model to generate masks, and then feeds these masks into a clustering engine. This engine analyzes voice characteristics to group speech segments by speaker identity.

The resulting tuples—containing `speakerID`, `start`, and `end` times—are consumed by `Sources/PalmierPro/Editor/ViewModel/EditorViewModel+Speakers.swift` to render timeline tints, speaker labels, and silence-wash visualizations in the user interface.

## Concurrency and Memory Management

Palmier Pro enforces strict resource limits to prevent unbounded memory usage during batch processing.

**ModelBox Actor for Thread Safety**

The `ModelBox` actor guarantees exclusive access to the Silero model, satisfying the library's requirement that inference calls must not overlap. Any concurrent requests to analyze multiple files are automatically serialized at the model access layer.

**AsyncSemaphore for Pipeline Gating**

An `AsyncSemaphore` named `pipelineGate` caps concurrent decode and inference operations to **two files maximum**. The semaphore is awaited before decoding begins and released after inference completes. This bounds memory consumption when processing large projects with numerous audio tracks.

## Caching Strategy and Edge Cases

The implementation includes defensive programming for production reliability.

**Disk Cache with Modification-Time Tagging**

Analysis results are persisted as JSON side-car files using `DiskCache`. The cache key combines the media reference with a `sizeMtimeTag` generated from the source file’s modification time. The method `cachedAnalysis(for:mediaRef:)` checks for existing results before re-running the expensive ML pipeline.

**Damaged Media and No-Audio Tracks**

- **No-audio tracks**: When a media asset lacks an audio track, the system caches a zero-length analysis via `cacheNoAudioAnalysis` to avoid redundant processing attempts.
- **Damaged files**: Errors from AVFoundation (specifically domain `AVFoundationErrorDomain`, code `-11829`) are recognized via the `isDamagedMedia` check and reported as failures without crashing the application.
- **Cancellation**: Both the chunking loop and the inference engine respect cooperative cancellation through `Task.checkCancellation()` and `MLXRuntime.shouldStop`.

## Practical Code Examples

The following patterns demonstrate how to integrate VAD and diarization into your own Palmier Pro workflows.

**Running Voice Activity Detection**

```swift
import PalmierPro

// Analyze an audio asset for speech segments
let mediaRef = "track_01"
let analysis = try await VoiceActivity.analysis(for: audioURL, mediaRef: mediaRef)

// Access the boolean mask (one value per 32ms chunk)
let mask = analysis.mask
let cellDuration = VoiceActivity.chunkDuration

// Map mask indices to timeline positions
for (index, isSpeech) in mask.enumerated() {
    let startTime = Double(index) * cellDuration
    if isSpeech {
        print("Speech detected at \(startTime)s")
    }
}

```

**Performing Speaker Diarization**

```swift
import SpeechVAD

// Cluster speech segments by speaker identity
let speakerSpans = try await SpeakerIdentity.analyze(
    sourceURL: audioURL,
    mediaRef: mediaRef
)

// Apply labels to timeline UI
for span in speakerSpans {
    timelineView.addSpeakerLabel(
        id: span.speakerID,
        start: span.start,
        end: span.end,
        color: speakerColorPalette[span.speakerID]
    )
}

```

## Summary

- **VoiceActivity.swift** serves as the central coordinator for VAD, managing the `AudioTrackReader`, `SileroVADModel`, and result caching.
- The **Silero VAD model** runs on the MLX engine and requires the `ModelBox` actor to serialize access due to thread-safety constraints.
- Audio is processed in **512-sample chunks (~32 ms)** at 16 kHz, with probabilities converted to speech spans via `VADPipeline.binarize()`.
- **SpeakerIdentity.swift** extends VAD results with diarization by clustering voice characteristics into speaker-specific ranges.
- An **AsyncSemaphore** limits concurrent processing to two files, preventing memory exhaustion during large imports.
- The **DiskCache** system uses modification-time tagging to avoid recomputing analyses when source files remain unchanged.

## Frequently Asked Questions

### How does Palmier Pro handle thread safety when running the Silero VAD model?

The `SileroVADModel` is wrapped in a `ModelBox` actor that serialize all inference calls. Since the underlying MLX implementation is not thread-safe, this actor ensures that only one chunk is processed at a time, even when multiple audio files are being analyzed concurrently.

### What happens if the source audio file is corrupted or contains no audio tracks?

For files lacking audio tracks, the system immediately caches a zero-length analysis via `cacheNoAudioAnalysis()`. For corrupted media detected via AVFoundation error code `-11829` (resource validation failure), the `isDamagedMedia` check returns a specific error without attempting further processing, preventing application crashes.

### Why does Palmier Pro limit concurrent audio processing to two files?

The `AsyncSemaphore` named `pipelineGate` caps concurrency at two simultaneous decode/inference operations. This prevents unbounded heap growth when importing large projects, as each audio buffer and ML tensor consumes significant memory. The limit ensures the application remains responsive while maximizing throughput through pipelining.

### How does the caching mechanism know when to invalidate stored VAD results?

The `DiskCache` generates a `sizeMtimeTag` by combining the media reference identifier with the file system modification time of the source audio. When `VoiceActivity.analysis(for:mediaRef:)` is called, it compares this tag against the cached JSON side-car. If the tags differ—indicating the file has changed since the last analysis—the cached result is discarded and the pipeline runs fresh.