Implementing Voice Activity Detection and Speaker Diarization in Palmier Pro
Palmier Pro performs both voice activity detection (VAD) and speaker diarization using the open-source SpeechVAD package and a custom MLX-powered analysis pipeline that processes audio in 32 ms chunks while enforcing strict concurrency limits.
The palmier-io/palmier-pro repository implements a production-ready audio analysis stack that combines on-device machine learning with robust error handling. This article examines the Swift implementation responsible for identifying speech segments and assigning speaker identities, referencing the actual source files and concurrency architecture used in the application.
Core VAD Implementation in VoiceActivity.swift
The primary entry point for speech detection is VoiceActivity.analysis(for:mediaRef:) in Sources/PalmierPro/Audio/Analysis/VoiceActivity.swift. This method orchestrates a multi-stage pipeline that transforms raw audio files into structured speech spans.
Audio Decoding with AudioTrackReader
The process begins with AudioTrackReader.readMonoFloats(from:sampleRate:), which loads the source file off the main actor into a mono float buffer sampled at 16 kHz. This decouples the heavy I/O operation from the UI thread and standardizes the input format for the neural network.
MLX-Powered Silero Model
The VAD engine uses a pre-trained Silero model accelerated via MLX. Because the underlying SileroVADModel is not thread-safe, the code wraps model access in a ModelBox actor. This actor serializes all inference calls, preventing race conditions when multiple audio files are processed concurrently.
The model loads lazily via SileroVADModel.fromPretrained(engine: .mlx) and processes audio in fixed 512-sample chunks—approximately 32 ms of audio per frame.
Chunk Processing and Probability Extraction
The detectSpeech(vad:samples:) method iterates through the float buffer using stride(from:to:by:) and feeds each chunk to SileroVADModel.processChunk. The method respects task cancellation via Task.checkCancellation() and aggregates per-frame speech probabilities into an array for downstream processing.
The VAD Pipeline and Configuration
Raw probabilities require transformation into boolean speech masks. The VADPipeline class, configured via VADConfig, handles this binarization.
The configuration dynamically sets windowDuration based on the number of probability frames multiplied by the chunk duration (Float(probabilities.count) * chunkDuration). Calling pipeline.binarize(probs:) produces an array of VoiceActivity.Span structs, each containing start and end times for contiguous speech segments.
This abstraction separates the probabilistic model output from the deterministic logic that merges adjacent frames into usable time ranges.
Speaker Diarization with SpeakerIdentity.swift
Building upon the VAD output, Sources/PalmierPro/Audio/Analysis/SpeakerIdentity.swift implements speaker diarization by importing the SpeechVAD package (available when BUNDLED_SPEECH is enabled during compilation).
The module follows the same architectural pattern as the VAD implementation: it reads the audio asset, runs the speech detection model to generate masks, and then feeds these masks into a clustering engine. This engine analyzes voice characteristics to group speech segments by speaker identity.
The resulting tuples—containing speakerID, start, and end times—are consumed by Sources/PalmierPro/Editor/ViewModel/EditorViewModel+Speakers.swift to render timeline tints, speaker labels, and silence-wash visualizations in the user interface.
Concurrency and Memory Management
Palmier Pro enforces strict resource limits to prevent unbounded memory usage during batch processing.
ModelBox Actor for Thread Safety
The ModelBox actor guarantees exclusive access to the Silero model, satisfying the library's requirement that inference calls must not overlap. Any concurrent requests to analyze multiple files are automatically serialized at the model access layer.
AsyncSemaphore for Pipeline Gating
An AsyncSemaphore named pipelineGate caps concurrent decode and inference operations to two files maximum. The semaphore is awaited before decoding begins and released after inference completes. This bounds memory consumption when processing large projects with numerous audio tracks.
Caching Strategy and Edge Cases
The implementation includes defensive programming for production reliability.
Disk Cache with Modification-Time Tagging
Analysis results are persisted as JSON side-car files using DiskCache. The cache key combines the media reference with a sizeMtimeTag generated from the source file’s modification time. The method cachedAnalysis(for:mediaRef:) checks for existing results before re-running the expensive ML pipeline.
Damaged Media and No-Audio Tracks
- No-audio tracks: When a media asset lacks an audio track, the system caches a zero-length analysis via
cacheNoAudioAnalysisto avoid redundant processing attempts. - Damaged files: Errors from AVFoundation (specifically domain
AVFoundationErrorDomain, code-11829) are recognized via theisDamagedMediacheck and reported as failures without crashing the application. - Cancellation: Both the chunking loop and the inference engine respect cooperative cancellation through
Task.checkCancellation()andMLXRuntime.shouldStop.
Practical Code Examples
The following patterns demonstrate how to integrate VAD and diarization into your own Palmier Pro workflows.
Running Voice Activity Detection
import PalmierPro
// Analyze an audio asset for speech segments
let mediaRef = "track_01"
let analysis = try await VoiceActivity.analysis(for: audioURL, mediaRef: mediaRef)
// Access the boolean mask (one value per 32ms chunk)
let mask = analysis.mask
let cellDuration = VoiceActivity.chunkDuration
// Map mask indices to timeline positions
for (index, isSpeech) in mask.enumerated() {
let startTime = Double(index) * cellDuration
if isSpeech {
print("Speech detected at \(startTime)s")
}
}
Performing Speaker Diarization
import SpeechVAD
// Cluster speech segments by speaker identity
let speakerSpans = try await SpeakerIdentity.analyze(
sourceURL: audioURL,
mediaRef: mediaRef
)
// Apply labels to timeline UI
for span in speakerSpans {
timelineView.addSpeakerLabel(
id: span.speakerID,
start: span.start,
end: span.end,
color: speakerColorPalette[span.speakerID]
)
}
Summary
- VoiceActivity.swift serves as the central coordinator for VAD, managing the
AudioTrackReader,SileroVADModel, and result caching. - The Silero VAD model runs on the MLX engine and requires the
ModelBoxactor to serialize access due to thread-safety constraints. - Audio is processed in 512-sample chunks (~32 ms) at 16 kHz, with probabilities converted to speech spans via
VADPipeline.binarize(). - SpeakerIdentity.swift extends VAD results with diarization by clustering voice characteristics into speaker-specific ranges.
- An AsyncSemaphore limits concurrent processing to two files, preventing memory exhaustion during large imports.
- The DiskCache system uses modification-time tagging to avoid recomputing analyses when source files remain unchanged.
Frequently Asked Questions
How does Palmier Pro handle thread safety when running the Silero VAD model?
The SileroVADModel is wrapped in a ModelBox actor that serialize all inference calls. Since the underlying MLX implementation is not thread-safe, this actor ensures that only one chunk is processed at a time, even when multiple audio files are being analyzed concurrently.
What happens if the source audio file is corrupted or contains no audio tracks?
For files lacking audio tracks, the system immediately caches a zero-length analysis via cacheNoAudioAnalysis(). For corrupted media detected via AVFoundation error code -11829 (resource validation failure), the isDamagedMedia check returns a specific error without attempting further processing, preventing application crashes.
Why does Palmier Pro limit concurrent audio processing to two files?
The AsyncSemaphore named pipelineGate caps concurrency at two simultaneous decode/inference operations. This prevents unbounded heap growth when importing large projects, as each audio buffer and ML tensor consumes significant memory. The limit ensures the application remains responsive while maximizing throughput through pipelining.
How does the caching mechanism know when to invalidate stored VAD results?
The DiskCache generates a sizeMtimeTag by combining the media reference identifier with the file system modification time of the source audio. When VoiceActivity.analysis(for:mediaRef:) is called, it compares this tag against the cached JSON side-car. If the tags differ—indicating the file has changed since the last analysis—the cached result is discarded and the pipeline runs fresh.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →