Implementing Dead Air Detection for Audio Enhancement in Palmier Pro
Palmier Pro detects and removes silent, speech‑free sections by combining on‑device Voice‑Activity Detection (VAD) with a specialized algorithm that identifies quiet spans below the speech floor, all while maintaining A/V sync and full undo support.
Palmier Pro is an open‑source video editing framework that leverages intelligent audio analysis to accelerate post‑production workflows. The dead air detection system automatically identifies non‑speech intervals in dialogue tracks, enabling filmmakers to tighten pacing through automated ripple‑deletes without manual scrubbing. This implementation runs entirely on‑device in Swift, utilizing session‑scoped caching and main‑actor isolation for UI consistency.
How Dead Air Detection Works
The system operates through two tightly coupled subsystems that process audio in 32‑millisecond chunks.
Voice‑Activity Detection (VAD) Subsystem
The first stage runs an on‑device neural network to classify each audio frame as speech or non‑speech. In Sources/PalmierPro/Audio/Analysis/VoiceActivity.swift, the method VoiceActivity.analysis(for:mediaRef:) produces a Boolean array called the speech mask, where true indicates speech presence and false indicates silence or background noise.
This analysis is invoked asynchronously through SpeechMaskStore.generate(for:), which queues a detached background task to keep heavy audio decoding off the main thread.
Deriving the Dead‑Air Mask from Speech Data
Once the VAD mask and normalized waveform samples are available, SpeechMaskStore derives a secondary dead‑air mask. The implementation in SpeechMaskStore.buildDeadAirMask(speech:samples:) (lines 92‑120) applies the following logic:
- Compute the peak amplitude for each 32 ms VAD cell (
cellPeak) - Calculate a quiet floor from the median of all speech‑cell peaks, then add a safety margin (
speechGap = 0.24) - Scan for consecutive non‑speech cells exceeding
minCells(approximately 0.26 seconds) whose median peak sits below the quiet floor - Mark qualifying cells as
truein the dead‑air mask
This approach distinguishes meaningful pauses from breathable mic noise by anchoring the threshold to actual speech levels in the specific clip.
The Dead Air Removal Workflow
The end‑to‑end pipeline moves from raw audio to timeline edits through five distinct stages.
Step 1: Requesting VAD Analysis
When a media asset enters the project, EditorViewModel triggers analysis by calling SpeechMaskStore.shared.generate(for: asset). This stores the resulting speechMasks in a dictionary keyed by the asset’s id, as implemented in Sources/PalmierPro/Audio/Analysis/SpeechMaskStore.swift (lines 17‑30).
Step 2: Building the Dead‑Air Mask
The mask generation is lazy. Calling deadAirMask(for:samples:) checks the cache first; if absent, it invokes buildDeadAirMask to compute and store the Boolean array according to the algorithm described above.
Step 3: Mapping Masks to Timeline Ranges
EditorViewModel+DeadAir.swift translates the abstract cell indices into concrete timeline ranges. The method deadAirRanges(for:) walks the dead‑air mask and:
- Converts cell indices to source‑frame times using
cellFrames = VoiceActivity.chunkDuration * timeline.fps - Maps source times to timeline frames via
timelineRange(clip:sourceStart:sourceEnd:)
For multicam clips, the system automatically adjusts for sync offsets (shift) before mapping (lines 7‑14).
Step 4: Performing Ripple‑Deletes
The public API exposes two removal entry points defined in EditorViewModel+DeadAir.swift:
removeDeadAir(clipId:atTimelineFrame:)– Removes the specific dead‑air span intersecting the playheadremoveAllDeadAir()– Removes every detected dead‑air section across the entire project in a single atomic operation
Both methods invoke rippleDeleteRangesOnTrack, which closes gaps while preserving linked audio‑video pairs and registers the operation under undo.perform("Remove Dead Air") for a single‑step undo (lines 59‑89).
Step 5: Agent Integration for Automated Editing
The AI Agent can invoke dead‑air removal through the tool defined in ToolDefinitions.swift (lines 718‑730). The description explains the VAD‑derived spans, allowing the language model to request targeted edits. Execution occurs in ToolExecutor+Words.swift, which throws an error if no dead‑air is present at the requested location.
Architectural Design Patterns
The Palmier Pro implementation prioritizes performance and editorial safety through several Swift‑specific patterns.
Threading and Actor Isolation
VAD processing runs inside Task.detached(priority: .utility) to prevent blocking the UI during audio decoding. However, all state mutations to speechMasks, deadAirMasks, and the inFlight tracking set are confined to @MainActor, guaranteeing that the UI never reads stale cache data or races with background writes.
Session‑Scoped Caching Strategy
SpeechMaskStore maintains in‑memory caches for both speech and dead‑air masks. These invalidate automatically when assets change via invalidate(_:) or globally through reset(). This design avoids recomputing expensive VAD results when users scrub back and forth across the same clip.
Multicam Offset Handling
When a clip belongs to a multicamera group, the dead‑air mask retrieval adjusts for the group’s synchronization offset. The shift parameter in deadAirMask(for:) (lines 7‑14 of EditorViewModel+DeadAir.swift) ensures that silence detected in the master audio aligns correctly with the corresponding video angle.
Undoability and Transaction Safety
Every deletion wraps its mutations in a named undo transaction. Whether removing a single span or processing the entire timeline, the user can revert the operation with one Cmd‑Z because rippleDeleteRangesOnTrack integrates with the editor’s command history.
Implementation Code Examples
Trigger analysis for a newly imported asset:
let asset = MediaAsset(id: "intro-audio", url: audioURL)
SpeechMaskStore.shared.generate(for: asset)
Query dead‑air spans for a specific clip to visualize gaps in the UI:
if let clip = editorViewModel.findClip(id: "clip-123") {
let deadAirRanges = editorViewModel.deadAirRanges(for: clip)
// deadAirRanges contains FrameRange objects mapped to the timeline
}
Remove dead‑air at the playhead location:
editorViewModel.removeDeadAir(clipId: "clip-123", atTimelineFrame: 250)
Perform a global dead‑air removal with atomic undo support:
if let report = editorViewModel.removeAllDeadAir() {
print("Removed \(report.sections) dead-air sections, totaling \(report.removedFrames) frames")
}
Summary
- Two‑stage detection: VAD produces a speech mask, then
buildDeadAirMaskderives quiet spans using median speech levels and a 0.24 safety gap. - Timeline integration:
EditorViewModel+DeadAir.swiftconverts 32 ms cells into frame‑accurate timeline ranges, handling multicam offsets automatically. - Performance safety: Analysis runs on utility‑priority detached tasks, while
@MainActorprotects cache state. - Editorial integrity: Ripple‑deletes preserve A/V sync and register as single undo events, accessible via both UI and Agent tools.
Frequently Asked Questions
How does Palmier Pro distinguish between intentional silence and dead air?
The algorithm in SpeechMaskStore.buildDeadAirMask compares the median amplitude of non‑speech cells against a quiet floor calculated from actual speech peaks in the clip. Only consecutive silent cells exceeding approximately 0.26 seconds with levels well below the speech floor (median speech peak + 0.24 gap) qualify as dead air, filtering out breath sounds and intentional dramatic pauses.
Where is the dead air detection code located in the repository?
The core logic resides in Sources/PalmierPro/Audio/Analysis/SpeechMaskStore.swift for mask generation and Sources/PalmierPro/Editor/ViewModel/EditorViewModel+DeadAir.swift for timeline mapping and deletion. Agent tool definitions live in Sources/PalmierPro/Agent/Tools/ToolDefinitions.swift with execution handled in Sources/PalmierPro/Agent/Tools/ToolExecutor+Words.swift.
Can dead air removal be automated via the AI Agent?
Yes. The Agent exposes a remove_silence tool defined in ToolDefinitions.swift that triggers the same removeDeadAir methods used by the UI. The Agent can query available dead‑air spans and request their removal, though the executor throws an error if no qualifying silence exists at the specified frame.
Does dead air detection work with multicam clips?
Yes. When retrieving masks for clips belonging to a multicam group, EditorViewModel+DeadAir.swift applies a shift offset to align the audio analysis with the correct video angle’s timeline position, ensuring that ripple‑deletes maintain synchronization across grouped angles.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →