Building an Audio Waveform Extraction and Visualization Pipeline in Palmier Pro: A Deep Dive

Palmier Pro generates audio waveforms on-demand using a background pipeline that caches results to disk, limits concurrent extractions with an AsyncSemaphore, and keeps the UI responsive via @MainActor isolation.

The waveform visualization system in Palmier Pro handles heavy audio decoding without blocking the timeline interface. At its core, the MediaVisualCache class manages the entire lifecycle—from triggering generation to persisting cached *.waveform2 files—ensuring that large media files don't freeze the editor during analysis.

The MediaVisualCache Architecture

The pipeline centers on MediaVisualCache, a @MainActor-isolated singleton defined in Sources/PalmierPro/Timeline/MediaVisualCache.swift. This class owns separate dictionaries for waveform samples, speech masks, beat analysis, and thumbnails, allowing independent cache eviction strategies for each visual asset type.

MainActor Isolation and Thread Safety

All mutable state lives on the main actor, including the waveformSamples dictionary and the waveformInFlight tracking set. However, the actual audio extraction runs in a detached utility task created via Task.detached(priority: .utility). This pattern guarantees that UI-critical code never blocks on I/O while ensuring that cache updates always occur on the main thread.

Disk Persistence Strategy

Extracted waveforms persist as *.waveform2 files managed by DiskCache (defined in Sources/PalmierPro/Utilities/DiskCache.swift). When loadOrGenerateWaveform(url:) executes, it first attempts to read from disk before falling back to AVAssetReader decoding. This enables fast app relaunch and avoids re-processing multi-gigabyte audio files.

How the Pipeline Works

The extraction process follows a strict lazy-generation pattern triggered only when the timeline view requests samples:

  1. Entry Point – generateWaveform(for:) checks if the asset is audio-only or video-with-audio, then verifies whether samples already exist in waveformSamples or if generation is already waveformInFlight.

  2. Concurrency Gating – A static AsyncSemaphore named waveformGate limits concurrent extractions to two instances. The semaphore is awaited before extraction begins and released in a defer block after completion, preventing the extractor from starving playback threads.

  3. Background Extraction – The actual decoding runs in Task.detached(priority: .utility), using AVAssetReader to normalize samples to a 0 (loud) to 1 (silence) range.

  4. Cache Update – Upon completion, the task returns to the main actor, removes the key from waveformInFlight, stores samples in waveformSamples, and marks the timeline view as needing display.

The UI accesses cached data through the non-isolated accessor samples(for:), which safely bridges main-actor state for SwiftUI draw calls:

// Request generation when an asset appears on timeline
let asset: MediaAsset = // audio or video-with-audio
MediaVisualCache.shared.generateWaveform(for: asset)

// Later, retrieve normalized samples for drawing
if let samples = MediaVisualCache.shared.samples(for: asset.id) {
    // samples: [Float] where 0 = loud, 1 = silence
    WaveformView(samples: samples)
}

Speech and Beat Analysis Integration

The waveform pipeline feeds downstream analysis modules without redundant decoding. Within generateWaveform(for:), the cache immediately spawns two additional tasks:

Both modules expose completion callbacks (onMaskReady, onBeatsReady) that trigger UI redraws independently, maintaining the cache's extensibility for future analysis types like pitch detection.

Cache Invalidation and User Control

Users can force regeneration through the storage settings UI. In Sources/PalmierPro/Settings/StoragePane.swift (line 28), the "waveforms, filmstrip thumbnails, and transcripts" option calls resetSessionState(), which wipes the in-memory waveformSamples and speech data dictionaries:

// Clear all cached visual data
MediaVisualCache.shared.resetSessionState()

This operation is side-effect-free regarding the editor's undo history, ensuring that clearing visual caches doesn't create undo groups or modify the actual project timeline.

Summary

  • Lazy Generation: Waveforms extract only when timeline views request them, conserving CPU and battery.
  • Concurrency Safety: AsyncSemaphore caps concurrent extractions at two, while @MainActor isolation protects UI state.
  • Dual Caching: In-memory dictionaries provide fast redraws; *.waveform2 disk files persist across sessions.
  • Modular Analysis: SpeechMaskStore and BeatStore consume waveform samples via callbacks, enabling extensible audio analysis.
  • User Control: Settings integration allows clearing cached visuals without affecting undo history.

Frequently Asked Questions

How does Palmier Pro prevent waveform extraction from freezing the UI?

The system uses Task.detached(priority: .utility) to run AVAssetReader decoding off the main thread, coordinated by an AsyncSemaphore that limits concurrent operations to two. The @MainActor isolation on MediaVisualCache ensures that only cache updates—not extraction logic—touch UI state.

What file format does Palmier Pro use to cache extracted waveforms?

Waveforms persist as binary *.waveform2 files stored via DiskCache in Sources/PalmierPro/Utilities/DiskCache.swift. These files contain normalized Float arrays where 0 represents maximum amplitude and 1 represents silence, enabling rapid reloading without re-decoding source media.

Can I access waveform samples for custom visualization or analysis?

Yes. Use the non-isolated samples(for:) method on MediaVisualCache.shared to retrieve the [Float] array for any cached asset ID. This accessor safely bridges the @MainActor state for drawing operations in SwiftUI or AppKit views.

Why does the pipeline limit concurrent extractions to two operations?

The static waveformGate semaphore restricts concurrent extractions to prevent audio decoding from starving playback threads or other real-time UI work. This cap ensures that waveform generation remains a background-friendly process even when multiple large files populate the timeline simultaneously.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →