Architecting Audio Beat Detection with AVFoundation in Palmier Pro

Palmier Pro performs on-device beat detection by decoding audio to 22 kHz mono PCM, running chunked Core ML inference with the BeatThis model, and caching results to disk for instant UI and Agent lookups.

The palmier-io/palmier-pro repository implements a complete audio beat detection with AVFoundation audio processing pipeline that analyzes media files entirely on-device. By combining AVAssetReader-based extraction with a custom Core ML inference strategy, the system delivers tempo and downbeat detection while respecting strict concurrency limits that preserve UI responsiveness on macOS.

Audio Extraction Pipeline

The beat detection workflow begins in Sources/PalmierPro/Audio/Beats/BeatDetector.swift with the decodeAudio(from:) method. This utility leverages AVFoundation to stream audio into a mono Float32 PCM buffer sampled at 22 kHz, the specific rate required by the bundled machine learning model.

All decoding operations run off the main actor to prevent frame drops in the SwiftUI interface. The method returns a flat [Float] array representing the normalized audio waveform, which serves as the input tensor for downstream inference.

Chunked Core ML Inference Architecture

Rather than processing entire files at once, the predictLogits(samples:model:) method slices the 22 kHz samples into overlapping chunks of 1500 frames with a 6-frame border padding. This padding ensures that interior frames remain free of edge artifacts during convolutional operations.

Each padded chunk feeds into BeatThis.mlmodelc, the on-device Core ML model that outputs per-frame logits for both beats and downbeats. The chunked approach balances memory pressure against throughput, allowing the system to process long audio files without loading the entire waveform into RAM simultaneously.

Post-Processing and Caching Strategy

After inference, the raw logits undergo probabilistic thresholding. The pickPeaks function applies a sigmoid activation followed by a 0.5 threshold to convert model outputs into discrete beat timestamps. Subsequently, estimateBPM calculates the tempo by measuring the median inter-beat interval across detected peaks.

To eliminate redundant computation, the system integrates DiskCache from Sources/PalmierPro/Utilities/DiskCache.swift. The cache maps each media file’s modification-time tag to a JSON file (*_beats.json) containing the serialized BeatAnalysis struct. Stale entries are automatically pruned after each analysis cycle, ensuring that renamed or edited files trigger fresh inference while unchanged assets load instantly from disk.

Concurrency and Threading Model

The BeatDetector class is marked @unchecked Sendable and utilizes an internal actor ModelBox to serialize access to the Core ML prediction interface. Because MLModel.prediction is not thread-safe, the actor wrapper guarantees that inference requests execute sequentially even when multiple analysis tasks run concurrently.

Resource exhaustion is prevented by two AsyncSemaphore gates:

  • pipelineGate throttles the end-to-end analysis pipeline
  • cacheLookupGate limits simultaneous cache write operations

Both gates cap concurrent operations at two simultaneous analyses, a tuning constant optimized for typical macOS hardware to prevent CPU and IO saturation. The Core ML model loads once via static let shared = try? BeatDetector() and remains resident in memory, avoiding the compilation cost of repeated model instantiation.

Implementation Examples

Request a fresh beat analysis for a media asset:

import PalmierPro

let mediaURL = URL(fileURLWithPath: "/path/to/clip.mov")
let mediaRef = "clip-01"

Task {
    do {
        let analysis = try await BeatDetector.analysis(
            for: mediaURL,
            mediaRef: mediaRef,
            force: false
        )
        
        print("BPM:", analysis.bpm)
        print("Beat timestamps:", analysis.beats)
        print("Downbeat timestamps:", analysis.downbeats)
    } catch {
        print("Beat detection failed:", error)
    }
}

Query the cache directly for instantaneous UI updates:

Task {
    if let entry = await BeatDetector.cachedAnalysis(
        for: mediaURL,
        mediaRef: mediaRef
    ) {
        print("Cached BPM:", entry.analysis.bpm)
    }
}

Integrate with Agent tools in Sources/PalmierPro/Agent/Tools/ToolExecutor+Beats.swift:

let analysis = try await BeatDetector.analysis(
    for: sourceURL,
    mediaRef: mediaKey,
    force: false
)
let bpm = analysis.bpm
// Use bpm to drive tempo-based editing workflows

Summary

  • AVAssetReader streams audio at 22 kHz mono Float32 PCM via decodeAudio(from:) in BeatDetector.swift
  • Chunked inference processes 1500-frame windows with 6-frame padding using the BeatThis.mlmodelc model
  • Peak picking applies sigmoid thresholding at 0.5, while BPM derives from median inter-beat intervals
  • DiskCache persists BeatAnalysis JSON alongside media file tags for sub-millisecond retrieval
  • Concurrency control uses AsyncSemaphore gates and a ModelBox actor to serialize Core ML predictions off the main thread

Frequently Asked Questions

What audio sampling rate does Palmier Pro require for beat detection?

The system decodes all input media to 22 kHz mono Float32 PCM before inference. This specific rate matches the training configuration of the BeatThis.mlmodelc model and ensures optimal prediction accuracy.

How does Palmier Pro prevent beat detection from blocking the UI?

All file I/O and model prediction run off the main actor. An internal actor ModelBox serializes Core ML calls, while two AsyncSemaphore gates limit the pipeline to two simultaneous analyses, preventing CPU starvation that could degrade interface responsiveness.

What is the chunk size for audio processing in the inference pipeline?

The predictLogits(samples:model:) method processes audio in overlapping chunks of 1500 frames with 6 frames of border padding. This window size balances computational efficiency against memory usage for long media files.

How does the caching mechanism store beat analysis results?

DiskCache writes a JSON representation of BeatAnalysis to a *_beats.json file keyed by the media file’s modification-time tag. The BeatDetector.cachedAnalysis(for:mediaRef:) method retrieves these entries instantly, avoiding redundant inference for unchanged assets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →