# Architecting Audio Beat Detection with AVFoundation in Palmier Pro

> Learn how Palmier Pro architects audio beat detection using AVFoundation and Core ML on-device. Discover efficient audio processing for instant UI and agent lookups.

- Repository: [Palmier/palmier-pro](https://github.com/palmier-io/palmier-pro)
- Tags: architecture
- Published: 2026-07-27

---

**Palmier Pro performs on-device beat detection by decoding audio to 22 kHz mono PCM, running chunked Core ML inference with the BeatThis model, and caching results to disk for instant UI and Agent lookups.**

The palmier-io/palmier-pro repository implements a complete **audio beat detection with AVFoundation audio processing** pipeline that analyzes media files entirely on-device. By combining `AVAssetReader`-based extraction with a custom Core ML inference strategy, the system delivers tempo and downbeat detection while respecting strict concurrency limits that preserve UI responsiveness on macOS.

## Audio Extraction Pipeline

The beat detection workflow begins in [`Sources/PalmierPro/Audio/Beats/BeatDetector.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Audio/Beats/BeatDetector.swift) with the `decodeAudio(from:)` method. This utility leverages **AVFoundation** to stream audio into a mono `Float32` PCM buffer sampled at **22 kHz**, the specific rate required by the bundled machine learning model.

All decoding operations run off the main actor to prevent frame drops in the SwiftUI interface. The method returns a flat `[Float]` array representing the normalized audio waveform, which serves as the input tensor for downstream inference.

## Chunked Core ML Inference Architecture

Rather than processing entire files at once, the `predictLogits(samples:model:)` method slices the 22 kHz samples into overlapping chunks of **1500 frames** with a **6-frame border padding**. This padding ensures that interior frames remain free of edge artifacts during convolutional operations.

Each padded chunk feeds into `BeatThis.mlmodelc`, the on-device Core ML model that outputs per-frame logits for both beats and downbeats. The chunked approach balances memory pressure against throughput, allowing the system to process long audio files without loading the entire waveform into RAM simultaneously.

## Post-Processing and Caching Strategy

After inference, the raw logits undergo probabilistic thresholding. The `pickPeaks` function applies a sigmoid activation followed by a **0.5 threshold** to convert model outputs into discrete beat timestamps. Subsequently, `estimateBPM` calculates the tempo by measuring the median inter-beat interval across detected peaks.

To eliminate redundant computation, the system integrates `DiskCache` from [`Sources/PalmierPro/Utilities/DiskCache.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Utilities/DiskCache.swift). The cache maps each media file’s modification-time tag to a JSON file (`*_beats.json`) containing the serialized `BeatAnalysis` struct. Stale entries are automatically pruned after each analysis cycle, ensuring that renamed or edited files trigger fresh inference while unchanged assets load instantly from disk.

## Concurrency and Threading Model

The `BeatDetector` class is marked `@unchecked Sendable` and utilizes an internal `actor ModelBox` to serialize access to the Core ML prediction interface. Because `MLModel.prediction` is not thread-safe, the actor wrapper guarantees that inference requests execute sequentially even when multiple analysis tasks run concurrently.

Resource exhaustion is prevented by two `AsyncSemaphore` gates:
- **`pipelineGate`** throttles the end-to-end analysis pipeline
- **`cacheLookupGate`** limits simultaneous cache write operations

Both gates cap concurrent operations at **two simultaneous analyses**, a tuning constant optimized for typical macOS hardware to prevent CPU and IO saturation. The Core ML model loads once via `static let shared = try? BeatDetector()` and remains resident in memory, avoiding the compilation cost of repeated model instantiation.

## Implementation Examples

Request a fresh beat analysis for a media asset:

```swift
import PalmierPro

let mediaURL = URL(fileURLWithPath: "/path/to/clip.mov")
let mediaRef = "clip-01"

Task {
    do {
        let analysis = try await BeatDetector.analysis(
            for: mediaURL,
            mediaRef: mediaRef,
            force: false
        )
        
        print("BPM:", analysis.bpm)
        print("Beat timestamps:", analysis.beats)
        print("Downbeat timestamps:", analysis.downbeats)
    } catch {
        print("Beat detection failed:", error)
    }
}

```

Query the cache directly for instantaneous UI updates:

```swift
Task {
    if let entry = await BeatDetector.cachedAnalysis(
        for: mediaURL,
        mediaRef: mediaRef
    ) {
        print("Cached BPM:", entry.analysis.bpm)
    }
}

```

Integrate with Agent tools in `Sources/PalmierPro/Agent/Tools/ToolExecutor+Beats.swift`:

```swift
let analysis = try await BeatDetector.analysis(
    for: sourceURL,
    mediaRef: mediaKey,
    force: false
)
let bpm = analysis.bpm
// Use bpm to drive tempo-based editing workflows

```

## Summary

- **AVAssetReader** streams audio at 22 kHz mono Float32 PCM via `decodeAudio(from:)` in [`BeatDetector.swift`](https://github.com/palmier-io/palmier-pro/blob/main/BeatDetector.swift)
- **Chunked inference** processes 1500-frame windows with 6-frame padding using the `BeatThis.mlmodelc` model
- **Peak picking** applies sigmoid thresholding at 0.5, while BPM derives from median inter-beat intervals
- **DiskCache** persists `BeatAnalysis` JSON alongside media file tags for sub-millisecond retrieval
- **Concurrency control** uses `AsyncSemaphore` gates and a `ModelBox` actor to serialize Core ML predictions off the main thread

## Frequently Asked Questions

### What audio sampling rate does Palmier Pro require for beat detection?

The system decodes all input media to **22 kHz** mono Float32 PCM before inference. This specific rate matches the training configuration of the `BeatThis.mlmodelc` model and ensures optimal prediction accuracy.

### How does Palmier Pro prevent beat detection from blocking the UI?

All file I/O and model prediction run off the main actor. An internal `actor ModelBox` serializes Core ML calls, while two `AsyncSemaphore` gates limit the pipeline to two simultaneous analyses, preventing CPU starvation that could degrade interface responsiveness.

### What is the chunk size for audio processing in the inference pipeline?

The `predictLogits(samples:model:)` method processes audio in overlapping chunks of **1500 frames** with **6 frames of border padding**. This window size balances computational efficiency against memory usage for long media files.

### How does the caching mechanism store beat analysis results?

`DiskCache` writes a JSON representation of `BeatAnalysis` to a `*_beats.json` file keyed by the media file’s modification-time tag. The `BeatDetector.cachedAnalysis(for:mediaRef:)` method retrieves these entries instantly, avoiding redundant inference for unchanged assets.