# How Vorssaint Implements Per-App Audio Mixing Using CoreAudio

> Learn how Vorssaint achieves per-app audio mixing with CoreAudio. Discover its innovative process tap and aggregate device routing for real-time buffer processing. Explore the vorssaint utils repository.

- Repository: [vorssaint/vorssaint-utils](https://github.com/vorssaint/vorssaint-utils)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Vorssaint captures per-app audio mixing by creating a CoreAudio process tap that excludes its own process from the system-wide mix, then routes that tap through an aggregate audio device for real-time buffer processing.**

The `vorssaint-utils` repository demonstrates a sophisticated approach to per-app audio mixing on macOS without requiring kernel extensions or third-party drivers. By leveraging private CoreAudio APIs, the implementation captures every application's audio output while deliberately omitting the host application's own stream, enabling clean isolation for recording and analysis workflows.

## CoreAudio Process Tap Architecture

The foundation of Vorssaint's per-app audio mixing resides in [`RecorderSystemAudioTap.swift`](https://github.com/vorssaint/vorssaint-utils/blob/main/RecorderSystemAudioTap.swift), where the system constructs a **process tap** using `CATapDescription`. This tap is configured with `stereoGlobalTapButExcludeProcesses`, listing the current process—obtained via `ownProcessObject()`—as the sole exclusion.

```swift
// Conceptual representation of the tap configuration
let tapDescription = CATapDescription(stereoGlobalTapButExcludeProcesses: [ownProcessObject()])

```

This configuration solves the "double-render" problem: by excluding the calling process from the global mix, Vorssaint prevents capturing its own audio generation while still receiving every other application's output. The tap description is then passed to `AudioObjectPropertyAddress` and `AudioObjectGetPropertyData` to instantiate the actual tap object, which serves as a virtual audio source representing the filtered system mix.

## Building the Aggregate Audio Device Pipeline

CoreAudio prohibits direct reading from process taps; instead, the tap must attach to an **aggregate audio device** that provides a stable hardware clock and routing infrastructure. The `buildPipeline()` method in [`RecorderSystemAudioTap.swift`](https://github.com/vorssaint/vorssaint-utils/blob/main/RecorderSystemAudioTap.swift) constructs this device by passing a configuration dictionary to `AudioHardwareCreateAggregateDevice`.

The aggregate configuration specifies the default system output as the host device and the process tap as a sub-tap, creating a unified audio object that merges the hardware clock with the virtual tap stream. This architectural requirement ensures that the tap's buffers can be read through standard CoreAudio IOProcs while maintaining synchronization with physical audio hardware.

## Real-Time Audio Processing with IOProc

Once the aggregate device initializes, Vorssaint establishes an `AudioDeviceIOProcID` using `AudioDeviceCreateIOProcIDWithBlock`. Inside this real-time callback block, the implementation performs four critical operations:

- **Output Silencing**: Calls `MixerRender.silence` on the aggregate device's output buffer to prevent audio feedback or stray playback
- **Tap Buffer Location**: Identifies the tap's specific buffer within the input `AudioBufferList` using `MixerRender.tapBufferIndex`
- **Sound Detection**: Applies `RecorderAudioProbe.containsSound` (vectorized peak detection) to determine if the buffer contains actual audio content, raising the `heard` flag when signal is present
- **Sample Buffer Construction**: Converts the buffer's host time to the recording timeline using `CMClockMakeHostTimeFromSystemUnits` and `CMSyncConvertTime`, then packages the audio into a `CMSampleBuffer` forwarded via the `onSample` closure

```swift
// Starting the audio tap with synchronization
if let tap = await RecorderSystemAudioTap.make() {
    tap.onSample = { sampleBuffer in
        // Process CMSampleBuffer for recording or analysis
    }
    
    let clock = CMClockGetHostTimeClock()
    let readerRunning = await tap.start(synchronizingTo: clock)
    // readerRunning is false if system audio permission is denied
}

```

## Handling Audio Device Changes

The implementation monitors the audio environment through two observation mechanisms in `watchDefaultOutputDevice` and `watchSampleRate`. When the default output device changes or the aggregate device's sample rate shifts, `rebuildPipelineIfChanged()` executes a teardown and reconstruction sequence on the dedicated `teardownQueue`.

This dynamic reconfiguration ensures that per-app audio mixing continues uninterrupted when users switch from headphones to speakers or when sample rate changes occur due to external audio interface adjustments. The `onReaderLost` closure notifies consuming code when the pipeline becomes invalid due to hardware changes, enabling graceful fallback behavior.

## Lifecycle and Permission Management

The `start(synchronizingTo:)` method returns a Boolean indicating whether the system granted audio capture permission, allowing the application to detect when macOS security policies block system audio recording. The `stop()` method performs clean resource deallocation by removing property listeners, stopping the IOProc, and destroying both the aggregate device and the process tap on the `teardownQueue` to avoid real-time thread blocking.

```swift
// Stopping the tap and releasing resources
await tap.stop()

// Checking if any audio was captured
let hadSound = tap.heardSound

```

## Summary

- **Process Tap Exclusion**: Vorssaint uses `CATapDescription` with `stereoGlobalTapButExcludeProcesses` to capture all system audio except the host application, located in [`Sources/Vorssaint/Services/Recorder/RecorderSystemAudioTap.swift`](https://github.com/vorssaint/vorssaint-utils/blob/main/Sources/Vorssaint/Services/Recorder/RecorderSystemAudioTap.swift)
- **Aggregate Device Requirement**: The tap must attach to an aggregate device via `AudioHardwareCreateAggregateDevice` to provide a readable, clock-synchronized audio stream
- **Real-Time Processing**: An `AudioDeviceIOProcIDWithBlock` callback extracts tap buffers, applies vectorized sound detection, and converts timestamps for `CMSampleBuffer` generation
- **Dynamic Reconfiguration**: The system watches default output devices and sample rates automatically, rebuilding the pipeline when hardware configurations change
- **Permission Handling**: The `start` method reports whether macOS granted system audio recording permission, while `stop` ensures proper resource cleanup on a dedicated queue

## Frequently Asked Questions

### What is a CoreAudio process tap and how does it enable per-app audio mixing?

A CoreAudio process tap is a private API mechanism that intercepts audio streams from specific processes or the entire system. Vorssaint creates a global tap that excludes only its own process using `CATapDescription(stereoGlobalTapButExcludeProcesses:)`, effectively receiving a mixed audio stream containing every other application's output. This approach provides per-app isolation without requiring individual process enumeration or multiple tap instances.

### Why does Vorssaint use an aggregate audio device instead of reading the tap directly?

CoreAudio process taps cannot be read directly through standard audio device interfaces. The implementation must wrap the tap in an aggregate device using `AudioHardwareCreateAggregateDevice`, which provides the necessary clock domain and `AudioDeviceIOProc` interface. This aggregate structure bridges the tap's virtual audio stream with CoreAudio's hardware abstraction layer, enabling standard buffer reading patterns while maintaining synchronization with physical audio output devices.

### How does the system handle permission requests for capturing system audio?

The `start(synchronizingTo:)` method attempts to initialize the audio pipeline and returns a Boolean indicating success. If macOS has not granted the application permission to record system audio, or if the user denies the request, the method returns `false` without throwing an exception. This allows the application to detect permission states and prompt users to enable system audio recording in Security & Privacy preferences.

### What happens when the default audio output device changes during recording?

Vorssaint monitors the default output device through `watchDefaultOutputDevice` and the aggregate's sample rate via `watchSampleRate`. When either changes, `rebuildPipelineIfChanged()` automatically tears down the existing aggregate device and tap on the `teardownQueue`, then reconstructs the pipeline with the new hardware configuration. This ensures continuous audio capture across hardware switches without dropping the recording session, though the `onReaderLost` callback notifies the UI if temporary interruption occurs.