How Screen Recording in Vorssaint-Utils Captures System Audio and Microphone on Separate Tracks

Vorssaint-utils captures system audio and microphone input as two independent tracks by configuring separate AVAssetWriterInput instances for each source, routing system audio through RecorderSystemAudioTap and microphone audio through RecorderMicrophoneCapture, then tagging each track with distinct metadata for downstream identification.

Screen recording utilities often struggle to isolate system sounds from voice commentary, but vorssaint-utils solves this with a dual-track architecture. By leveraging distinct capture paths and separate AVAssetWriter inputs, the library ensures that Mac system audio and microphone input remain isolated throughout the recording pipeline. This implementation allows editors to adjust volume levels independently or mute specific tracks during post-production.

The Dual-Track Architecture Overview

Audio Preferences with RecorderSelectionAudioOptions

The process begins in Sources/Vorssaint/Services/Recorder/ScreenRecorderService.swift with the RecorderSelectionAudioOptions class (lines 10‑21). This lightweight structure stores user preferences as two Boolean flags: systemAudio and microphone. These values persist in UserDefaults, allowing the application to remember audio configurations between sessions.

Orchestration via ScreenRecorderService

When a user initiates recording, the ScreenRecorderService.record(_:audioOptions:) method (lines 75‑84) reads these flags and instantiates a RecorderSession with the appropriate capture configurations. This separation of concerns ensures that audio source selection happens independently from the actual recording mechanics.

Separate Capture Paths for System Audio and Microphone

System Audio Capture via RecorderSystemAudioTap

When capturesSystemAudio is enabled, RecorderSession creates a RecorderSystemAudioTap instance. This component utilizes low-level audio tapping APIs to capture the Mac's output stream without interfering with the hardware microphone path. The tap generates CMSampleBuffer objects containing raw audio data from the system.

Microphone Input via RecorderMicrophoneCapture

Concurrently, if capturesMicrophone is true, the session instantiates RecorderMicrophoneCapture to access the built-in or external microphone. This creates a completely independent capture stream, ensuring that voice commentary and system sounds never mix before reaching the encoder.

Writing Dual Audio Tracks with AVAssetWriter

Configuring Independent Audio Inputs

In Sources/Vorssaint/Services/Recorder/RecorderWriter.swift, the initializer (lines 68‑90) constructs the AVAssetWriter and immediately adds two distinct AVAssetWriterInput instances: systemAudioInput and microphoneInput. Each input receives specific metadata through RecorderAudioSource.system.trackMetadata and RecorderAudioSource.microphone.trackMetadata, embedding identification tags directly into the track headers.

Routing Samples by Kind

The RecorderSession.append(_:kind:) method (lines 108‑122) acts as a traffic controller, receiving every CMSampleBuffer from the capture engines. The kind parameter—differentiating between video, system audio, and microphone—determines the routing destination. The session forwards these buffers to RecorderWriter.append(_:kind:), which switches on the kind value and writes system audio samples to systemAudioInput and microphone samples to microphoneInput.

Finalizing the Recording

When the user stops the capture, the finish(at:) method (lines 14‑20) in RecorderWriter marks both audio inputs as finished by calling systemAudioInput?.markAsFinished() and microphoneInput?.markAsFinished() before finalizing the file. This ensures the resulting .mov container contains properly terminated tracks that maintain their separation throughout the file structure.

Implementation Examples

Starting a Recording with Both Tracks

import Vorssaint

// 1. Choose audio options (e.g. from UI)
let audioOptions = RecorderSelectionAudioOptions()
audioOptions.systemAudio = true            // capture system output
audioOptions.microphone = true            // capture mic

// 2. Define the region you want to record (full screen in this case)
let region = RecorderSupport.Region.fullScreen

// 3. Ask the shared service to start recording
ScreenRecorderService.shared.record(region, audioOptions: audioOptions)

Accessing Tracks After Recording

// The editor receives a `RecorderTakeStore.Take` whose video URL points to a .mov file.
// You can inspect the file with AVFoundation:

let asset = AVAsset(url: take.videoURL)
let audioTracks = asset.tracks(withMediaType: .audio)

// Expect exactly two tracks if both options were enabled
print("Audio track count:", audioTracks.count)   // → 2

// Identify them by metadata
for track in audioTracks {
    if let meta = track.metadata.first(where: { $0.identifier == .commonIdentifierTitle }) {
        print("Track title:", meta.stringValue ?? "unknown")
    }
}

Disabling One Track

// In the UI, toggle the switches and persist them:
UserDefaults.standard.set(false, forKey: DefaultsKey.recorderSystemAudio)   // no system audio
UserDefaults.standard.set(true,  forKey: DefaultsKey.recorderMicrophone)  // mic only

Summary

  • Dual-track architecture: Vorssaint-utils configures separate AVAssetWriterInput instances for system audio and microphone, ensuring physical separation in the output file.
  • Independent capture paths: RecorderSystemAudioTap captures Mac output while RecorderMicrophoneCapture handles input devices, preventing signal mixing at the source.
  • Metadata tagging: Each track receives distinct metadata identifiers (RecorderAudioSource.system.trackMetadata and RecorderAudioSource.microphone.trackMetadata) for programmatic identification.
  • Clean finalization: The finish(at:) method properly closes both audio tracks before sealing the container, preserving track integrity.

Frequently Asked Questions

How does vorssaint-utils prevent system audio and microphone from mixing during recording?

The library maintains separate capture instances—RecorderSystemAudioTap for system output and RecorderMicrophoneCapture for microphone input—each feeding distinct AVAssetWriterInput streams. This architectural separation ensures samples never merge until final multiplexing, keeping channels isolated in the output MOV file.

Can I enable only one audio track instead of both?

Yes. The RecorderSelectionAudioOptions class stores Boolean flags for each source. Setting systemAudio or microphone to false prevents the corresponding RecorderSession from instantiating that capture component, resulting in a single-track recording.

How are the audio tracks identified in the final video file?

During initialization in RecorderWriter (lines 68‑90), each audio input receives metadata through trackMetadata properties of the RecorderAudioSource enum. These embedded tags allow downstream tools to distinguish tracks by inspecting the asset's metadata using AVFoundation.

Which source files control the audio routing logic?

The primary orchestration happens in Sources/Vorssaint/Services/Recorder/ScreenRecorderService.swift, where RecorderSession manages capture lifecycles. The actual writing logic resides in Sources/Vorssaint/Services/Recorder/RecorderWriter.swift, which handles AVAssetWriter configuration and sample routing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →