# How the Media Tool's Video Compression Utilizes AVFoundation in vorssaint-utils

> Discover how vorssaint-utils media tool uses AVFoundation for efficient video compression. Learn about its multi-pass H.264 bitrate adjustment for target file sizes and audio-video sync.

- Repository: [vorssaint/vorssaint-utils](https://github.com/vorssaint/vorssaint-utils)
- Tags: deep-dive
- Published: 2026-09-11

---

**The media tool's video compression leverages AVFoundation through a multi-pass read-encode-write pipeline that dynamically adjusts H.264 bitrate to hit specific target file sizes while maintaining audio-video synchronization.**

The vorssaint-utils repository provides a Swift-based media processing toolkit that implements precise video compression using Apple's AVFoundation framework. At the heart of this implementation, the `MediaVideoTargetEncoder` class orchestrates a sophisticated encoding pipeline designed to compress video files to exact byte targets through iterative bitrate refinement.

## Core Architecture: Source Preparation and Size Planning

The compression workflow begins with structured source preparation and mathematical target planning encapsulated in two primary components.

### The Source Struct and AVAsset Wrapping

In [`MediaVideoTargetEncoder.swift`](https://github.com/vorssaint/vorssaint-utils/blob/main/MediaVideoTargetEncoder.swift), the `Source` struct wraps an `AVAsset` together with its essential metadata. This includes the video track, audio tracks, natural dimensions, preferred transform, and frame rate—providing the encoder with complete context for the re-encoding operation.

```swift
let source = MediaVideoTargetEncoder.Source(
    asset: asset,
    videoTrack: videoTrack,
    audioTracks: audioTracks,
    naturalSize: videoTrack.naturalSize,
    preferredTransform: videoTrack.preferredTransform,
    frameRate: videoTrack.nominalFrameRate
)

```

### Dynamic Bitrate Planning with MediaSupport

Before encoding begins, `MediaSupport.videoSizePlan` (called from the `encode` method) computes a `MediaVideoSizePlan` that estimates the required bitrate to keep the output under the desired byte budget. This plan serves as the baseline for the first encoding pass, with provisions for scaling in subsequent iterations if the initial calculation proves insufficient.

## AVFoundation Pipeline: Asset Reading and Writing

The actual transcoding leverages AVFoundation's `AVAssetReader` and `AVAssetWriter` classes to establish a high-performance data flow between source and destination.

### Configuring AVAssetReader for Input

The encoder initializes an `AVAssetReader` to extract raw sample data from the source asset. It configures separate outputs for video and audio tracks using `AVAssetReaderTrackOutput` for video and `AVAssetReaderAudioMixOutput` for audio components, ensuring both streams are available for simultaneous processing.

### AVAssetWriter Configuration for H.264 Output

The output generation relies on an `AVAssetWriter` configured to produce MP4 files (`.mp4` extension). According to the source code in [`MediaVideoTargetEncoder.swift`](https://github.com/vorssaint/vorssaint-utils/blob/main/MediaVideoTargetEncoder.swift) (lines 78-84 and 88-121), the video settings specify H.264 codec compression with explicit width, height, and compression properties including:

- **Average bit rate** derived from the size plan
- **Expected frame rate** matching the source
- **Key frame interval** for efficient seekability

```swift
// Simplified representation of the writer configuration
let writer = AVAssetWriter(url: destination, fileType: .mp4)
let videoInput = AVAssetWriterInput(mediaType: .video, outputSettings: [
    AVVideoCodecKey: AVVideoCodecType.h264,
    AVVideoWidthKey: naturalSize.width,
    AVVideoHeightKey: naturalSize.height,
    AVVideoCompressionPropertiesKey: [
        AVVideoAverageBitRateKey: plannedBitrate,
        AVVideoExpectedSourceFrameRateKey: frameRate
    ]
])

```

## Multi-Pass Bitrate Refinement Strategy

The media tool's video compression implements an intelligent retry mechanism when initial estimates fail to meet size constraints.

### Iterative Scaling Logic

After each encoding pass completes, the system measures the actual output file size. If the result exceeds the target, `MediaSupport.targetRetryScale` calculates a new scaling factor to reduce the bitrate for the subsequent attempt. This process repeats for up to three passes, progressively tightening the compression until the file fits within the specified byte budget or the maximum attempt threshold is reached.

### Failure Handling for Impossible Targets

If the encoder cannot achieve the target size after three compression passes, the system throws `EncodeError.targetTooSmall`, indicating that the requested byte limit is infeasible for the given video duration and quality requirements.

## Audio-Video Interleaving and Synchronization

Proper synchronization between audio and video streams relies on timestamp-based sample interleaving that mirrors AVFoundation's expectations for live sources.

The encoder pulls samples from both the video and audio reader outputs, comparing presentation timestamps to determine which sample occurs earliest. It then appends that sample to the corresponding `AVAssetWriterInput`, ensuring that interleaved writes maintain precise temporal alignment throughout the output file (as implemented in lines 61-71 of [`MediaVideoTargetEncoder.swift`](https://github.com/vorssaint/vorssaint-utils/blob/main/MediaVideoTargetEncoder.swift)).

## Progress Tracking and Cancellation Support

The encoding API provides real-time feedback mechanisms essential for responsive user interfaces.

### Progress Callbacks

A user-provided `progress` closure receives updates after each processed video frame, with values capped at 0.99 until the final write completes. This allows applications to display accurate compression progress without premature completion indication.

### Cooperative Cancellation

The `isCancelled` closure enables immediate operation abortion. When this returns `true`, the encoder triggers cancellation on both the `AVAssetReader` and `AVAssetWriter`, ensuring proper resource cleanup and preventing partial file corruption.

```swift
let writtenBytes = try MediaVideoTargetEncoder.encode(
    source: source,
    trim: trim,
    targetBytes: 5 * 1_048_576, // 5 MiB target
    destination: destination,
    isCancelled: { return userPressedCancelButton },
    progress: { percent in
        progressBar.value = Float(percent)
    })

```

## Summary

- **Source Wrapping**: The `Source` struct in [`MediaVideoTargetEncoder.swift`](https://github.com/vorssaint/vorssaint-utils/blob/main/MediaVideoTargetEncoder.swift) encapsulates `AVAsset` metadata including tracks, transforms, and frame rates.
- **AVFoundation Pipeline**: Uses `AVAssetReader` with track-specific outputs and `AVAssetWriter` configured for H.264 MP4 generation.
- **Multi-Pass Adjustment**: Implements up to three encoding passes with bitrate scaling via `MediaSupport.videoSizePlan` to hit target file sizes.
- **Sync Maintenance**: Interleaves audio and video samples based on presentation timestamps to ensure proper synchronization.
- **Operational Control**: Provides progress callbacks capped at 0.99 and cooperative cancellation through the `isCancelled` closure.

## Frequently Asked Questions

### How does the media tool determine the initial bitrate for video compression?

The encoder calls `MediaSupport.videoSizePlan` to calculate a `MediaVideoSizePlan` that estimates the required bitrate based on the target byte count, video duration, and frame rate. This plan provides the initial compression settings for the first encoding pass, with subsequent passes adjusting the scale if the output remains too large.

### What happens if the video cannot be compressed to the target size?

If the file size exceeds the target after three compression attempts, the encoder throws `EncodeError.targetTooSmall`. This error indicates that the requested byte limit is technically infeasible given the source video's resolution, frame rate, and duration constraints.

### Which AVFoundation classes handle the actual reading and writing of video samples?

The implementation uses `AVAssetReader` paired with `AVAssetReaderTrackOutput` (for video) and `AVAssetReaderAudioMixOutput` (for audio) to read source data. For output, `AVAssetWriter` with `AVAssetWriterInput` handles the H.264-encoded MP4 file generation, as configured in [`MediaVideoTargetEncoder.swift`](https://github.com/vorssaint/vorssaint-utils/blob/main/MediaVideoTargetEncoder.swift).

### Can the compression process be cancelled mid-operation?

Yes. The `encode` method accepts an `isCancelled` closure that the encoder checks periodically. When this returns `true`, the method triggers cancellation on both the reader and writer instances, aborting the operation cleanly and preventing partial file writes.