Performance Implications of Using AVFoundation for Video Editing in Palmier Pro

Using AVFoundation for video editing in Palmier Pro delivers hardware-accelerated GPU rendering through Metal-backed CIContext and custom video compositing, but requires careful memory management of uncompressed BGRA pixel buffers and non-blocking async request handling to avoid pipeline stalls and thermal throttling during high-resolution timeline playback.

Palmier Pro builds its editing pipeline on top of AVFoundation, Apple’s low-level media framework that provides direct access to hardware-accelerated decoding, encoding, and rendering. While AVFoundation offloads heavy media processing to dedicated video codecs and the GPU, the implementation in CustomVideoCompositor.swift reveals specific performance constraints around memory pressure, threading saturation, and thermal limits that shape how the editor handles real-time preview and export.

Hardware Acceleration and GPU Offloading

Custom Video Compositor Implementation

At the core of Palmier Pro’s performance profile is a custom implementation of AVVideoCompositing in Sources/PalmierPro/Compositing/CustomVideoCompositor.swift. The compositor drives video import, export, and preview by processing AVAsynchronousVideoCompositionRequest instances through AVAssetExportSession and AVPlayerItem. By implementing the startRequest method, the app intercepts frame composition and redirects processing to the GPU, leveraging dedicated HEVC and H.264 hardware encoders on Apple silicon. This offloads the heavy lifting from the CPU, which is essential for maintaining real-time preview of high-resolution timelines without dropping frames.

Metal-Backed Rendering Context

The compositor creates a shared CIContext that initializes with a MTLDevice when available, ensuring all Core Image operations execute on the GPU. In CustomVideoCompositor.swift, the context configures cacheIntermediates: true to retain intermediate CI images between frames. This caching dramatically reduces per-frame CPU cost when chaining multiple effects, as the Metal-backed context avoids repeated texture uploads. However, the persistent cache increases memory footprint, requiring careful invalidation when switching between timelines.

Memory Management and Pixel Buffer Formats

BGRA Buffer Requirements

Palmier Pro insists on specific pixel buffer attributes to eliminate format conversion overhead. Both sourcePixelBufferAttributes and requiredPixelBufferAttributesForRenderContext request kCVPixelFormatType_32BGRA with kCVPixelBufferMetalCompatibilityKey set to true. This guarantees that buffers are Metal-compatible throughout the pipeline, avoiding costly CPU-side conversions that would otherwise negate GPU acceleration benefits. The single uncompressed BGRA format simplifies the compositor logic but consumes significantly more memory than compressed YUV alternatives.

VRAM Constraints on Long Timelines

The use of uncompressed BGRA buffers means that large timelines with multiple video tracks can rapidly exhaust system RAM and VRAM. Each frame at 1920×1080 resolution requires approximately 8MB of buffer space, and the compositor maintains multiple buffers simultaneously during effect chaining. Palmier Pro mitigates out-of-memory crashes by monitoring buffer pool usage in FrameRenderer.swift and limiting the number of concurrent composition requests through a lock-protected pending dictionary that de-duplicates rapid scrubbing operations.

Threading and Concurrency

DispatchQueue Configuration

The compositor queues each request on a private DispatchQueue initialized with QoS .userInteractive and the label "io.palmier.compositor". This prioritization ensures rendering work takes precedence over background tasks, delivering smooth UI updates during playback. However, the high QoS level can saturate the GPU compute queue if users rapidly scrub through the timeline, triggering multiple concurrent render requests for adjacent frames.

Request De-duplication

To prevent GPU starvation, CustomVideoCompositor.swift implements a de-duplication mechanism using a synchronized pending dictionary that tracks in-flight requests by composition time. When a new request arrives for a frame already being processed, the compositor returns the cached result rather than spawning duplicate Metal command buffers. This pattern is critical for maintaining responsive playback when users scrub through 4K footage with multiple effects applied.

Asynchronous Pipeline Architecture

Non-blocking Frame Processing

AVFoundation’s async compositing model processes AVAsynchronousVideoCompositionRequest instances in the background, allowing the export pipeline to continue while frames render. The process(_:) method in CustomVideoCompositor.swift must return quickly to avoid stalling the pipeline, as the system expects immediate completion of frame processing. Any blocking operation—such as heavy image analysis or synchronous file I/O—inside the compositor creates a bottleneck that manifests as dropped frames or frozen progress bars during export.

Export Path Inheritance

ExportService constructs an AVMutableComposition in CompositionBuilder.swift and attaches the custom AVVideoComposition that references the Metal-based compositor. Because the export path inherits the same GPU-accelerated rendering as the preview, export speeds often approach real-time playback for timelines with moderate effects. However, the final transcode phase remains bound by the encoder’s bitrate constraints and the device’s thermal envelope, which can throttle performance on sustained exports longer than ten minutes.

Core Image Effects and Kernel Performance

GPU Kernel Parallelization

Effects such as Vignette, Hue Curves, and Grain are implemented as Core Image kernels in the Kernels/*.swift files. These kernels execute on the GPU using highly parallelized Metal compute shaders, providing near-zero latency for simple color adjustments. Complex kernels that require multiple texture samples or iterative processing increase per-frame latency exponentially, particularly when chained together in CompositorInstruction.swift.

Effect Chaining Optimization

Palmier Pro balances kernel complexity by limiting the number of active effects per clip and caching intermediate results in the shared CIContext. The FrameRenderer.render method batches compatible effects into single CIImage pipeline operations where possible, reducing the number of Metal command buffer submissions. This optimization prevents the GPU command queue from becoming a bottleneck when applying heavy LUTs or temporal noise reduction filters.

Code Examples

The following patterns demonstrate how Palmier Pro configures AVFoundation for high-performance editing.

Creating a Video Composition with Custom Compositor

import AVFoundation
import CoreImage

// 1️⃣ Build a mutable composition of source tracks
let composition = AVMutableComposition()
let videoTrack = composition.addMutableTrack(
    withMediaType: .video,
    preferredTrackID: kCMPersistentTrackID_Invalid)

// … add source assets to `videoTrack` …

// 2️⃣ Create a video composition and assign the custom compositor
let videoComposition = AVMutableVideoComposition()
videoComposition.frameDuration = CMTime(value: 1, timescale: 30)   // 30 fps
videoComposition.renderSize = CGSize(width: 1920, height: 1080)
videoComposition.customVideoCompositorClass = CustomVideoCompositor.self
videoComposition.instructions = [/* array of CompositorInstruction */]

// 3️⃣ Export the composition
let exportSession = AVAssetExportSession(asset: composition,
                                         presetName: AVAssetExportPresetHEVC1920x1080)!
exportSession.videoComposition = videoComposition
exportSession.outputFileType = .mov
exportSession.outputURL = URL(fileURLWithPath: "/tmp/export.mov")
exportSession.exportAsynchronously {
    switch exportSession.status {
    case .completed: print("Export succeeded")
    case .failed:    print("Export failed: \(exportSession.error!)")
    default: break
    }
}

Inside the Custom Compositor

final class CustomVideoCompositor: NSObject, AVVideoCompositing {
    // … attributes omitted for brevity …

    private static func process(_ request: AVAsynchronousVideoCompositionRequest) {
        guard let instruction = request.videoCompositionInstruction as? CompositorInstruction else {
            // fallback: copy first source buffer
            if let first = request.sourceTrackIDs.first,
               let buffer = request.sourceFrame(byTrackID: first.int32Value) {
                request.finish(withComposedVideoFrame: buffer)
                return
            }
            request.finish(with: RenderError())
            return
        }

        guard let output = request.renderContext.newPixelBuffer() else {
            request.finish(with: RenderError())
            return
        }

        // Core Image rendering (GPU‑accelerated)
        FrameRenderer.render(
            instruction: instruction,
            sourceFrame: { request.sourceFrame(byTrackID: $0) },
            compositionTime: request.compositionTime,
            into: output,
            context: ciContext
        )
        request.finish(withComposedVideoFrame: output)
    }
}

Summary

  • GPU-first design – By insisting on Metal-compatible pixel buffers and using a shared CIContext with cacheIntermediates: true, Palmier Pro obtains high-throughput rendering, but must carefully manage GPU memory to avoid OOM crashes on long timelines.
  • Asynchronous request handling – AVFoundation’s async compositing model fits well with Swift’s concurrency model, yet developers must avoid blocking calls inside process(_:) to keep the pipeline fluid and prevent frame drops.
  • Memory-vs-speed trade-off – Storing intermediate CI results speeds up effect chaining but raises memory pressure; the codebase mitigates this by limiting concurrent request count and using BGRA buffers exclusively.
  • Thermal and encoder limits – Even with hardware acceleration, final export speed can be throttled by the device’s thermal budget and the chosen output codec bitrate, particularly for exports exceeding ten minutes.

Frequently Asked Questions

Does AVFoundation use the GPU or CPU for video editing in Palmier Pro?

AVFoundation uses the GPU for the majority of video processing in Palmier Pro. The app implements AVVideoCompositing in CustomVideoCompositor.swift to force rendering onto a Metal-backed CIContext, while hardware encoders handle HEVC and H.264 transcoding. The CPU primarily manages composition instructions and queue coordination, leaving pixel processing to dedicated video codecs and Metal compute shaders.

Why does Palmier Pro use uncompressed BGRA pixel buffers instead of compressed formats?

Palmier Pro uses kCVPixelFormatType_32BGRA with kCVPixelBufferMetalCompatibilityKey enabled to ensure buffers remain Metal-compatible throughout the pipeline without costly format conversions. While compressed YUV formats would reduce memory footprint, they would require CPU-side decompression before GPU filtering, negating the performance benefits of the Metal-based compositor. The trade-off favors speed over memory efficiency.

How does Palmier Pro prevent the video compositor from blocking the main thread?

The compositor runs on a private DispatchQueue with QoS .userInteractive rather than the main thread, processing AVAsynchronousVideoCompositionRequest instances asynchronously. The implementation includes a lock-protected pending dictionary that de-duplicates incoming requests, preventing GPU saturation and ensuring the UI remains responsive during rapid timeline scrubbing or heavy effect processing.

Why does video export slow down even with hardware acceleration?

Export performance degrades when the device hits thermal limits or when encoding at high bitrates that exceed the hardware encoder’s sustained throughput. While AVAssetExportSession leverages the same GPU-accelerated pipeline as preview, the final transcode phase must wait for the hardware encoder to flush bitstreams, which throttles under sustained load. Palmier Pro detects these conditions through ExportServiceRoundTripTests.swift to warn users when thermal throttling is likely.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →