How ipsw Handles LZFSE Compression and Extracts PBZX Streams: A Technical Deep Dive

The ipsw tool detects LZFSE-compressed data by scanning for the bvx2 magic header and decompresses it using a pure-Go FSE decoder, while PBZX streams are extracted via a concurrent pipeline that inflates XZ-compressed chunks and reassembles them in order using a min-heap.

Apple's firmware update formats rely on proprietary compression schemes that standard open-source tools cannot process. The blacktop/ipsw repository provides complete Go implementations for both LZFSE compression—found in kernelcaches, IMG3 payloads, and device trees—and PBZX stream extraction—the container format used in iOS OTA updates. This article examines the detection logic, decoder architecture, and concurrent extraction pipelines implemented in the source code.

Detecting LZFSE Compression via the bvx2 Magic Header

How ipsw Identifies LZFSE Blocks

ipsw detects LZFSE compression by checking the first four bytes of a payload for the ASCII string bvx2. This check appears in pkg/img3/img3.go within the DecryptData function, where decrypted payloads are scanned before decompression.

// Detect LZFSE-compressed payload (magic "bvx2")
if len(decrypted) >= 4 && bytes.Contains(decrypted[:4], []byte("bvx2")) {
    decompressed, err := lzfse.NewDecoder(decrypted).DecodeBuffer()
    // … error handling …
}

Source: img3.go lines 294-300

The same detection pattern appears in pkg/ftab/ftab.go and pkg/devicetree/devicetree.go for other firmware components.

The Decoder Architecture in pkg/lzfse

The LZFSE decoder is implemented as a pure-Go package in pkg/lzfse. The entry point NewDecoder initializes a scratch buffer and a growing output buffer, while DecodeBuffer orchestrates the decompression.

// NewDecoder creates a decoder with a scratch buffer.
func NewDecoder(data []byte) *Decoder {
    var dst bytes.Buffer
    dst.Grow(4 * len(data))
    scratch := make([]byte, 2*binary.Size(compressedBlockHeaderV1{}))
    return &Decoder{
        src: bytes.NewReader(append(data, scratch...)),
        dst: dst,
    }
}

// DecodeBuffer runs the full decode pipeline.
func (s *Decoder) DecodeBuffer() ([]byte, error) {
    if err := s.decode(); err != nil { return nil, err }
    return s.dst.Bytes(), nil
}

Source: lzfse.go lines 27-46

The core decompression logic resides in decodeV1, decodeV2HeaderSize, and decodeLmdExecuteMatch. These functions perform the following operations:

  • Parse the block header to extract magic, raw-bytes count, and literal/match tables
  • Rebuild the FSE (Finite State Entropy) tables for literal and match streams
  • Execute sliding-window copy operations to reconstruct the original data

Key constants defining hash sizes, symbol counts, and state structures are centralized in pkg/lzfse/types.go to maintain compatibility with Apple's reference implementation.

Source: types.go lines 11-57

Extracting PBZX Streams with Concurrent XZ Decompression

Detecting PBZX Containers

ipsw identifies PBZX streams by checking the first four bytes for the ASCII string "pbzx". The detection logic resides in internal/magic/magic.go, which provides the IsPBZX and IsPBZXData helper functions.

func IsPBZX(filePath string) (bool, error) {
    f, err := os.Open(filePath)
    // … read 4 bytes …
    return IsPBZXData(f)
}

func IsPBZXData(r io.Reader) (bool, error) {
    magic := make([]byte, 4)
    _, err := io.ReadFull(r, magic)
    // …
    if string(magic) == "pbzx" { return true, nil }
    return false, nil
}

Source: magic.go lines 33-36 and 318-334

The CLI command ipsw pbzx (implemented in cmd/ipsw/cmd/pbzx.go) uses magic.IsPBZX to validate input files before invoking the extraction pipeline.

The Three-Stage Pipeline: Read, Inflate, and Ordered Write

The PBZX extraction engine is implemented in pkg/ota/pbzx as a concurrent pipeline that decompresses XZ-compressed chunks while maintaining output order. The public API is the Extract function:

func Extract(ctx context.Context, src io.Reader, dst io.Writer, numWorker int) error

Source: extract.go lines 10-21

The pipeline consists of three coordinated stages:

1. Reader Stage (read.go)

The reader parses the PBZX container header and iterates through chunks, each defined by an inflated (uncompressed) size and a deflated (compressed) size.

// Logic from read.go:
// If deflateSize < inflateSize → compressed chunk → inflateCh
// If deflateSize == inflateSize → raw chunk → writeCh

Source: read.go lines 13-70

2. Inflate Workers (inflate.go)

A configurable pool of workers (defaulting to runtime.NumCPU()) receives compressed chunks from the inflateCh channel. Each worker uses github.com/xi2/xz to decompress the XZ payload, validates the output length against the expected inflated size, and forwards the result to the writeCh.

// Worker logic from inflate.go:
// xz.NewReader(chunkData) → decompress → validate length → writeCh

Source: inflate.go lines 12-40

3. Ordered Writer (write.go)

Because XZ decompression happens concurrently, chunks may complete out of order. The writer stage uses a min-heap (_Heap) keyed by chunk index to buffer out-of-order blocks. It writes chunks to the output stream only when the next expected index arrives, guaranteeing that the final extracted file maintains the correct byte order.

// Heap-based ordering from write.go:
// out-of-order chunks buffered in _Heap
// write when idx == nextExpected

Source: write.go lines 10-49

The Extract function orchestrates these stages by creating buffered channels, launching the reader goroutine, spawning the worker pool, and running the writer. A shared context.Context enables cancellation if any stage encounters an error.

Practical Code Examples

Decompressing Raw LZFSE Data in Go

The following example demonstrates how to use ipsw's LZFSE decoder to decompress a raw buffer:

package main

import (
	"fmt"
	"os"

	"github.com/blacktop/ipsw/pkg/lzfse"
)

func main() {
	// Load a file containing raw LZFSE-compressed data.
	data, err := os.ReadFile("kernelcache.lzfse")
	if err != nil {
		panic(err)
	}

	// Initialize the decoder with a scratch buffer.
	dec := lzfse.NewDecoder(data)
	
	// Decompress the entire buffer.
	plain, err := dec.DecodeBuffer()
	if err != nil {
		panic(err)
	}

	fmt.Printf("Decompressed %d → %d bytes\n", len(data), len(plain))
	// Use plain (e.g., parse Mach-O, write to disk, etc.)
}

Key API: lzfse.NewDecoder(data).DecodeBuffer() implemented in pkg/lzfse/lzfse.go.

Extracting PBZX Streams Programmatically

This example shows how to extract a PBZX container using the concurrent pipeline:

package main

import (
	"bytes"
	"context"
	"fmt"
	"os"
	"runtime"

	"github.com/blacktop/ipsw/pkg/ota/pbzx"
	"github.com/blacktop/ipsw/internal/magic"
)

func main() {
	inPath := "AppleSoftwareUpdate.pbzx"

	// Verify the file is a valid PBZX container.
	ok, err := magic.IsPBZX(inPath)
	if err != nil || !ok {
		panic("not a PBZX file")
	}

	f, err := os.Open(inPath)
	if err != nil {
		panic(err)
	}
	defer f.Close()

	var out bytes.Buffer
	
	// Extract using all available CPU cores.
	err = pbzx.Extract(context.Background(), f, &out, runtime.NumCPU())
	if err != nil {
		panic(err)
	}

	// Write the concatenated output.
	outputPath := inPath + ".extracted"
	err = os.WriteFile(outputPath, out.Bytes(), 0o644)
	if err != nil {
		panic(err)
	}
	
	fmt.Printf("Extracted %d bytes to %s\n", out.Len(), outputPath)
}

Key API: pbzx.Extract(ctx, src, dst, workers) defined in pkg/ota/pbzx/extract.go.

Summary

  • LZFSE Detection: ipsw scans the first four bytes for the bvx2 magic string in pkg/img3/img3.go and other firmware parsers.
  • LZFSE Decoding: A pure-Go implementation in pkg/lzfse uses NewDecoder and DecodeBuffer to rebuild FSE tables and execute sliding-window matches.
  • PBZX Detection: The internal/magic/magic.go package validates the pbzx header before processing.
  • PBZX Extraction: pkg/ota/pbzx implements a concurrent three-stage pipeline (read, inflate via github.com/xi2/xz, ordered write) that preserves chunk sequence using a min-heap.

Frequently Asked Questions

How does ipsw detect LZFSE-compressed data?

ipsw detects LZFSE compression by checking the first four bytes of a payload for the ASCII string bvx2. This check appears in pkg/img3/img3.go within the DecryptData function, where decrypted payloads are scanned before decompression. The same detection logic is reused in pkg/ftab/ftab.go and pkg/devicetree/devicetree.go for other firmware components.

What is the architecture of the LZFSE decoder in ipsw?

The LZFSE decoder is implemented as a pure-Go package in pkg/lzfse. The entry point NewDecoder initializes a scratch buffer and a growing output buffer, while DecodeBuffer orchestrates the decompression. Internally, decodeV1, decodeV2HeaderSize, and decodeLmdExecuteMatch parse block headers, rebuild Finite State Entropy (FSE) tables for literal and match streams, and execute sliding-window copy operations to reconstruct the original data.

How does ipsw handle out-of-order chunks during PBZX extraction?

ipsw handles out-of-order decompression using a min-heap in the writer stage (pkg/ota/pbzx/write.go). Because XZ decompression happens concurrently across multiple workers, chunks may complete out of sequence. The writer buffers completed chunks in the heap keyed by their original index, writing them to the output stream only when the next expected index arrives. This guarantees that the final extracted file maintains the correct byte order despite parallel processing.

Can I use ipsw's LZFSE and PBZX packages in my own Go programs?

Yes, both packages are designed as reusable Go modules. For LZFSE, import github.com/blacktop/ipsw/pkg/lzfse and use lzfse.NewDecoder(data).DecodeBuffer() to decompress buffers. For PBZX, import github.com/blacktop/ipsw/pkg/ota/pbzx and call pbzx.Extract(ctx, src, dst, workers) with an io.Reader and io.Writer. The internal/magic package provides magic.IsPBZX for validation before extraction.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →