How Dive Identifies Layers That Contribute Most to Image Size

Dive determines which layers contribute most to image size by extracting raw byte counts from the image manifest, aggregating cumulative sizes during the analysis phase, and applying a descending sort strategy to rank layers from largest to smallest.

When analyzing container images with the wagoodman/dive repository, pinpointing which layers bloat your Docker image is essential for optimization. The tool implements a deterministic three-phase pipeline that identifies size-heavy layers through metadata extraction, cumulative accounting, and strategic sorting. This process operates entirely within the Go codebase before any UI rendering occurs.

Step 1: Extracting Raw Layer Sizes from Image Metadata

The identification process begins when Dive resolves an image from Docker, OCI, or Podman sources. In dive/image/layer.go (lines 16-21), the tool defines the Layer struct that captures essential metadata for each filesystem layer.

When parsing the image manifest, Dive creates a *image.Layer for every entry and populates the Size field using the history.Size value supplied by the Docker or OCI API. For saved image archives, Dive extracts the size directly from the tar header. This raw byte count represents the uncompressed contribution of that specific layer to the overall image.

// From dive/image/layer.go
type Layer struct {
    Index  int
    Id     string
    Size   int64  // Raw size in bytes
    // ... other fields
}

Step 2: Accumulating Per-Layer Size Contributions

After extracting individual layer metadata, Dive processes the complete image stack in dive/image/analysis.go (lines 20-30). The image.Analyze function iterates over all layers to compute running totals while building the Analysis struct.

During this pass, the analyzer maintains two critical counters:

  • sizeBytes: The cumulative total of all layers
  • userSizeBytes: The cumulative size excluding the base layer

The resulting *image.Analysis object contains a slice Layers []*Layer where each element retains its individual Size value, enabling precise attribution of image bloat to specific layers.

// Conceptual flow from dive/image/analysis.go
analysis := &Analysis{
    Layers: layers,
    // Size fields populated during iteration
}

Step 3: Ranking Layers by Size and Efficiency

The final phase transforms raw size data into actionable insights using sorting strategies defined in dive/filetree/order_strategy.go (lines 45-58). The BySizeDesc implementation sorts layer nodes according to their GetSize() values, ordering them from largest to smallest contributor.

When users launch the TUI and press Ctrl-O, the viewmodel invokes ToggleSortOrder (located in cmd/dive/cli/internal/ui/v1/viewmodel/filetree.go) to switch to the BySizeDesc strategy. This immediately reorders the layer list to surface the most impactful layers at the top of the interface.

Beyond raw size, Dive also identifies wasteful layers through the efficiency analysis in dive/filetree/efficiency.go (lines 35-44). This pass walks the stacked file tree to detect duplicated or white-outed files, aggregating their cumulative size per path into the Inefficiencies slice. These entries highlight layers that waste space rather than simply being large.

// From filetree/order_strategy.go
func BySizeDesc(a, b *Node) bool {
    return a.GetSize() > b.GetSize()
}

Implementing Layer Size Analysis in Go

You can replicate Dive's layer identification logic programmatically by importing the package and following the same three-phase pattern. The example below demonstrates resolving an image, analyzing its layers, and extracting the top contributors by size.

package main

import (
    "context"
    "fmt"
    "sort"
    
    "github.com/wagoodman/dive/dive/image"
)

func main() {
    // Phase 1: Resolve the image
    resolver, _ := image.GetResolver(image.DockerResolver)
    img, _ := resolver.Resolve(context.Background(), "alpine:latest")
    
    // Phase 2: Analyze to populate layer sizes
    analysis, _ := image.Analyze(context.Background(), img)
    
    // Phase 3: Sort by size descending
    sort.SliceStable(analysis.Layers, func(i, j int) bool {
        return analysis.Layers[i].Size > analysis.Layers[j].Size
    })
    
    // Display top 3 contributors
    for i := 0; i < 3 && i < len(analysis.Layers); i++ {
        l := analysis.Layers[i]
        fmt.Printf("Layer %d – %s – %d bytes\n", l.Index, l.ShortId(), l.Size)
    }
}

To analyze wasted space rather than raw size, inspect the Inefficiencies slice from the analysis result:

// Inefficiencies sorted by cumulative wasted size (ascending)
for i := len(analysis.Inefficiencies) - 1; i >= 0; i-- {
    e := analysis.Inefficiencies[i]
    fmt.Printf("Path %s wastes %d bytes across %d layers\n",
        e.Path, e.CumulativeSize, len(e.Nodes))
}

Summary

  • Metadata Extraction: Dive captures raw layer sizes from Docker/OCI API responses or tar headers in dive/image/layer.go.
  • Cumulative Accounting: The image.Analyze function in dive/image/analysis.go aggregates individual layer sizes while tracking total and user-space byte counts.
  • Size-Based Ranking: The BySizeDesc strategy in dive/filetree/order_strategy.go sorts layers from largest to smallest for UI presentation.
  • Efficiency Detection: The efficiency analyzer in dive/filetree/efficiency.go identifies layers containing duplicated or removed files that waste storage.
  • Programmatic Access: Developers can import the dive package to resolve images, analyze layers, and sort by size using the same Go APIs that power the CLI.

Frequently Asked Questions

How does Dive calculate layer size for OCI-compliant images?

Dive extracts the layer size from the history.Size field provided by the OCI image manifest API. When analyzing exported tar archives, it reads the size directly from the tar header metadata. This value represents the uncompressed bytes added by that specific layer, not the compressed size in the registry.

What is the difference between raw layer size and wasted space in Dive?

Raw layer size represents the total bytes added by a layer's filesystem changes, recorded in the Layer.Size field. Wasted space refers to bytes consumed by files that are duplicated across multiple layers or removed (white-outed) in subsequent layers, calculated by the efficiency analyzer and stored in analysis.Inefficiencies. A layer can be small in raw size but high in wasted space if it contains redundant data.

Can I sort layers by criteria other than size in the Dive interface?

Yes. The TUI supports toggling between different sort strategies using Ctrl-O. While BySizeDesc ranks layers by byte contribution, other strategies allow sorting by file tree structure or efficiency metrics. The viewmodel in cmd/dive/cli/internal/ui/v1/viewmodel/filetree.go handles these transitions dynamically without re-analyzing the image.

Does Dive identify which specific files within a layer contribute to its size?

Yes. Within the file tree view, Dive applies the same BySizeDesc ordering strategy to individual file nodes using their GetSize() values. This allows you to drill down from a large layer into its constituent files to identify specific bloated artifacts, dependencies, or logs consuming excessive storage.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →