# How Dive Identifies Layers That Contribute Most to Image Size

> Discover how Dive identifies image layers that contribute most to size. Learn about byte counts, cumulative analysis, and size ranking for efficient Docker image optimization.

- Repository: [Alex Goodman/dive](https://github.com/wagoodman/dive)
- Tags: internals
- Published: 2026-03-07

---

**Dive determines which layers contribute most to image size by extracting raw byte counts from the image manifest, aggregating cumulative sizes during the analysis phase, and applying a descending sort strategy to rank layers from largest to smallest.**

When analyzing container images with the wagoodman/dive repository, pinpointing which layers bloat your Docker image is essential for optimization. The tool implements a deterministic three-phase pipeline that identifies size-heavy layers through metadata extraction, cumulative accounting, and strategic sorting. This process operates entirely within the Go codebase before any UI rendering occurs.

## Step 1: Extracting Raw Layer Sizes from Image Metadata

The identification process begins when Dive resolves an image from Docker, OCI, or Podman sources. In [`dive/image/layer.go`](https://github.com/wagoodman/dive/blob/main/dive/image/layer.go) (lines 16-21), the tool defines the `Layer` struct that captures essential metadata for each filesystem layer.

When parsing the image manifest, Dive creates a `*image.Layer` for every entry and populates the `Size` field using the `history.Size` value supplied by the Docker or OCI API. For saved image archives, Dive extracts the size directly from the tar header. This raw byte count represents the uncompressed contribution of that specific layer to the overall image.

```go
// From dive/image/layer.go
type Layer struct {
    Index  int
    Id     string
    Size   int64  // Raw size in bytes
    // ... other fields
}

```

## Step 2: Accumulating Per-Layer Size Contributions

After extracting individual layer metadata, Dive processes the complete image stack in [`dive/image/analysis.go`](https://github.com/wagoodman/dive/blob/main/dive/image/analysis.go) (lines 20-30). The `image.Analyze` function iterates over all layers to compute running totals while building the `Analysis` struct.

During this pass, the analyzer maintains two critical counters:
- `sizeBytes`: The cumulative total of all layers
- `userSizeBytes`: The cumulative size excluding the base layer

The resulting `*image.Analysis` object contains a slice `Layers []*Layer` where each element retains its individual `Size` value, enabling precise attribution of image bloat to specific layers.

```go
// Conceptual flow from dive/image/analysis.go
analysis := &Analysis{
    Layers: layers,
    // Size fields populated during iteration
}

```

## Step 3: Ranking Layers by Size and Efficiency

The final phase transforms raw size data into actionable insights using sorting strategies defined in [`dive/filetree/order_strategy.go`](https://github.com/wagoodman/dive/blob/main/dive/filetree/order_strategy.go) (lines 45-58). The `BySizeDesc` implementation sorts layer nodes according to their `GetSize()` values, ordering them from largest to smallest contributor.

When users launch the TUI and press **Ctrl-O**, the viewmodel invokes `ToggleSortOrder` (located in [`cmd/dive/cli/internal/ui/v1/viewmodel/filetree.go`](https://github.com/wagoodman/dive/blob/main/cmd/dive/cli/internal/ui/v1/viewmodel/filetree.go)) to switch to the `BySizeDesc` strategy. This immediately reorders the layer list to surface the most impactful layers at the top of the interface.

Beyond raw size, Dive also identifies wasteful layers through the efficiency analysis in [`dive/filetree/efficiency.go`](https://github.com/wagoodman/dive/blob/main/dive/filetree/efficiency.go) (lines 35-44). This pass walks the stacked file tree to detect duplicated or white-outed files, aggregating their cumulative size per path into the `Inefficiencies` slice. These entries highlight layers that waste space rather than simply being large.

```go
// From filetree/order_strategy.go
func BySizeDesc(a, b *Node) bool {
    return a.GetSize() > b.GetSize()
}

```

## Implementing Layer Size Analysis in Go

You can replicate Dive's layer identification logic programmatically by importing the package and following the same three-phase pattern. The example below demonstrates resolving an image, analyzing its layers, and extracting the top contributors by size.

```go
package main

import (
    "context"
    "fmt"
    "sort"
    
    "github.com/wagoodman/dive/dive/image"
)

func main() {
    // Phase 1: Resolve the image
    resolver, _ := image.GetResolver(image.DockerResolver)
    img, _ := resolver.Resolve(context.Background(), "alpine:latest")
    
    // Phase 2: Analyze to populate layer sizes
    analysis, _ := image.Analyze(context.Background(), img)
    
    // Phase 3: Sort by size descending
    sort.SliceStable(analysis.Layers, func(i, j int) bool {
        return analysis.Layers[i].Size > analysis.Layers[j].Size
    })
    
    // Display top 3 contributors
    for i := 0; i < 3 && i < len(analysis.Layers); i++ {
        l := analysis.Layers[i]
        fmt.Printf("Layer %d – %s – %d bytes\n", l.Index, l.ShortId(), l.Size)
    }
}

```

To analyze wasted space rather than raw size, inspect the `Inefficiencies` slice from the analysis result:

```go
// Inefficiencies sorted by cumulative wasted size (ascending)
for i := len(analysis.Inefficiencies) - 1; i >= 0; i-- {
    e := analysis.Inefficiencies[i]
    fmt.Printf("Path %s wastes %d bytes across %d layers\n",
        e.Path, e.CumulativeSize, len(e.Nodes))
}

```

## Summary

- **Metadata Extraction**: Dive captures raw layer sizes from Docker/OCI API responses or tar headers in [`dive/image/layer.go`](https://github.com/wagoodman/dive/blob/main/dive/image/layer.go).
- **Cumulative Accounting**: The `image.Analyze` function in [`dive/image/analysis.go`](https://github.com/wagoodman/dive/blob/main/dive/image/analysis.go) aggregates individual layer sizes while tracking total and user-space byte counts.
- **Size-Based Ranking**: The `BySizeDesc` strategy in [`dive/filetree/order_strategy.go`](https://github.com/wagoodman/dive/blob/main/dive/filetree/order_strategy.go) sorts layers from largest to smallest for UI presentation.
- **Efficiency Detection**: The efficiency analyzer in [`dive/filetree/efficiency.go`](https://github.com/wagoodman/dive/blob/main/dive/filetree/efficiency.go) identifies layers containing duplicated or removed files that waste storage.
- **Programmatic Access**: Developers can import the dive package to resolve images, analyze layers, and sort by size using the same Go APIs that power the CLI.

## Frequently Asked Questions

### How does Dive calculate layer size for OCI-compliant images?

Dive extracts the layer size from the `history.Size` field provided by the OCI image manifest API. When analyzing exported tar archives, it reads the size directly from the tar header metadata. This value represents the uncompressed bytes added by that specific layer, not the compressed size in the registry.

### What is the difference between raw layer size and wasted space in Dive?

Raw layer size represents the total bytes added by a layer's filesystem changes, recorded in the `Layer.Size` field. Wasted space refers to bytes consumed by files that are duplicated across multiple layers or removed (white-outed) in subsequent layers, calculated by the efficiency analyzer and stored in `analysis.Inefficiencies`. A layer can be small in raw size but high in wasted space if it contains redundant data.

### Can I sort layers by criteria other than size in the Dive interface?

Yes. The TUI supports toggling between different sort strategies using **Ctrl-O**. While `BySizeDesc` ranks layers by byte contribution, other strategies allow sorting by file tree structure or efficiency metrics. The viewmodel in [`cmd/dive/cli/internal/ui/v1/viewmodel/filetree.go`](https://github.com/wagoodman/dive/blob/main/cmd/dive/cli/internal/ui/v1/viewmodel/filetree.go) handles these transitions dynamically without re-analyzing the image.

### Does Dive identify which specific files within a layer contribute to its size?

Yes. Within the file tree view, Dive applies the same `BySizeDesc` ordering strategy to individual file nodes using their `GetSize()` values. This allows you to drill down from a large layer into its constituent files to identify specific bloated artifacts, dependencies, or logs consuming excessive storage.