How Dive Calculates Image Efficiency Metrics and Wasted Space

Dive calculates image efficiency by comparing the cumulative size of files across all layers against their minimum discovered size, while wasted space represents the cumulative size of files that appear in exactly two layers (duplicates or whiteouts).

Dive, the popular container image exploration tool from wagoodman/dive, provides developers with critical insights into Docker image bloat. Understanding how Dive calculates image efficiency metrics and wasted space helps you optimize layer composition and reduce deployment artifacts. The analysis engine traverses layered file trees to identify duplicate files and compute precise efficiency scores.

The Efficiency Score Algorithm

Walking the Layered File Tree

Dive computes the efficiency score in [dive/filetree/efficiency.go](https://github.com/wagoodman/dive/blob/main/dive/filetree/efficiency.go) through the Efficiency function. This implementation walks every file node across all image layers using VisitDepthChildFirst, tracking two critical metrics for each unique path:

  • CumulativeSize: The total bytes occupied by that path across every layer where it appears
  • minDiscoveredSize: The smallest size observed for that path (typically the first occurrence)

Detecting Inefficient Matches

When the algorithm identifies a path that appears in exactly two layers, it flags this as an inefficient match. These represent duplicates, modifications, or whiteout files that consume extra storage without adding unique value to the final image.

The Efficiency Formula

After scanning all layers, Dive calculates the score using:

[ \text{score} = \frac{\sum\text{minDiscoveredSize}}{\sum\text{CumulativeSize}} ]

If no bytes are discovered (empty image), the score defaults to 1.0 (perfect efficiency). A score of 1.0 indicates optimal layer composition, while lower values reveal storage overhead from redundant file versions.

Calculating Wasted Space Metrics

The wasted space analysis occurs in [dive/image/analysis.go](https://github.com/wagoodman/dive/blob/main/dive/image/analysis.go) within the Analyze function.

Aggregating Wasted Bytes

When Analyze invokes filetree.Efficiency, it receives a slice of inefficiencies (the duplicate paths identified during the efficiency scan). The function sums the CumulativeSize of each inefficient match to derive WastedBytes—the total storage consumed by redundant file versions across layers.

User-Layer Percentage

Dive contextualizes waste through WastedUserPercent, calculated as the ratio of WastedBytes to the total size of user-added layers (all layers except the base image). This metric answers: "What percentage of the space I added is actually wasted?"

Key Implementation Files

File Purpose
[dive/filetree/efficiency.go](https://github.com/wagoodman/dive/blob/main/dive/filetree/efficiency.go) Implements the core algorithm that walks layered file trees, detects duplicate/white-out paths, and calculates the efficiency score.
[dive/image/analysis.go](https://github.com/wagoodman/dive/blob/main/dive/image/analysis.go) Orchestrates the analysis by calling filetree.Efficiency, aggregates layer sizes, computes wasted bytes and percentages, and returns the Analysis struct consumed by the UI and CLI.
[dive/cmd/dive/cli/internal/command/ci/rules.go](https://github.com/wagoodman/dive/blob/main/cmd/dive/cli/internal/command/ci/rules.go) Provides CI validation rules that enforce minimum efficiency or maximum wasted space thresholds.

Programmatically Accessing Efficiency Data

You can leverage Dive's analysis engine directly in your own Go applications to audit container images without invoking the TUI.

// Example: programmatically obtain efficiency and wasted space for a local image
package main

import (
	"context"
	"fmt"
	"github.com/wagoodman/dive/dive/image"
)

func main() {
	// Resolve the image (Docker, Podman, or a tar archive)
	img, err := image.Resolve(context.Background(), "myimage:latest")
	if err != nil {
		panic(err)
	}

	// Run the deep analysis
	report, err := image.Analyze(context.Background(), img)
	if err != nil {
		panic(err)
	}

	fmt.Printf("Image: %s\n", report.Image)
	fmt.Printf("Overall efficiency: %.2f%%\n", report.Efficiency*100)
	fmt.Printf("Total size: %d bytes\n", report.SizeBytes)
	fmt.Printf("User‑added size: %d bytes\n", report.UserSizeByes)
	fmt.Printf("Wasted space: %d bytes (%.2f%% of user size)\n",
		report.WastedBytes, report.WastedUserPercent*100)

	// List the top 5 most wasteful paths
	fmt.Println("\nMost wasteful files:")
	for i, data := range report.Inefficiencies {
		if i >= 5 {
			break
		}
		fmt.Printf("- %s (duplicate size: %d bytes)\n",
			data.Path, data.CumulativeSize)
	}
}

Running this snippet prints the efficiency score, total/user size, wasted space, and the worst offenders—exactly the metrics displayed by the Dive CLI.

Summary

  • Dive calculates efficiency by dividing the sum of minimum file sizes by the sum of cumulative file sizes across all layers, defaulting to 1.0 for empty images.
  • The algorithm in dive/filetree/efficiency.go identifies inefficient matches as files appearing in exactly two layers.
  • Wasted space equals the cumulative size of these duplicate files, computed in dive/image/analysis.go.
  • WastedUserPercent contextualizes waste against only user-added layers, excluding the base image.
  • You can programmatically access these metrics via the image.Analyze API in your own Go tools.

Frequently Asked Questions

What does an efficiency score of 100% mean in Dive?

An efficiency score of 1.0 (or 100%) indicates that every byte in the container image represents essential, non-duplicated content. According to the wagoodman/dive source code, this occurs when the sum of minimum discovered sizes equals the sum of cumulative sizes, meaning no files were overwritten, duplicated, or deleted across layers.

How does Dive distinguish between base image layers and user-added layers?

Dive categorizes layers based on their position in the image history. The analysis engine in dive/image/analysis.go separates the base image (typically the initial layers from the parent image) from user-added layers (subsequent RUN, COPY, and ADD instructions). This distinction allows Dive to calculate WastedUserPercent specifically against the storage you added, providing actionable optimization targets.

Can I use Dive's efficiency calculation in CI/CD pipelines?

Yes. The repository includes CI validation rules in dive/cmd/dive/cli/internal/command/ci/rules.go that enforce minimum efficiency thresholds or maximum wasted space limits. You can configure these rules in your .dive-ci file to fail builds when images exceed specified inefficiency bounds, ensuring automated quality gates for container optimization.

Why does Dive flag files appearing in exactly two layers as inefficient?

Files appearing in exactly two layers typically indicate a modification pattern: the first layer adds the file, and the second either overwrites or deletes it (whiteout). According to the implementation in dive/filetree/efficiency.go, these represent storage waste because the first version contributes to the image size without appearing in the final filesystem. Files present in only one layer are efficient, while those in more than two layers represent multi-version bloat.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →