How Dive Resolves and Analyzes Docker Images: A Deep Dive into the wagoodman/dive Source Code

Dive resolves Docker images through a pluggable resolver architecture that supports Docker Engine, local tar archives, and Podman, then analyzes them by parsing layer manifests into file trees and calculating efficiency metrics based on duplicated files and whiteouts.

Dive is an open-source tool for exploring Docker images, layer contents, and discovering ways to shrink your image size. Understanding how Dive resolves and analyzes Docker images requires examining its modular architecture, which abstracts image retrieval behind a common interface and performs deep filesystem analysis on the resulting layer data. The source code in wagoodman/dive demonstrates a clear separation between image resolution (fetching and parsing) and image analysis (efficiency calculation and waste detection).

Image Resolution Architecture

Dive supports three primary image sources: the Docker Engine daemon, local Docker tar archives, and the Podman Engine. The resolution process is orchestrated by GetImageResolver in dive/get_image_resolver.go, which parses the image URI (e.g., docker://myimage:latest or docker-archive:///path/to.tar) and returns the appropriate resolver implementation.

The Resolver Interface

All image sources implement the Resolver interface defined in dive/image/resolver.go. This contract ensures consistent behavior across different backends:

type Resolver interface {
    Name() string
    Fetch(ctx context.Context, id string) (*Image, error)
    Build(ctx context.Context, options []string) (*Image, error)
    ContentReader
}
  • Fetch() retrieves an existing image by ID or path.
  • Build() creates an image from a Dockerfile (only the Docker-engine resolver implements this).
  • Extract() (from the embedded ContentReader) allows Dive to pull individual files from specific layers.

Docker Engine Resolution

The engineResolver (dive/image/docker/engine_resolver.go) handles live Docker daemon interactions through the following steps:

  1. Determine Docker host – determineDockerHost() checks the DOCKER_HOST environment variable, Docker contexts, or falls back to the OS default.
  2. Create client – Initializes a Docker client with proper host and TLS configuration.
  3. Inspect and pull – If docker inspect reports the image as not found, Dive executes docker pull to retrieve it.
  4. Stream the image – dockerClient.ImageSave streams the image as a tar archive.
  5. Parse the archive – The tar stream feeds into NewImageArchive to construct an in-memory image.Image structure.
reader, err := r.fetchArchive(ctx, id) // ImageSave → tar stream
img, err := NewImageArchive(reader)   // parse tar layers
return img.ToImage(id)                // final Image struct

Local Archive Resolution

For offline analysis, the archiveResolver (dive/image/docker/archive_resolver.go) processes local tarball files without daemon interaction. The Fetch method opens the file directly, passes the handle to NewImageArchive, and returns the parsed image:

reader, err := os.Open(path)
img, err := NewImageArchive(reader)
return img.ToImage(path)

This approach is ideal for CI/CD pipelines or air-gapped environments where Docker daemon access is unavailable.

Building the Image Model

Once a resolver obtains the image data, NewImageArchive (located in dive/image/docker/image_archive.go) processes the tar archive to construct the internal representation. It extracts layer manifests and creates a slice of Layer objects, each building a FileTree that represents the filesystem snapshot for that layer.

The resulting image.Image struct (defined in dive/image/image.go) contains:

type Image struct {
    Request string          // original user input
    Trees   []*filetree.FileTree
    Layers  []*Layer
}

Each layer maintains metadata about its size and commands, while the file trees enable efficient diffing and analysis across the image history.

Image Analysis Pipeline

After resolution, Dive performs deep analysis via image.Analyze in dive/image/analysis.go. This pipeline aggregates size metrics and calculates efficiency scores by examining how files change across layers.

Efficiency Calculation

The analysis performs two primary operations:

  1. Size aggregation – Calculates total size and user-added size across all layers (excluding the base layer).
  2. File-level efficiency – filetree.Efficiency walks every FileTree, tracking duplicated file paths and whiteouts (files removed in later layers). It calculates a score where minimum possible size is divided by actual discovered size, producing a sorted list of inefficient files.

Key implementation details from dive/filetree/efficiency.go:

efficiency, inefficiencies := filetree.Efficiency(img.Trees)
for i, v := range img.Layers {
    sizeBytes += v.Size
    if i != 0 { userSizeBytes += v.Size }
}

The final Analysis struct contains the efficiency ratio, wasted bytes count, and a detailed breakdown of duplicated files consuming unnecessary space.

Code Examples

Programmatically Resolve and Analyze a Docker Image

You can use Dive as a library to resolve images and run analysis programmatically:

package main

import (
    "context"
    "fmt"
    "log"

    "github.com/wagoodman/dive/dive"
    "github.com/wagoodman/dive/dive/image"
)

func main() {
    // 1️⃣ Choose an image source string (Docker-engine URI)
    src := "docker://alpine:latest"

    // 2️⃣ Derive the source type and raw identifier
    srcType, identifier := dive.DeriveImageSource(src)

    // 3️⃣ Get the appropriate resolver
    resolver, err := dive.GetImageResolver(srcType)
    if err != nil {
        log.Fatalf("resolver error: %v", err)
    }

    // 4️⃣ Fetch the image (this may pull from the daemon)
    img, err := resolver.Fetch(context.Background(), identifier)
    if err != nil {
        log.Fatalf("fetch error: %v", err)
    }

    // 5️⃣ Run the analysis
    analysis, err := image.Analyze(context.Background(), img)
    if err != nil {
        log.Fatalf("analysis error: %v", err)
    }

    // 6️⃣ Print a summary
    fmt.Printf("Image: %s\n", analysis.Image)
    fmt.Printf("Total size: %.2f MB\n", float64(analysis.SizeBytes)/(1024*1024))
    fmt.Printf("Efficiency score: %.2f\n", analysis.Efficiency)
    fmt.Printf("Wasted bytes: %d (%.2f%% of user size)\n",
        analysis.WastedBytes,
        analysis.WastedUserPercent*100)
}

The resolver automatically handles image pulling if the image is missing locally, and image.Analyze returns comprehensive efficiency metrics.

Analyzing a Local Docker Tarball

For offline analysis of exported images:

src := "docker-archive:///tmp/myimage.tar"
srcType, identifier := dive.DeriveImageSource(src)

resolver, _ := dive.GetImageResolver(srcType)
img, _ := resolver.Fetch(context.Background(), identifier)

analysis, _ := image.Analyze(context.Background(), img)

fmt.Println("Inefficient files (top 5):")
for i, f := range analysis.Inefficiencies[:5] {
    fmt.Printf("%d. %s – duplicated %d bytes\n",
        i+1, f.Path, f.CumulativeSize)
}

The archiveResolver reads the tar directly, completely avoiding Docker daemon interaction while providing identical analysis capabilities.

CLI Usage

Behind the scenes, the Dive CLI (cmd/dive/main.go) implements the same workflow:


# Inspect a remote image (Docker engine)

dive docker://nginx:alpine

# Inspect a saved tarball

dive docker-archive:///home/user/nginx.tar

Summary

  • Dive uses a resolver pattern to abstract image retrieval across Docker Engine, Podman, and local tar archives, with logic centralized in dive/get_image_resolver.go.
  • The Resolver interface requires Fetch() and Build() methods, allowing uniform handling of different image sources through the image.Resolver contract in dive/image/resolver.go.
  • Docker Engine resolution involves host detection, client initialization, conditional pulling, and tar streaming via ImageSave, implemented in dive/image/docker/engine_resolver.go.
  • Archive resolution bypasses the daemon entirely by opening local tarballs directly in dive/image/docker/archive_resolver.go.
  • Image construction relies on NewImageArchive to parse layer manifests and build FileTree structures representing filesystem snapshots.
  • Analysis calculates efficiency by detecting duplicated files and whiteouts across layers in dive/filetree/efficiency.go, producing metrics like wasted bytes and efficiency scores.

Frequently Asked Questions

Can Dive analyze Docker images without the Docker daemon running?

Yes. While the engineResolver requires a Docker daemon to fetch images, the archiveResolver can analyze saved Docker images from local tarballs using the docker-archive:// URI scheme. Additionally, Dive supports Podman as an alternative container engine, allowing analysis on systems where Docker is not installed.

How does Dive determine image efficiency?

Dive calculates efficiency by comparing the minimum possible size of an image to its actual size. The filetree.Efficiency function in dive/filetree/efficiency.go walks all layer file trees to identify files that appear in multiple layers (duplication) and whiteouts (files that were added then removed). An efficiency score of 1.0 (100%) indicates no wasted space, while lower scores reveal optimization opportunities.

What image formats and sources does Dive support?

Dive supports three primary sources as implemented in dive/get_image_resolver.go: Docker Engine images (via docker:// URIs or bare image names), local Docker tar archives (via docker-archive:// paths), and Podman Engine images. All sources resolve to the same internal image.Image structure, ensuring consistent analysis regardless of origin.

Does Dive pull images automatically if they are not present locally?

Yes. When using the Docker Engine resolver (dive/image/docker/engine_resolver.go), Dive first attempts to inspect the image. If the daemon reports the image as not found, Dive automatically executes docker pull to retrieve it before proceeding with the analysis. This behavior ensures seamless analysis of remote images without manual pre-fetching.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →