# How Dive Resolves and Analyzes Docker Images: A Deep Dive into the wagoodman/dive Source Code

> Discover how Dive resolves Docker images using its pluggable architecture and analyzes them by parsing layer manifests and calculating efficiency metrics. Learn from the wagoodman/dive source code.

- Repository: [Alex Goodman/dive](https://github.com/wagoodman/dive)
- Tags: deep-dive
- Published: 2026-03-07

---

**Dive resolves Docker images through a pluggable resolver architecture that supports Docker Engine, local tar archives, and Podman, then analyzes them by parsing layer manifests into file trees and calculating efficiency metrics based on duplicated files and whiteouts.**

Dive is an open-source tool for exploring Docker images, layer contents, and discovering ways to shrink your image size. Understanding how Dive resolves and analyzes Docker images requires examining its modular architecture, which abstracts image retrieval behind a common interface and performs deep filesystem analysis on the resulting layer data. The source code in `wagoodman/dive` demonstrates a clear separation between image resolution (fetching and parsing) and image analysis (efficiency calculation and waste detection).

## Image Resolution Architecture

Dive supports three primary image sources: the Docker Engine daemon, local Docker tar archives, and the Podman Engine. The resolution process is orchestrated by `GetImageResolver` in [`dive/get_image_resolver.go`](https://github.com/wagoodman/dive/blob/main/dive/get_image_resolver.go), which parses the image URI (e.g., `docker://myimage:latest` or `docker-archive:///path/to.tar`) and returns the appropriate resolver implementation.

### The Resolver Interface

All image sources implement the **Resolver interface** defined in [`dive/image/resolver.go`](https://github.com/wagoodman/dive/blob/main/dive/image/resolver.go). This contract ensures consistent behavior across different backends:

```go
type Resolver interface {
    Name() string
    Fetch(ctx context.Context, id string) (*Image, error)
    Build(ctx context.Context, options []string) (*Image, error)
    ContentReader
}

```

- `Fetch()` retrieves an existing image by ID or path.
- `Build()` creates an image from a Dockerfile (only the Docker-engine resolver implements this).
- `Extract()` (from the embedded `ContentReader`) allows Dive to pull individual files from specific layers.

### Docker Engine Resolution

The **engineResolver** ([`dive/image/docker/engine_resolver.go`](https://github.com/wagoodman/dive/blob/main/dive/image/docker/engine_resolver.go)) handles live Docker daemon interactions through the following steps:

1. **Determine Docker host** – `determineDockerHost()` checks the `DOCKER_HOST` environment variable, Docker contexts, or falls back to the OS default.
2. **Create client** – Initializes a Docker client with proper host and TLS configuration.
3. **Inspect and pull** – If `docker inspect` reports the image as not found, Dive executes `docker pull` to retrieve it.
4. **Stream the image** – `dockerClient.ImageSave` streams the image as a tar archive.
5. **Parse the archive** – The tar stream feeds into `NewImageArchive` to construct an in-memory `image.Image` structure.

```go
reader, err := r.fetchArchive(ctx, id) // ImageSave → tar stream
img, err := NewImageArchive(reader)   // parse tar layers
return img.ToImage(id)                // final Image struct

```

### Local Archive Resolution

For offline analysis, the **archiveResolver** ([`dive/image/docker/archive_resolver.go`](https://github.com/wagoodman/dive/blob/main/dive/image/docker/archive_resolver.go)) processes local tarball files without daemon interaction. The `Fetch` method opens the file directly, passes the handle to `NewImageArchive`, and returns the parsed image:

```go
reader, err := os.Open(path)
img, err := NewImageArchive(reader)
return img.ToImage(path)

```

This approach is ideal for CI/CD pipelines or air-gapped environments where Docker daemon access is unavailable.

## Building the Image Model

Once a resolver obtains the image data, `NewImageArchive` (located in [`dive/image/docker/image_archive.go`](https://github.com/wagoodman/dive/blob/main/dive/image/docker/image_archive.go)) processes the tar archive to construct the internal representation. It extracts layer manifests and creates a slice of `Layer` objects, each building a **FileTree** that represents the filesystem snapshot for that layer.

The resulting `image.Image` struct (defined in [`dive/image/image.go`](https://github.com/wagoodman/dive/blob/main/dive/image/image.go)) contains:

```go
type Image struct {
    Request string          // original user input
    Trees   []*filetree.FileTree
    Layers  []*Layer
}

```

Each layer maintains metadata about its size and commands, while the file trees enable efficient diffing and analysis across the image history.

## Image Analysis Pipeline

After resolution, Dive performs deep analysis via `image.Analyze` in [`dive/image/analysis.go`](https://github.com/wagoodman/dive/blob/main/dive/image/analysis.go). This pipeline aggregates size metrics and calculates efficiency scores by examining how files change across layers.

### Efficiency Calculation

The analysis performs two primary operations:

1. **Size aggregation** – Calculates total size and user-added size across all layers (excluding the base layer).
2. **File-level efficiency** – `filetree.Efficiency` walks every `FileTree`, tracking duplicated file paths and **whiteouts** (files removed in later layers). It calculates a score where *minimum possible size* is divided by *actual discovered size*, producing a sorted list of inefficient files.

Key implementation details from [`dive/filetree/efficiency.go`](https://github.com/wagoodman/dive/blob/main/dive/filetree/efficiency.go):

```go
efficiency, inefficiencies := filetree.Efficiency(img.Trees)
for i, v := range img.Layers {
    sizeBytes += v.Size
    if i != 0 { userSizeBytes += v.Size }
}

```

The final `Analysis` struct contains the efficiency ratio, wasted bytes count, and a detailed breakdown of duplicated files consuming unnecessary space.

## Code Examples

### Programmatically Resolve and Analyze a Docker Image

You can use Dive as a library to resolve images and run analysis programmatically:

```go
package main

import (
    "context"
    "fmt"
    "log"

    "github.com/wagoodman/dive/dive"
    "github.com/wagoodman/dive/dive/image"
)

func main() {
    // 1️⃣ Choose an image source string (Docker-engine URI)
    src := "docker://alpine:latest"

    // 2️⃣ Derive the source type and raw identifier
    srcType, identifier := dive.DeriveImageSource(src)

    // 3️⃣ Get the appropriate resolver
    resolver, err := dive.GetImageResolver(srcType)
    if err != nil {
        log.Fatalf("resolver error: %v", err)
    }

    // 4️⃣ Fetch the image (this may pull from the daemon)
    img, err := resolver.Fetch(context.Background(), identifier)
    if err != nil {
        log.Fatalf("fetch error: %v", err)
    }

    // 5️⃣ Run the analysis
    analysis, err := image.Analyze(context.Background(), img)
    if err != nil {
        log.Fatalf("analysis error: %v", err)
    }

    // 6️⃣ Print a summary
    fmt.Printf("Image: %s\n", analysis.Image)
    fmt.Printf("Total size: %.2f MB\n", float64(analysis.SizeBytes)/(1024*1024))
    fmt.Printf("Efficiency score: %.2f\n", analysis.Efficiency)
    fmt.Printf("Wasted bytes: %d (%.2f%% of user size)\n",
        analysis.WastedBytes,
        analysis.WastedUserPercent*100)
}

```

The resolver automatically handles image pulling if the image is missing locally, and `image.Analyze` returns comprehensive efficiency metrics.

### Analyzing a Local Docker Tarball

For offline analysis of exported images:

```go
src := "docker-archive:///tmp/myimage.tar"
srcType, identifier := dive.DeriveImageSource(src)

resolver, _ := dive.GetImageResolver(srcType)
img, _ := resolver.Fetch(context.Background(), identifier)

analysis, _ := image.Analyze(context.Background(), img)

fmt.Println("Inefficient files (top 5):")
for i, f := range analysis.Inefficiencies[:5] {
    fmt.Printf("%d. %s – duplicated %d bytes\n",
        i+1, f.Path, f.CumulativeSize)
}

```

The `archiveResolver` reads the tar directly, completely avoiding Docker daemon interaction while providing identical analysis capabilities.

### CLI Usage

Behind the scenes, the Dive CLI ([`cmd/dive/main.go`](https://github.com/wagoodman/dive/blob/main/cmd/dive/main.go)) implements the same workflow:

```bash

# Inspect a remote image (Docker engine)

dive docker://nginx:alpine

# Inspect a saved tarball

dive docker-archive:///home/user/nginx.tar

```

## Summary

- **Dive uses a resolver pattern** to abstract image retrieval across Docker Engine, Podman, and local tar archives, with logic centralized in [`dive/get_image_resolver.go`](https://github.com/wagoodman/dive/blob/main/dive/get_image_resolver.go).
- **The `Resolver` interface** requires `Fetch()` and `Build()` methods, allowing uniform handling of different image sources through the `image.Resolver` contract in [`dive/image/resolver.go`](https://github.com/wagoodman/dive/blob/main/dive/image/resolver.go).
- **Docker Engine resolution** involves host detection, client initialization, conditional pulling, and tar streaming via `ImageSave`, implemented in [`dive/image/docker/engine_resolver.go`](https://github.com/wagoodman/dive/blob/main/dive/image/docker/engine_resolver.go).
- **Archive resolution** bypasses the daemon entirely by opening local tarballs directly in [`dive/image/docker/archive_resolver.go`](https://github.com/wagoodman/dive/blob/main/dive/image/docker/archive_resolver.go).
- **Image construction** relies on `NewImageArchive` to parse layer manifests and build `FileTree` structures representing filesystem snapshots.
- **Analysis calculates efficiency** by detecting duplicated files and whiteouts across layers in [`dive/filetree/efficiency.go`](https://github.com/wagoodman/dive/blob/main/dive/filetree/efficiency.go), producing metrics like wasted bytes and efficiency scores.

## Frequently Asked Questions

### Can Dive analyze Docker images without the Docker daemon running?

Yes. While the `engineResolver` requires a Docker daemon to fetch images, the `archiveResolver` can analyze saved Docker images from local tarballs using the `docker-archive://` URI scheme. Additionally, Dive supports Podman as an alternative container engine, allowing analysis on systems where Docker is not installed.

### How does Dive determine image efficiency?

Dive calculates efficiency by comparing the *minimum possible size* of an image to its *actual size*. The `filetree.Efficiency` function in [`dive/filetree/efficiency.go`](https://github.com/wagoodman/dive/blob/main/dive/filetree/efficiency.go) walks all layer file trees to identify files that appear in multiple layers (duplication) and **whiteouts** (files that were added then removed). An efficiency score of 1.0 (100%) indicates no wasted space, while lower scores reveal optimization opportunities.

### What image formats and sources does Dive support?

Dive supports three primary sources as implemented in [`dive/get_image_resolver.go`](https://github.com/wagoodman/dive/blob/main/dive/get_image_resolver.go): Docker Engine images (via `docker://` URIs or bare image names), local Docker tar archives (via `docker-archive://` paths), and Podman Engine images. All sources resolve to the same internal `image.Image` structure, ensuring consistent analysis regardless of origin.

### Does Dive pull images automatically if they are not present locally?

Yes. When using the Docker Engine resolver ([`dive/image/docker/engine_resolver.go`](https://github.com/wagoodman/dive/blob/main/dive/image/docker/engine_resolver.go)), Dive first attempts to inspect the image. If the daemon reports the image as not found, Dive automatically executes `docker pull` to retrieve it before proceeding with the analysis. This behavior ensures seamless analysis of remote images without manual pre-fetching.