How Dive Resolves and Analyzes Docker Images: A Deep Dive into the wagoodman/dive Source Code
Dive resolves Docker images through a pluggable resolver architecture that supports Docker Engine, local tar archives, and Podman, then analyzes them by parsing layer manifests into file trees and calculating efficiency metrics based on duplicated files and whiteouts.
Dive is an open-source tool for exploring Docker images, layer contents, and discovering ways to shrink your image size. Understanding how Dive resolves and analyzes Docker images requires examining its modular architecture, which abstracts image retrieval behind a common interface and performs deep filesystem analysis on the resulting layer data. The source code in wagoodman/dive demonstrates a clear separation between image resolution (fetching and parsing) and image analysis (efficiency calculation and waste detection).
Image Resolution Architecture
Dive supports three primary image sources: the Docker Engine daemon, local Docker tar archives, and the Podman Engine. The resolution process is orchestrated by GetImageResolver in dive/get_image_resolver.go, which parses the image URI (e.g., docker://myimage:latest or docker-archive:///path/to.tar) and returns the appropriate resolver implementation.
The Resolver Interface
All image sources implement the Resolver interface defined in dive/image/resolver.go. This contract ensures consistent behavior across different backends:
type Resolver interface {
Name() string
Fetch(ctx context.Context, id string) (*Image, error)
Build(ctx context.Context, options []string) (*Image, error)
ContentReader
}
Fetch()retrieves an existing image by ID or path.Build()creates an image from a Dockerfile (only the Docker-engine resolver implements this).Extract()(from the embeddedContentReader) allows Dive to pull individual files from specific layers.
Docker Engine Resolution
The engineResolver (dive/image/docker/engine_resolver.go) handles live Docker daemon interactions through the following steps:
- Determine Docker host –
determineDockerHost()checks theDOCKER_HOSTenvironment variable, Docker contexts, or falls back to the OS default. - Create client – Initializes a Docker client with proper host and TLS configuration.
- Inspect and pull – If
docker inspectreports the image as not found, Dive executesdocker pullto retrieve it. - Stream the image –
dockerClient.ImageSavestreams the image as a tar archive. - Parse the archive – The tar stream feeds into
NewImageArchiveto construct an in-memoryimage.Imagestructure.
reader, err := r.fetchArchive(ctx, id) // ImageSave → tar stream
img, err := NewImageArchive(reader) // parse tar layers
return img.ToImage(id) // final Image struct
Local Archive Resolution
For offline analysis, the archiveResolver (dive/image/docker/archive_resolver.go) processes local tarball files without daemon interaction. The Fetch method opens the file directly, passes the handle to NewImageArchive, and returns the parsed image:
reader, err := os.Open(path)
img, err := NewImageArchive(reader)
return img.ToImage(path)
This approach is ideal for CI/CD pipelines or air-gapped environments where Docker daemon access is unavailable.
Building the Image Model
Once a resolver obtains the image data, NewImageArchive (located in dive/image/docker/image_archive.go) processes the tar archive to construct the internal representation. It extracts layer manifests and creates a slice of Layer objects, each building a FileTree that represents the filesystem snapshot for that layer.
The resulting image.Image struct (defined in dive/image/image.go) contains:
type Image struct {
Request string // original user input
Trees []*filetree.FileTree
Layers []*Layer
}
Each layer maintains metadata about its size and commands, while the file trees enable efficient diffing and analysis across the image history.
Image Analysis Pipeline
After resolution, Dive performs deep analysis via image.Analyze in dive/image/analysis.go. This pipeline aggregates size metrics and calculates efficiency scores by examining how files change across layers.
Efficiency Calculation
The analysis performs two primary operations:
- Size aggregation – Calculates total size and user-added size across all layers (excluding the base layer).
- File-level efficiency –
filetree.Efficiencywalks everyFileTree, tracking duplicated file paths and whiteouts (files removed in later layers). It calculates a score where minimum possible size is divided by actual discovered size, producing a sorted list of inefficient files.
Key implementation details from dive/filetree/efficiency.go:
efficiency, inefficiencies := filetree.Efficiency(img.Trees)
for i, v := range img.Layers {
sizeBytes += v.Size
if i != 0 { userSizeBytes += v.Size }
}
The final Analysis struct contains the efficiency ratio, wasted bytes count, and a detailed breakdown of duplicated files consuming unnecessary space.
Code Examples
Programmatically Resolve and Analyze a Docker Image
You can use Dive as a library to resolve images and run analysis programmatically:
package main
import (
"context"
"fmt"
"log"
"github.com/wagoodman/dive/dive"
"github.com/wagoodman/dive/dive/image"
)
func main() {
// 1️⃣ Choose an image source string (Docker-engine URI)
src := "docker://alpine:latest"
// 2️⃣ Derive the source type and raw identifier
srcType, identifier := dive.DeriveImageSource(src)
// 3️⃣ Get the appropriate resolver
resolver, err := dive.GetImageResolver(srcType)
if err != nil {
log.Fatalf("resolver error: %v", err)
}
// 4️⃣ Fetch the image (this may pull from the daemon)
img, err := resolver.Fetch(context.Background(), identifier)
if err != nil {
log.Fatalf("fetch error: %v", err)
}
// 5️⃣ Run the analysis
analysis, err := image.Analyze(context.Background(), img)
if err != nil {
log.Fatalf("analysis error: %v", err)
}
// 6️⃣ Print a summary
fmt.Printf("Image: %s\n", analysis.Image)
fmt.Printf("Total size: %.2f MB\n", float64(analysis.SizeBytes)/(1024*1024))
fmt.Printf("Efficiency score: %.2f\n", analysis.Efficiency)
fmt.Printf("Wasted bytes: %d (%.2f%% of user size)\n",
analysis.WastedBytes,
analysis.WastedUserPercent*100)
}
The resolver automatically handles image pulling if the image is missing locally, and image.Analyze returns comprehensive efficiency metrics.
Analyzing a Local Docker Tarball
For offline analysis of exported images:
src := "docker-archive:///tmp/myimage.tar"
srcType, identifier := dive.DeriveImageSource(src)
resolver, _ := dive.GetImageResolver(srcType)
img, _ := resolver.Fetch(context.Background(), identifier)
analysis, _ := image.Analyze(context.Background(), img)
fmt.Println("Inefficient files (top 5):")
for i, f := range analysis.Inefficiencies[:5] {
fmt.Printf("%d. %s – duplicated %d bytes\n",
i+1, f.Path, f.CumulativeSize)
}
The archiveResolver reads the tar directly, completely avoiding Docker daemon interaction while providing identical analysis capabilities.
CLI Usage
Behind the scenes, the Dive CLI (cmd/dive/main.go) implements the same workflow:
# Inspect a remote image (Docker engine)
dive docker://nginx:alpine
# Inspect a saved tarball
dive docker-archive:///home/user/nginx.tar
Summary
- Dive uses a resolver pattern to abstract image retrieval across Docker Engine, Podman, and local tar archives, with logic centralized in
dive/get_image_resolver.go. - The
Resolverinterface requiresFetch()andBuild()methods, allowing uniform handling of different image sources through theimage.Resolvercontract indive/image/resolver.go. - Docker Engine resolution involves host detection, client initialization, conditional pulling, and tar streaming via
ImageSave, implemented indive/image/docker/engine_resolver.go. - Archive resolution bypasses the daemon entirely by opening local tarballs directly in
dive/image/docker/archive_resolver.go. - Image construction relies on
NewImageArchiveto parse layer manifests and buildFileTreestructures representing filesystem snapshots. - Analysis calculates efficiency by detecting duplicated files and whiteouts across layers in
dive/filetree/efficiency.go, producing metrics like wasted bytes and efficiency scores.
Frequently Asked Questions
Can Dive analyze Docker images without the Docker daemon running?
Yes. While the engineResolver requires a Docker daemon to fetch images, the archiveResolver can analyze saved Docker images from local tarballs using the docker-archive:// URI scheme. Additionally, Dive supports Podman as an alternative container engine, allowing analysis on systems where Docker is not installed.
How does Dive determine image efficiency?
Dive calculates efficiency by comparing the minimum possible size of an image to its actual size. The filetree.Efficiency function in dive/filetree/efficiency.go walks all layer file trees to identify files that appear in multiple layers (duplication) and whiteouts (files that were added then removed). An efficiency score of 1.0 (100%) indicates no wasted space, while lower scores reveal optimization opportunities.
What image formats and sources does Dive support?
Dive supports three primary sources as implemented in dive/get_image_resolver.go: Docker Engine images (via docker:// URIs or bare image names), local Docker tar archives (via docker-archive:// paths), and Podman Engine images. All sources resolve to the same internal image.Image structure, ensuring consistent analysis regardless of origin.
Does Dive pull images automatically if they are not present locally?
Yes. When using the Docker Engine resolver (dive/image/docker/engine_resolver.go), Dive first attempts to inspect the image. If the daemon reports the image as not found, Dive automatically executes docker pull to retrieve it before proceeding with the analysis. This behavior ensures seamless analysis of remote images without manual pre-fetching.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →