How Dive Extracts Layer Contents: A Deep Dive into the File Extraction Pipeline

Dive extracts layer contents by streaming the image tarball from the container engine, locating the specific layer blob, and writing the requested files to the host filesystem through a pipeline involving UI listeners, controllers, and resolver implementations.

Extracting layer contents in Dive is a core functionality that allows developers to inspect and recover files from specific image layers without running the container. The wagoodman/dive repository implements this through a sophisticated pipeline that abstracts different container engines while providing a consistent interface for file retrieval.

The Layer Extraction Pipeline in Dive

The process for extracting layer contents in Dive follows a seven-step pipeline that bridges the terminal UI with low-level container engine operations. Each stage is decoupled through interfaces, allowing the tool to support Docker, Podman, and archive-based images.

UI Event Handling and Controller Delegation

The extraction journey begins in the file tree view. When a user initiates an extract action (typically by pressing Ctrl+e), the UI listener registered in cmd/dive/cli/internal/ui/v1/view/filetree.go captures the event via AddViewExtractListener.

The controller at cmd/dive/cli/internal/ui/v1/app/controller.go receives this signal through its onFileTreeViewExtract method. This method extracts the target path from the UI state and delegates the actual work to the configured content extractor, ensuring the UI remains responsive during the potentially long-running extraction process.

Resolver Interface Abstraction

Dive abstracts container engine differences through the image.Resolver interface defined in dive/image/resolver.go. This interface declares the generic Extract method that all engine implementations must satisfy:

type Resolver interface {
    Extract(ctx context.Context, image string, layerID string, destination string) error
    // ... other methods
}

When the controller calls Extract, it passes the image reference, specific layer identifier, and destination path. The resolver implementation handles the engine-specific mechanics of locating and retrieving the layer contents.

Docker-Engine Implementation

The Docker-engine resolver at dive/image/docker/engine_resolver.go implements the extraction logic for running Docker daemons. Its Extract method follows a two-phase approach:

First, it pulls the raw image tarball from the Docker daemon using ImageSave, which streams the entire image archive across the Docker socket. Second, it delegates to image/docker.ExtractFromImage, passing the tar stream, layer identifier, and destination path.

The fetchArchive helper manages the temporary storage of this stream, ensuring efficient processing without exhausting memory on large images. This implementation handles the complexity of Docker's layer storage format, converting the engine-specific representation into a standard tar stream for further processing.

Tar Stream Processing and File Extraction

The final stage occurs in dive/image/docker/image_archive.go, where ExtractFromImage processes the tar stream. This function scans top-level entries until it locates the file matching the requested layer name. Upon finding the layer blob, it delegates to extractInner.

The extractInner function walks the layer's tar contents, creating necessary directories on the host filesystem and writing each regular file to the specified destination. This low-level tar handling ensures that file permissions, timestamps, and directory structures are preserved during extraction, providing an accurate representation of the layer contents as they would appear in a running container.

Extracting Files via the Dive UI

For interactive use, extracting layer contents in Dive requires no programming knowledge. The tool provides a keyboard-driven workflow:

  1. Launch Dive with your target image: dive nginx:latest
  2. Navigate to the desired layer using arrow keys or j/k
  3. Move the cursor in the file tree to the specific file or directory
  4. Press Ctrl+e and enter the host path where content should be written
  5. Dive extracts the selected files and displays a confirmation status

This UI workflow triggers the same pipeline described above, abstracting the complexity of tar stream handling and engine communication behind a simple key binding.

Programmatic Layer Extraction

Developers integrating Dive's functionality into their own tools can use the resolver interface directly. The following Go example demonstrates extracting a specific file from layer 3 of an image:

package main

import (
    "context"
    "log"
    
    "github.com/wagoodman/dive/dive/image"
    "github.com/wagoodman/dive/dive/image/docker"
)

func extractFromLayer() error {
    // Initialize Docker-engine resolver
    resolver := docker.NewResolverFromEngine()
    
    // Fetch image metadata
    ctx := context.Background()
    img, err := resolver.Fetch(ctx, "nginx:latest")
    if err != nil {
        return err
    }
    
    // Extract /etc/nginx/nginx.conf from layer index 3
    targetLayer := img.Layers[3].Id
    destination := "/tmp/nginx.conf"
    
    return resolver.Extract(ctx, img.Request, targetLayer, destination)
}

func main() {
    if err := extractFromLayer(); err != nil {
        log.Fatal(err)
    }
}

This approach leverages the same Extract method used by the CLI, ensuring consistent behavior across Docker and Podman backends.

Summary

Extracting layer contents in Dive involves a sophisticated pipeline that bridges user interactions with low-level container engine operations:

  • The UI layer captures extraction requests via AddViewExtractListener in filetree.go and delegates to the controller
  • The controller orchestrates the operation through onFileTreeViewExtract, handling user input and status updates
  • The resolver interface abstracts engine differences, with implementations for Docker, Podman, and archives
  • The Docker resolver streams the image tarball via ImageSave, then processes it through ExtractFromImage and extractInner to write files to the host filesystem

This architecture ensures that extracting layer contents works consistently across different container runtimes while maintaining clean separation between the terminal UI and backend operations.

Frequently Asked Questions

How do I extract a file from a specific layer in Dive?

Navigate to the desired layer in the Dive UI using the arrow keys or j/k, move the cursor to the file in the tree view, and press Ctrl+e. Enter the destination path on your host machine, and Dive will write the file there using the extraction pipeline defined in dive/image/docker/image_archive.go.

What container engines does Dive support for file extraction?

Dive supports Docker and Podman for full file extraction, with both implementing the image.Resolver interface defined in dive/image/resolver.go. The Docker-archive resolver (archive_resolver.go) currently returns a "not implemented" error for extraction operations, limiting extraction to live engine connections.

Why does Dive need to pull a tarball to extract a single file?

According to the implementation in dive/image/docker/engine_resolver.go, Dive must call ImageSave to stream the entire image archive from the Docker daemon because Docker's API does not provide a direct endpoint to extract individual files from specific layers. The ExtractFromImage function then scans this stream to locate and extract the requested layer contents.

Can I extract files from Docker archive files using Dive?

No, the Docker-archive resolver at dive/image/docker/archive_resolver.go currently returns a "not implemented" error for the Extract method. To extract layer contents, you must use Dive with a live Docker or Podman engine connection rather than static archive files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →