How Jenkins Implements Build Fingerprinting and Artifact Tracking: Core Source Code Analysis

Jenkins implements build fingerprinting and artifact tracking by computing MD5 hashes of files and storing immutable Fingerprint records that map file origins to every downstream build that consumes them, enabling full artifact provenance tracing through the Fingerprint, Fingerprinter, and ArtifactArchiver components.

Jenkins build fingerprinting and artifact tracking provides a robust mechanism for tracing file origins across distributed pipelines. The jenkinsci/jenkins repository implements this through three tightly integrated core components that persist file hashes and build usage relationships. Understanding these internals helps administrators optimize storage and developers leverage the fingerprinting API for custom plugins.

Core Architecture: The Three Pillars of Fingerprinting

The fingerprinting system rests on three primary classes that handle distinct responsibilities from hash computation to persistent storage.

Fingerprint Model

The Fingerprint class in hudson/model/Fingerprint.java serves as the immutable data model for artifact tracking. Each instance stores the MD5 hash, original build reference (BuildPtr), file name, and a usage map tracking every subsequent build that references the file. The constructor new Fingerprint(build, fileName, md5) captures provenance metadata and immediately triggers persistence via save().

Fingerprinter Build Step

The Fingerprinter class in hudson/tasks/Fingerprinter.java functions as a build step that scans the workspace for configured patterns. It computes MD5 hashes using Util.toHexString(md5sum), builds a Map<String,String> mapping relative file paths to hashes, and attaches a Fingerprinter.FingerprintAction to the build record. This action bridges the build execution to the persistent fingerprint storage.

ArtifactArchiver Integration

The ArtifactArchiver class in hudson/tasks/ArtifactArchiver.java handles the publisher role that copies files to the build’s artifact directory. When setFingerprint(true) is invoked—or when the pipeline specifies fingerprint: true—the archiver delegates to Fingerprinter to record hashes after copying files to run.getArtifactsDir().

The Fingerprint Lifecycle

Understanding the end-to-end flow from hash computation to usage tracking reveals how Jenkins maintains artifact provenance across builds.

Creation and Hash Computation

During execution, Fingerprinter performs an Ant-style FileSet scan of the workspace. For each matched file, it calculates the MD5 hash and constructs a map where keys represent relative file paths and values contain the hexadecimal hash string. This map feeds into FingerprintAction creation, which instantiates Fingerprint objects for each unique file.

Persistence and Storage Abstraction

The Fingerprint.save() method delegates to the configured FingerprintStorage implementation (default: FileFingerprintStorage). The default storage writes XML representations to $JENKINS_HOME/fingerprints/, organized by the hash value. This abstraction allows administrators to plugin alternative backends like S3 or database storage without modifying the core fingerprint logic.

Usage Tracking and Provenance

Every time a subsequent build records the same file hash, Fingerprint.addFor(build) (or add(jobFullName, buildNumber)) updates the internal usages map—a Hashtable<String,RangeSet> that records which job builds have consumed the artifact. This creates a bidirectional relationship between the original build and all downstream consumers.

Query and Permission Checks

When rendering the UI or responding to API calls, Fingerprint._getUsages() filters the usage map based on the current user’s permissions. The method checks Item.DISCOVER and Item.READ permissions via canDiscoverItem before returning each Fingerprint.RangeItem, ensuring users cannot infer the existence of jobs they cannot access.

Artifact Archiving Workflow

The artifact archiving process operates independently but can integrate fingerprinting to establish file lineage.

File Selection and Copying

ArtifactArchiver receives artifact patterns (including optional exclusion patterns) and uses Ant-style globbing to locate files in the workspace. The selected files are copied to the build-specific artifacts directory via the ArtifactManager (typically StandardArtifactManager), making them available at URLs like job/.../lastSuccessfulBuild/artifact/....

Optional Fingerprinting Integration

When fingerprinting is enabled for archived artifacts, the archiver creates an internal Fingerprinter instance after copying files. This ensures that artifacts stored in the build archive directory also have corresponding fingerprint records, enabling downstream builds to trace the origin of consumed artifacts even when those artifacts are later deleted from the build directory.

Dependency Graph Integration

When Fingerprinter.enableFingerprintsInDependencyGraph is set to true, the system registers synthetic dependencies between builds. Each fingerprint creates a DependencyGraph.Dependency edge from the original build (the producer) to any later build that references the same hash (the consumer). This visualization appears in Jenkins’ Dependency Graph view, showing artifact flow across the CI/CD pipeline.

Security Considerations

Both Fingerprint and ArtifactArchiver enforce Jenkins’ permission model. Access to fingerprint metadata requires Item.DISCOVER and Item.READ permissions. The canDiscoverItem check prevents unauthorized users from enumerating job names or build numbers through fingerprint usage queries, effectively blocking information disclosure attacks that could reveal private project structures.

Practical Implementation Examples

Declarative Pipeline with Fingerprinting

Archive all JAR files and automatically record fingerprints for provenance tracking:

pipeline {
    agent any
    stages {
        stage('Build') {
            steps {
                sh 'mvn package'
                archiveArtifacts artifacts: '**/target/*.jar',
                                 fingerprint: true,
                                 allowEmptyArchive: false
            }
        }
    }
}

Freestyle Job Configuration

In the Jenkins UI for a freestyle project:

  1. Add the "Archive the artifacts" post-build action
  2. Set the pattern to **/target/*.jar
  3. Enable the "Fingerprint all archived artifacts" checkbox
  4. Alternatively, add the "Record fingerprints of files" build step to fingerprint files without archiving them

Programmatic Fingerprinting in a Plugin

Compute and record fingerprints from within a custom plugin:

// Inside a plugin's perform() method
FilePath workspace = ...;
Run<?,?> build = ...;
Map<String,String> record = new HashMap<>();

// Compute fingerprints for JAR files
new Fingerprinter("**/*.jar").record(build, workspace, listener, record,
                                    listener.getLogger());

// Attach the fingerprint action to the build
build.addAction(new Fingerprinter.FingerprintAction(build, record));

Reading Fingerprint Usage Data

Query fingerprint provenance from plugin code:

Fingerprint fp = Fingerprint.load(hash);
if (fp != null) {
    List<Fingerprint.RangeItem> usages = fp._getUsages(); // Permission-filtered
    for (Fingerprint.RangeItem u : usages) {
        System.out.println(u.name + " used in builds: " + u.ranges);
    }
}

Summary

  • Jenkins implements build fingerprinting and artifact tracking through three core components: the immutable Fingerprint model, the Fingerprinter build step, and the ArtifactArchiver publisher.
  • Fingerprint persistence uses a pluggable FingerprintStorage abstraction, with FileFingerprintStorage writing XML to $JENKINS_HOME/fingerprints/ by default.
  • Usage tracking maintains a Hashtable<String,RangeSet> mapping jobs to build numbers that consumed each file, enabling full provenance queries via Fingerprint._getUsages().
  • Security gates ensure fingerprint metadata respects Item.DISCOVER and Item.READ permissions, preventing unauthorized job enumeration.
  • Dependency graph integration creates synthetic edges between artifact producers and consumers when enableFingerprintsInDependencyGraph is active.

Frequently Asked Questions

What is the difference between artifact archiving and fingerprinting in Jenkins?

Artifact archiving copies files from the workspace to the build’s permanent artifact directory (run.getArtifactsDir()), making them accessible via the UI and API. Fingerprinting computes an MD5 hash of files and stores a persistent record (Fingerprint) tracking which builds produced and consumed those files. While archiving preserves the file content, fingerprinting preserves provenance metadata, allowing builds to trace artifact origins even after the original files are deleted.

Where does Jenkins store fingerprint data?

By default, Jenkins stores fingerprint data as XML files in $JENKINS_HOME/fingerprints/, managed by FileFingerprintStorage. The storage location is abstracted through the FingerprintStorage interface, allowing administrators to configure alternative backends such as databases or cloud storage. Each fingerprint file is named according to its MD5 hash, creating a content-addressed storage system.

How can I query fingerprint usage programmatically in a Jenkins plugin?

Use Fingerprint.load(hash) to retrieve a fingerprint record by its MD5 hash, then call fp._getUsages() to obtain a list of Fingerprint.RangeItem objects containing job names and build ranges. Note that _getUsages() automatically filters results based on the current user’s Item.DISCOVER and Item.READ permissions, returning only entries the user is authorized to see.

What hashing algorithm does Jenkins use for fingerprinting?

Jenkins uses MD5 for fingerprinting, computing hashes via standard Java MessageDigest utilities and converting them to hexadecimal strings using Util.toHexString(md5sum). While MD5 is considered cryptographically weak for security purposes, it remains sufficient for Jenkins’ content-addressing and deduplication requirements, providing a 128-bit identifier for artifact tracking.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →