How Jenkins Node Monitors Track Agent Health and Resources: A Deep Dive into the jenkinsci/jenkins Core

Jenkins node monitors use a pluggable extension framework where NodeMonitor subclasses asynchronously collect health metrics from agents via the Computer API, cache results in AbstractNodeMonitorDescriptor, and automatically mark agents offline when configured thresholds are breached.

Jenkins node monitors continuously evaluate agent health and resource availability to prevent builds from running on unhealthy machines. In the jenkinsci/jenkins repository, this framework relies on the NodeMonitor extension point and the AbstractNodeMonitorDescriptor base class to orchestrate periodic checks across the master-agent channel. Understanding how these components interact reveals how Jenkins maintains cluster reliability through automated health monitoring.

Core Architecture of the Monitoring Framework

The node monitoring system in core/src/main/java/hudson/node_monitors/ consists of several coordinated components that handle data collection, caching, and offline decisions.

The NodeMonitor Extension Point

The abstract class NodeMonitor defines the contract for all health monitors according to the source code. Each implementation overrides the data(Computer) method to return a monitor-specific value object representing the current metric. The class also provides triggerUpdate() to initiate asynchronous data collection and getColumnCaption() to supply the display text shown on the Manage Nodes page.

Key source: core/src/main/java/hudson/node_monitors/NodeMonitor.java

AbstractNodeMonitorDescriptor and Caching

AbstractNodeMonitorDescriptor serves as the base descriptor for all monitors and maintains a cached map of Computer to value objects. It defines the default update interval of 5 minutes and provides the monitor() method executed asynchronously to gather data from agents. This descriptor handles the background task scheduling that prevents the master from blocking during remote metric collection.

Key source: core/src/main/java/hudson/node_monitors/AbstractNodeMonitorDescriptor.java

NodeMonitorUpdater Event Listener

The NodeMonitorUpdater class implements ComputerListener to react to agent lifecycle events. When an agent connects (onOnline or onTemporarilyOnline), it schedules a bulk update via triggerUpdate() after a short quiet period, preventing a flood of simultaneous checks when many agents start together.

Key source: core/src/main/java/hudson/node_monitors/NodeMonitorUpdater.java

MonitorOfflineCause Hierarchy

When a monitor detects an unhealthy condition, it creates a MonitorOfflineCause subclass representing the specific reason (e.g., low disk space, high response time). The descriptor's markNodeOfflineOrOnline method uses these cause objects to transition agents offline or bring them back online when metrics recover.

Key source: core/src/main/java/hudson/node_monitors/MonitorOfflineCause.java

Central Registry via ComputerSet

ComputerSet.getMonitors() provides the central registry returning all live NodeMonitor instances. Both the updater framework and UI components rely on this method to access the current list of configured monitors.

Key source: core/src/main/java/hudson/model/ComputerSet.java

How Jenkins Node Monitors Track Agent Health: The Execution Flow

The monitoring process follows a distinct lifecycle from registration to UI rendering.

1. Extension Registration

Each concrete monitor is annotated with @Extension (or @Symbol) so Jenkins discovers it at startup. The corresponding descriptor registers with the system, making the monitor available via ComputerSet.getMonitors().

2. Event-Driven Scheduling

When an agent connects, NodeMonitorUpdater receives the lifecycle callback and schedules a MONITOR_UPDATER runnable. This runnable iterates over all monitors from ComputerSet.getMonitors() and invokes nm.triggerUpdate() to initiate data collection.

3. Asynchronous Data Collection

The triggerUpdate() method delegates to the monitor's descriptor, which spawns a background MonitorTask. This task runs on the agent via the master-to-slave channel to gather metrics such as:

  • Free disk space via DiskSpaceMonitorDescriptor.monitorDetailed()
  • Swap memory via MemoryMonitor.get()
  • Response time via ping operations

Results are stored in the descriptor's internal cache map.

4. Offline Decision Logic

After data collection, the descriptor calls markNodeOfflineOrOnline. It compares the metric against user-configured thresholds (e.g., freeSpaceThreshold, warningThreshold). If the metric falls below the threshold, the monitor marks the agent offline with a specific MonitorOfflineCause subclass. If the metric recovers, the agent returns to online status.

5. UI Rendering

The Manage Nodes page displays monitor columns using the caption from NodeMonitor.getColumnCaption(). The Jelly view (column.jelly) reads monitor.data(computer) to render current values such as "15 GB free" or "25ms response time".

Resource-Specific Monitor Implementations

Jenkins ships with several concrete monitors targeting specific system resources.

DiskSpaceMonitorDescriptor

Tracks: Free space of a configured directory (default $JENKINS_HOME).

Value object: DiskSpace (extends MonitorOfflineCause).

Threshold logic: AbstractDiskSpaceMonitor.getThresholdBytes() parses human-readable strings (e.g., "1GiB"), and markNodeOfflineOrOnline checks if size <= threshold. When breached, the agent is marked offline with a DiskSpace cause.

Key source: core/src/main/java/hudson/node_monitors/DiskSpaceMonitorDescriptor.java

SwapSpaceMonitor

Tracks: Available swap memory on the agent.

Value object: MemoryUsage2 (extends MemoryUsage).

Threshold logic: Unlike disk monitoring, this monitor's canTakeOffline() returns false. It displays warnings in the UI when free swap is low but does not automatically take agents offline.

Key source: core/src/main/java/hudson/node_monitors/SwapSpaceMonitor.java

ResponseTimeMonitor

Tracks: Round-trip ping latency to the agent.

Value object: ResponseTime (extends MonitorOfflineCause).

Threshold logic: Exceeding the configured response time threshold triggers an offline status via the corresponding cause object.

Key source: core/src/main/java/hudson/node_monitors/ResponseTimeMonitor.java

ClockMonitor

Tracks: Clock drift between master and agent.

Value object: Clock (extends MonitorOfflineCause).

Threshold logic: Significant clock drift triggers an offline status to prevent build timestamp inconsistencies.

Key source: core/src/main/java/hudson/node_monitors/ClockMonitor.java

ArchitectureMonitor

Tracks: CPU architecture string reported by the agent.

Value object: Architecture (extends MonitorOfflineCause).

Threshold logic: Mismatches between expected and reported architecture can mark the node offline.

Key source: core/src/main/java/hudson/node_monitors/ArchitectureMonitor.java

Accessing Node Monitor Data Programmatically

You can interact with the monitoring framework directly from Jenkins scripts or plugins.

Retrieving Current Metrics for All Agents

This example demonstrates accessing the disk space monitor data for every connected agent:

import hudson.model.Computer;
import hudson.node_monitors.NodeMonitor;
import jenkins.model.Jenkins;

// Iterate through all agents
for (Computer c : Jenkins.get().getComputers()) {
    // Retrieve Disk Space monitor data (may be null if not yet collected)
    Object diskInfo = NodeMonitor.getAll()
        .stream()
        .filter(m -> m instanceof hudson.node_monitors.DiskSpaceMonitor)
        .findFirst()
        .orElseThrow()
        .data(c);
    System.out.println(c.getDisplayName() + " – disk: " + diskInfo);
}

Source reference: NodeMonitor.getAll() → ComputerSet.getMonitors() → NodeMonitor.data(Computer)

Triggering Manual Updates

While the framework automatically updates monitors every 5 minutes, you can force a refresh programmatically:

import hudson.node_monitors.NodeMonitor;

// Force asynchronous refresh of all node monitors
for (NodeMonitor m : NodeMonitor.getAll()) {
    m.triggerUpdate();
}

Source reference: NodeMonitor.triggerUpdate() → AbstractNodeMonitorDescriptor.triggerUpdate()

Summary

Jenkins node monitors track agent health and resources through a sophisticated asynchronous framework:

  • NodeMonitor defines the extension point for health checks, with concrete implementations targeting specific resources like disk space, swap memory, and response time.
  • AbstractNodeMonitorDescriptor manages cached data and schedules asynchronous collection every 5 minutes via background tasks.
  • NodeMonitorUpdater listens for agent connection events to trigger bulk updates without overwhelming the system.
  • MonitorOfflineCause objects represent specific failure reasons, allowing automated offline/online transitions when thresholds are breached or resolved.
  • ComputerSet.getMonitors() provides the central registry for accessing all configured monitors.

Frequently Asked Questions

How often do Jenkins node monitors check agent health?

By default, AbstractNodeMonitorDescriptor schedules updates every 5 minutes. Additionally, NodeMonitorUpdater triggers immediate checks when agents connect, subject to a short quiet period to prevent flooding.

Can I customize the thresholds for taking agents offline?

Yes. Monitors like DiskSpaceMonitorDescriptor expose threshold configuration via AbstractDiskSpaceMonitor.getThresholdBytes(), which parses human-readable values such as "1GiB" or "500MiB". When the monitored metric falls below this threshold, markNodeOfflineOrOnline transitions the agent offline.

What happens if a node monitor fails to collect data?

The AbstractNodeMonitorDescriptor handles exceptions during asynchronous collection gracefully. Failed checks typically result in null or stale data in the cache, and the UI may display empty values. The agent remains online unless the monitor explicitly implements error handling that triggers an offline cause.

How do I add a custom node monitor to Jenkins?

Extend the NodeMonitor abstract class and annotate your implementation with @Extension. Override data(Computer) to return your metric value, and optionally extend AbstractNodeMonitorDescriptor to implement custom threshold logic and offline decisions. Follow the pattern established by SwapSpaceMonitor for monitors that only display warnings, or DiskSpaceMonitorDescriptor for monitors that enforce offline states.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →