How to Integrate and Use Prometheus Metrics Exported by Dragonboat for Cluster Monitoring

Set NodeHostConfig.EnableMetrics to true when creating your NodeHost, then expose the metrics via HTTP using dragonboat.WriteHealthMetrics(w) for Prometheus to scrape.

Dragonboat is a high-performance Go implementation of the Raft consensus protocol. When running production clusters, you need visibility into node health, leadership state, and transport activity. The library exposes a comprehensive set of Prometheus metrics that track everything from election campaigns to message connection failures. This guide shows you how to enable, access, and monitor these metrics using the actual source implementation from the lni/dragonboat repository.

Enabling Metrics in NodeHostConfig

Metrics collection is disabled by default to eliminate overhead in performance-critical deployments. To activate the instrumentation, you must explicitly enable it in your configuration.

In config/config.go (lines 49-51), the EnableMetrics boolean flag controls whether Dragonboat instantiates counters and gauges:

// EnableMetrics determines whether health metrics in Prometheus format should be enabled.
EnableMetrics bool

When you create your NodeHost, pass this configuration:

nhConfig := config.NodeHostConfig{
    WALDir:         "./wal",
    NodeHostDir:    "./nh",
    RTTMillisecond: 200,
    EnableMetrics:  true,  // Critical: enables all metric collection
}

Setting this flag to true triggers the creation of metric registries in the Raft event listener, transport layer, and LogDB components.

Understanding the Metric Architecture

Dragonboat uses the VictoriaMetrics/metrics library for efficient, lock-free metric collection. The architecture splits instrumentation across three primary components, each prefixed with dragonboat_ to ensure unique Prometheus series names.

Raft Event Listener Metrics

When EnableMetrics is true, newRaftEventListener in node.go (lines 9-11) initializes counters for core Raft state machine events:

rn.raftEvents = newRaftEventListener(config.ShardID,
    config.ReplicaID, nhConfig.EnableMetrics, liQueue)

These metrics track:

  • dragonboat_raftnode_campaign_launched_total – Leader elections initiated
  • dragonboat_raftnode_campaign_skipped_total – Elections aborted due to existing leader
  • dragonboat_raftnode_proposal_dropped_total – Proposals rejected due to queue overflow
  • dragonboat_raftnode_read_index_dropped_total – Read index requests dropped
  • dragonboat_raftnode_has_leader{shardid,replicaid} – Gauge indicating current leadership state (1 or 0)
  • dragonboat_raftnode_term{shardid,replicaid} – Gauge showing current Raft term

Transport Layer Metrics

The transport layer in internal/transport/metrics.go exposes network health indicators. The newTransportMetrics function registers gauges and counters only when metrics are enabled:

tm.messageConns = metrics.GetOrCreateGauge(name, msgCount)

Key transport metrics include:

  • dragonboat_transport_message_connections – Active TCP connections for Raft messages
  • dragonboat_transport_snapshot_connections – Active snapshot streaming connections
  • dragonboat_transport_message_send_success_total – Successfully sent messages
  • dragonboat_transport_message_send_failure_total – Failed send attempts
  • dragonboat_transport_failed_message_connection_attempt_total – TCP connection failures
  • dragonboat_transport_received_message_total – Inbound messages processed
  • dragonboat_transport_received_message_dropped_total – Messages dropped due to processing backpressure
  • dragonboat_transport_received_snapshot_total – Inbound snapshots received

LogDB Busy State

In nodehost.go (lines 12-14), Dragonboat tracks LogDB pressure via the logDBMetrics struct:

lm.update(info.Busy)

This updates an internal gauge indicating whether the LogDB shard is busy, which is useful for detecting disk I/O bottlenecks. While this metric is primarily internal, it influences the proposal_dropped counters visible in the Raft metrics when backpressure occurs.

Exposing Metrics via HTTP

Dragonboat does not start its own HTTP server. You must create an endpoint in your application that calls WriteHealthMetrics from event.go (lines 28-33):

// WriteHealthMetrics writes all health metrics in Prometheus format to the specified writer.
func WriteHealthMetrics(w io.Writer) {
    metrics.WritePrometheus(w, false)
}

Here is a complete, minimal example that starts a NodeHost with metrics enabled and serves them on port 8000:

package main

import (
	"context"
	"log"
	"net/http"

	"github.com/lni/dragonboat/v4"
	"github.com/lni/dragonboat/v4/config"
	"github.com/lni/dragonboat/v4/logger"
)

func main() {
	// Configure NodeHost with metrics enabled
	nhConfig := config.NodeHostConfig{
		WALDir:         "./wal",
		NodeHostDir:    "./nh",
		RTTMillisecond: 200,
		DeploymentID:   1,
		EnableMetrics:  true, // Critical for Prometheus export
	}

	// Initialize default logger
	logger.NewDefaultZapLogger(logger.InfoLevel)

	// Create NodeHost
	nh, err := dragonboat.NewNodeHost(nhConfig, dragonboat.DefaultNodeHostOptions())
	if err != nil {
		log.Fatalf("failed to create NodeHost: %v", err)
	}
	defer nh.Close()

	// Expose Prometheus metrics endpoint
	http.HandleFunc("/metrics", func(w http.ResponseWriter, r *http.Request) {
		dragonboat.WriteHealthMetrics(w)
	})

	go func() {
		log.Println("Prometheus metrics available at :8000/metrics")
		if err := http.ListenAndServe(":8000", nil); err != nil {
			log.Fatalf("metrics server error: %v", err)
		}
	}()

	// Keep application running
	select {}
}

When you run this application, Prometheus can scrape http://localhost:8000/metrics to receive all dragonboat_* series.

Prometheus Scrape Configuration

Add the following job to your prometheus.yml to collect metrics from your Dragonboat cluster:

scrape_configs:
  - job_name: 'dragonboat_cluster'
    static_configs:
      - targets: 
        - 'node1.example.com:8000'
        - 'node2.example.com:8000'
        - 'node3.example.com:8000'
    metrics_path: /metrics
    scheme: http
    scrape_interval: 15s

Replace the targets with the actual hostnames or IP addresses of your NodeHost instances. The dragonboat_raftnode_has_leader and dragonboat_raftnode_term gauges are particularly useful for alerting on split-brain scenarios or leader election storms.

Key Source Files Reference

File Purpose Location
config/config.go EnableMetrics flag definition github.com/lni/dragonboat/blob/master/config/config.go#L49-L51
node.go Raft event listener creation with metrics support github.com/lni/dragonboat/blob/master/node.go#L9-L11
event.go Public WriteHealthMetrics exporter function github.com/lni/dragonboat/blob/master/event.go#L28-L33
internal/transport/metrics.go Transport layer counters and gauges github.com/lni/dragonboat/blob/master/internal/transport/metrics.go
nodehost.go LogDB busy state tracking github.com/lni/dragonboat/blob/master/nodehost.go#L12-L14

Troubleshooting Common Issues

Symptom Root Cause Solution
No metrics appear at /metrics EnableMetrics is false (default) Explicitly set EnableMetrics: true in NodeHostConfig before creating the NodeHost.
Missing dragonboat_ prefix Custom event listener or transport implementation not passing the useMetrics flag Verify you are using the default constructors (newRaftEventListener, newTransportMetrics) with the metrics boolean propagated from NodeHostConfig.
High latency on metrics endpoint Large number of shards creating excessive metric cardinality Reduce scrape frequency in Prometheus or aggregate metrics per NodeHost rather than per-shard in your dashboards.
Transport counters stay at zero Using in-process loopback transport (common in unit tests) Deploy across actual network interfaces or simulate network failures to observe transport error counters incrementing.

Summary

  • Enable metrics by setting NodeHostConfig.EnableMetrics = true before instantiating your NodeHost.
  • Dragonboat exposes dragonboat_raftnode_* for Raft state, dragonboat_transport_* for network health, and internal LogDB gauges for backpressure monitoring.
  • Export metrics by calling dragonboat.WriteHealthMetrics(w) in your HTTP handler; the library writes Prometheus exposition format directly.
  • Source files defining this behavior include config/config.go, node.go, event.go, and internal/transport/metrics.go.

Frequently Asked Questions

How do I disable metrics to reduce overhead?

Set EnableMetrics: false (or omit the field, as false is the default) in your NodeHostConfig. When disabled, Dragonboat skips instantiation of all Counter and Gauge objects in newRaftEventListener and newTransportMetrics, resulting in zero memory or CPU overhead from metric collection.

Can I customize the metric names or add custom labels?

The built-in metrics use hardcoded prefixes (dragonboat_) and labels (e.g., shardid, replicaid) defined in node.go and internal/transport/metrics.go. To add custom business metrics, import the github.com/VictoriaMetrics/metrics package (the same library Dragonboat uses) and register your own counters alongside the built-in ones. The WriteHealthMetrics function exports all registered metrics, both internal and custom.

Why are my transport metrics showing zero values despite network activity?

Transport metrics only increment for actual TCP network operations. If your cluster runs on a single host using loopback interfaces or in-memory transports (common in integration tests), the transport layer bypasses the metric instrumentation paths in internal/transport/metrics.go. Deploy your nodes on separate hosts or distinct network interfaces to observe non-zero values for dragonboat_transport_message_send_success_total and connection failure counters.

What is the performance impact of enabling metrics?

When EnableMetrics is true, Dragonboat creates metrics.Counter and metrics.Gauge objects from the VictoriaMetrics library. These use atomic operations for increments, which are generally low-overhead (nanoseconds per operation). However, high-cardinality labels (unique combinations of shardid and replicaid) can increase memory usage. If you run thousands of shards per NodeHost, monitor memory consumption or consider aggregating metrics at the Prometheus level rather than per-shard.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →