# How to Integrate and Use Prometheus Metrics Exported by Dragonboat for Cluster Monitoring

> Learn to integrate Prometheus metrics from Dragonboat for robust cluster monitoring. Enable metrics in NodeHost and expose them via HTTP for Prometheus scraping. Monitor your Dragonboat cluster effectively.

- Repository: [lni/dragonboat](https://github.com/lni/dragonboat)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Set `NodeHostConfig.EnableMetrics` to `true` when creating your NodeHost, then expose the metrics via HTTP using `dragonboat.WriteHealthMetrics(w)` for Prometheus to scrape.**

Dragonboat is a high-performance Go implementation of the Raft consensus protocol. When running production clusters, you need visibility into node health, leadership state, and transport activity. The library exposes a comprehensive set of **Prometheus metrics** that track everything from election campaigns to message connection failures. This guide shows you how to enable, access, and monitor these metrics using the actual source implementation from the `lni/dragonboat` repository.

## Enabling Metrics in NodeHostConfig

Metrics collection is disabled by default to eliminate overhead in performance-critical deployments. To activate the instrumentation, you must explicitly enable it in your configuration.

In [`config/config.go`](https://github.com/lni/dragonboat/blob/main/config/config.go) (lines 49-51), the `EnableMetrics` boolean flag controls whether Dragonboat instantiates counters and gauges:

```go
// EnableMetrics determines whether health metrics in Prometheus format should be enabled.
EnableMetrics bool

```

When you create your `NodeHost`, pass this configuration:

```go
nhConfig := config.NodeHostConfig{
    WALDir:         "./wal",
    NodeHostDir:    "./nh",
    RTTMillisecond: 200,
    EnableMetrics:  true,  // Critical: enables all metric collection
}

```

Setting this flag to `true` triggers the creation of metric registries in the Raft event listener, transport layer, and LogDB components.

## Understanding the Metric Architecture

Dragonboat uses the `VictoriaMetrics/metrics` library for efficient, lock-free metric collection. The architecture splits instrumentation across three primary components, each prefixed with `dragonboat_` to ensure unique Prometheus series names.

### Raft Event Listener Metrics

When `EnableMetrics` is true, `newRaftEventListener` in [`node.go`](https://github.com/lni/dragonboat/blob/main/node.go) (lines 9-11) initializes counters for core Raft state machine events:

```go
rn.raftEvents = newRaftEventListener(config.ShardID,
    config.ReplicaID, nhConfig.EnableMetrics, liQueue)

```

These metrics track:

- **`dragonboat_raftnode_campaign_launched_total`** – Leader elections initiated
- **`dragonboat_raftnode_campaign_skipped_total`** – Elections aborted due to existing leader
- **`dragonboat_raftnode_proposal_dropped_total`** – Proposals rejected due to queue overflow
- **`dragonboat_raftnode_read_index_dropped_total`** – Read index requests dropped
- **`dragonboat_raftnode_has_leader{shardid,replicaid}`** – Gauge indicating current leadership state (1 or 0)
- **`dragonboat_raftnode_term{shardid,replicaid}`** – Gauge showing current Raft term

### Transport Layer Metrics

The transport layer in [`internal/transport/metrics.go`](https://github.com/lni/dragonboat/blob/main/internal/transport/metrics.go) exposes network health indicators. The `newTransportMetrics` function registers gauges and counters only when metrics are enabled:

```go
tm.messageConns = metrics.GetOrCreateGauge(name, msgCount)

```

Key transport metrics include:

- **`dragonboat_transport_message_connections`** – Active TCP connections for Raft messages
- **`dragonboat_transport_snapshot_connections`** – Active snapshot streaming connections
- **`dragonboat_transport_message_send_success_total`** – Successfully sent messages
- **`dragonboat_transport_message_send_failure_total`** – Failed send attempts
- **`dragonboat_transport_failed_message_connection_attempt_total`** – TCP connection failures
- **`dragonboat_transport_received_message_total`** – Inbound messages processed
- **`dragonboat_transport_received_message_dropped_total`** – Messages dropped due to processing backpressure
- **`dragonboat_transport_received_snapshot_total`** – Inbound snapshots received

### LogDB Busy State

In [`nodehost.go`](https://github.com/lni/dragonboat/blob/main/nodehost.go) (lines 12-14), Dragonboat tracks LogDB pressure via the `logDBMetrics` struct:

```go
lm.update(info.Busy)

```

This updates an internal gauge indicating whether the LogDB shard is busy, which is useful for detecting disk I/O bottlenecks. While this metric is primarily internal, it influences the `proposal_dropped` counters visible in the Raft metrics when backpressure occurs.

## Exposing Metrics via HTTP

Dragonboat does not start its own HTTP server. You must create an endpoint in your application that calls `WriteHealthMetrics` from [`event.go`](https://github.com/lni/dragonboat/blob/main/event.go) (lines 28-33):

```go
// WriteHealthMetrics writes all health metrics in Prometheus format to the specified writer.
func WriteHealthMetrics(w io.Writer) {
    metrics.WritePrometheus(w, false)
}

```

Here is a complete, minimal example that starts a NodeHost with metrics enabled and serves them on port 8000:

```go
package main

import (
	"context"
	"log"
	"net/http"

	"github.com/lni/dragonboat/v4"
	"github.com/lni/dragonboat/v4/config"
	"github.com/lni/dragonboat/v4/logger"
)

func main() {
	// Configure NodeHost with metrics enabled
	nhConfig := config.NodeHostConfig{
		WALDir:         "./wal",
		NodeHostDir:    "./nh",
		RTTMillisecond: 200,
		DeploymentID:   1,
		EnableMetrics:  true, // Critical for Prometheus export
	}

	// Initialize default logger
	logger.NewDefaultZapLogger(logger.InfoLevel)

	// Create NodeHost
	nh, err := dragonboat.NewNodeHost(nhConfig, dragonboat.DefaultNodeHostOptions())
	if err != nil {
		log.Fatalf("failed to create NodeHost: %v", err)
	}
	defer nh.Close()

	// Expose Prometheus metrics endpoint
	http.HandleFunc("/metrics", func(w http.ResponseWriter, r *http.Request) {
		dragonboat.WriteHealthMetrics(w)
	})

	go func() {
		log.Println("Prometheus metrics available at :8000/metrics")
		if err := http.ListenAndServe(":8000", nil); err != nil {
			log.Fatalf("metrics server error: %v", err)
		}
	}()

	// Keep application running
	select {}
}

```

When you run this application, Prometheus can scrape `http://localhost:8000/metrics` to receive all `dragonboat_*` series.

## Prometheus Scrape Configuration

Add the following job to your [`prometheus.yml`](https://github.com/lni/dragonboat/blob/main/prometheus.yml) to collect metrics from your Dragonboat cluster:

```yaml
scrape_configs:
  - job_name: 'dragonboat_cluster'
    static_configs:
      - targets: 
        - 'node1.example.com:8000'
        - 'node2.example.com:8000'
        - 'node3.example.com:8000'
    metrics_path: /metrics
    scheme: http
    scrape_interval: 15s

```

Replace the targets with the actual hostnames or IP addresses of your NodeHost instances. The `dragonboat_raftnode_has_leader` and `dragonboat_raftnode_term` gauges are particularly useful for alerting on split-brain scenarios or leader election storms.

## Key Source Files Reference

| File | Purpose | Location |
|------|---------|----------|
| [`config/config.go`](https://github.com/lni/dragonboat/blob/main/config/config.go) | `EnableMetrics` flag definition | `github.com/lni/dragonboat/blob/master/config/config.go#L49-L51` |
| [`node.go`](https://github.com/lni/dragonboat/blob/main/node.go) | Raft event listener creation with metrics support | `github.com/lni/dragonboat/blob/master/node.go#L9-L11` |
| [`event.go`](https://github.com/lni/dragonboat/blob/main/event.go) | Public `WriteHealthMetrics` exporter function | `github.com/lni/dragonboat/blob/master/event.go#L28-L33` |
| [`internal/transport/metrics.go`](https://github.com/lni/dragonboat/blob/main/internal/transport/metrics.go) | Transport layer counters and gauges | [`github.com/lni/dragonboat/blob/master/internal/transport/metrics.go`](https://github.com/lni/dragonboat/blob/main/github.com/lni/dragonboat/blob/master/internal/transport/metrics.go) |
| [`nodehost.go`](https://github.com/lni/dragonboat/blob/main/nodehost.go) | LogDB busy state tracking | `github.com/lni/dragonboat/blob/master/nodehost.go#L12-L14` |

## Troubleshooting Common Issues

| Symptom | Root Cause | Solution |
|---------|------------|----------|
| **No metrics appear at `/metrics`** | `EnableMetrics` is `false` (default) | Explicitly set `EnableMetrics: true` in `NodeHostConfig` before creating the NodeHost. |
| **Missing `dragonboat_` prefix** | Custom event listener or transport implementation not passing the `useMetrics` flag | Verify you are using the default constructors (`newRaftEventListener`, `newTransportMetrics`) with the metrics boolean propagated from `NodeHostConfig`. |
| **High latency on metrics endpoint** | Large number of shards creating excessive metric cardinality | Reduce scrape frequency in Prometheus or aggregate metrics per NodeHost rather than per-shard in your dashboards. |
| **Transport counters stay at zero** | Using in-process loopback transport (common in unit tests) | Deploy across actual network interfaces or simulate network failures to observe transport error counters incrementing. |

## Summary

- **Enable metrics** by setting `NodeHostConfig.EnableMetrics = true` before instantiating your NodeHost.
- **Dragonboat exposes** `dragonboat_raftnode_*` for Raft state, `dragonboat_transport_*` for network health, and internal LogDB gauges for backpressure monitoring.
- **Export metrics** by calling `dragonboat.WriteHealthMetrics(w)` in your HTTP handler; the library writes Prometheus exposition format directly.
- **Source files** defining this behavior include [`config/config.go`](https://github.com/lni/dragonboat/blob/main/config/config.go), [`node.go`](https://github.com/lni/dragonboat/blob/main/node.go), [`event.go`](https://github.com/lni/dragonboat/blob/main/event.go), and [`internal/transport/metrics.go`](https://github.com/lni/dragonboat/blob/main/internal/transport/metrics.go).

## Frequently Asked Questions

### How do I disable metrics to reduce overhead?

Set `EnableMetrics: false` (or omit the field, as `false` is the default) in your `NodeHostConfig`. When disabled, Dragonboat skips instantiation of all `Counter` and `Gauge` objects in `newRaftEventListener` and `newTransportMetrics`, resulting in zero memory or CPU overhead from metric collection.

### Can I customize the metric names or add custom labels?

The built-in metrics use hardcoded prefixes (`dragonboat_`) and labels (e.g., `shardid`, `replicaid`) defined in [`node.go`](https://github.com/lni/dragonboat/blob/main/node.go) and [`internal/transport/metrics.go`](https://github.com/lni/dragonboat/blob/main/internal/transport/metrics.go). To add custom business metrics, import the `github.com/VictoriaMetrics/metrics` package (the same library Dragonboat uses) and register your own counters alongside the built-in ones. The `WriteHealthMetrics` function exports all registered metrics, both internal and custom.

### Why are my transport metrics showing zero values despite network activity?

Transport metrics only increment for actual TCP network operations. If your cluster runs on a single host using loopback interfaces or in-memory transports (common in integration tests), the transport layer bypasses the metric instrumentation paths in [`internal/transport/metrics.go`](https://github.com/lni/dragonboat/blob/main/internal/transport/metrics.go). Deploy your nodes on separate hosts or distinct network interfaces to observe non-zero values for `dragonboat_transport_message_send_success_total` and connection failure counters.

### What is the performance impact of enabling metrics?

When `EnableMetrics` is true, Dragonboat creates `metrics.Counter` and `metrics.Gauge` objects from the VictoriaMetrics library. These use atomic operations for increments, which are generally low-overhead (nanoseconds per operation). However, high-cardinality labels (unique combinations of `shardid` and `replicaid`) can increase memory usage. If you run thousands of shards per NodeHost, monitor memory consumption or consider aggregating metrics at the Prometheus level rather than per-shard.