Configuring ElectionRTT and HeartbeatRTT for Optimal Leader Election Stability in High-Latency Networks

Set ElectionRTT to at least 10–20× the measured network RTT and HeartbeatRTT to 2–4× RTT in Dragonboat’s RaftConfig, ensuring the election timeout remains significantly longer than the heartbeat interval to prevent false leader elections in high-latency networks.

Dragonboat is a high-performance Raft consensus implementation in Go. Tuning the ElectionRTT and HeartbeatRTT parameters in the configuration is critical for maintaining stable leader election when operating across wide-area networks or high-latency environments.

Understanding ElectionRTT and HeartbeatRTT in Dragonboat

In config/config.go, the RaftConfig struct defines these two critical timing parameters as unsigned integers representing multiples of the network round-trip time:

// ElectionRTT is the number of RTT (round‑trip times) that a follower
// must not receive a valid leader heartbeat before starting a new
// election.
ElectionRTT uint64

// HeartbeatRTT is the number of RTT (round‑trip times) between
// successive heartbeat messages.
HeartbeatRTT uint64

Dragonboat’s Raft engine dynamically converts these RTT counts into actual time durations by multiplying them against the measured network RTT. The leader uses HeartbeatRTT to schedule heartbeat transmissions, while followers use ElectionRTT to determine when to initiate a new election.

Optimal Configuration Strategy for High-Latency Networks

The RTT Multiplier Rule

For stable leader election across high-latency links, maintain a ratio where ElectionRTT is 2–3× larger than HeartbeatRTT, but scaled up proportionally to account for network jitter. The core logic resides in the Raft engine’s timer calculations (see raftpb/raft.go), where:


electionTimeout   = ElectionRTT * measuredRTT
heartbeatInterval = HeartbeatRTT * measuredRTT

If the network RTT is 200ms and HeartbeatRTT is set to 2, the leader transmits heartbeats every 400ms. Consequently, ElectionRTT should be set to at least 10 (2 seconds), giving the system sufficient buffer to absorb delayed packets without triggering unnecessary elections.

Network Profile Typical RTT HeartbeatRTT ElectionRTT Approximate Timeout
Local LAN 1–10ms 1 5–10 5–100ms
Regional WAN 50–100ms 2 10–15 1–1.5s
High-Latency/Intercontinental 200–300ms 2–4 15–25 3–7.5s

Always profile your actual network RTT using Dragonboat’s internal RTT measurements or external tools before finalizing these values.

Implementation Details and Source Code References

The configuration parameters are consumed during node initialization in node.go, where the RaftConfig is passed to the underlying Raft state machine. The actual timer management occurs in raftpb/raft.go, where the engine:

  1. Samples current RTT to each peer
  2. Computes effective intervals using the configured multipliers
  3. Schedules heartbeat transmissions and election timeouts accordingly

Refer to the struct definition in config/config.go (lines 38–46) when constructing your configuration to ensure type safety and valid ranges.

Practical Configuration Example

The following Go code demonstrates initializing a Dragonboat node host with RTT values tuned for a 250ms WAN environment:

package main

import (
    "github.com/lni/dragonboat/v4"
    "github.com/lni/dragonboat/v4/config"
)

func main() {
    // Configure for high-latency network (250ms RTT)
    rc := config.RaftConfig{
        NodeID:        1,
        ClusterID:     100,
        HeartbeatRTT:  2,  // ~500ms heartbeat interval
        ElectionRTT:   10, // ~2.5s election timeout
    }

    nhc := config.NodeHostConfig{
        NodeHostDir: "/var/lib/dragonboat",
        LogDir:      "/var/log/dragonboat",
        RTTMillisecond: 250, // Base RTT estimate for initialization
    }

    nh, err := dragonboat.NewNodeHost(nhc)
    if err != nil {
        panic(err)
    }

    err = nh.StartReplica(rc, nil, false)
    if err != nil {
        panic(err)
    }
}

Adjust the HeartbeatRTT and ElectionRTT values based on your observed network latency to achieve the desired balance between failure detection speed and election stability.

Summary

  • ElectionRTT defines the timeout multiplier (in RTTs) before a follower starts a new election, defined in config/config.go.
  • HeartbeatRTT controls the frequency of leader heartbeats as a multiplier of measured RTT.
  • Maintain a minimum 2:1 ratio of ElectionRTT to HeartbeatRTT, with higher multipliers (10–25) recommended for high-latency networks.
  • The Raft engine in raftpb/raft.go automatically converts these RTT counts into actual timeouts by multiplying against real-time RTT measurements.
  • Increase both values proportionally as network latency grows to prevent false leader elections while maintaining fault detection capability.

Frequently Asked Questions

What is the relationship between ElectionRTT and HeartbeatRTT?

ElectionRTT must always be significantly larger than HeartbeatRTT—typically 2 to 3 times larger at minimum. This ensures that a single lost heartbeat or network jitter does not immediately trigger a new election. In high-latency networks, increase both values proportionally while maintaining this ratio to absorb variable delay spikes.

How do I calculate the correct values for my network?

First measure your baseline RTT between nodes using ping or Dragonboat’s internal metrics. Then set HeartbeatRTT to 2–4× the RTT (for sub-second to second-level detection) and ElectionRTT to 10–20× the RTT. For a 300ms RTT, use HeartbeatRTT: 3 (900ms) and ElectionRTT: 15 (4.5s) to ensure stability.

Can I change these settings dynamically after the cluster is started?

No, these values are immutable for a running replica. The RaftConfig is locked at startup in node.go when the replica is initialized. To change RTT settings, you must stop the node, update the configuration, and restart. In a running cluster, perform a rolling restart with the new values, one node at a time, to maintain availability.

What happens if I set ElectionRTT too low in a high-latency network?

If ElectionRTT is set too low relative to the actual network RTT, followers will repeatedly time out and trigger unnecessary leader elections. This creates a "split-brain" scenario where leadership oscillates frequently, degrading cluster throughput and potentially causing log inconsistency. Monitor metrics for frequent term increments as an indicator of misconfigured timeouts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →