# How to Safely Perform Dynamic Membership Changes in a Dragonboat Cluster

> Safely add or remove nodes in a Dragonboat cluster. Learn the step by step procedure for dynamic membership changes to ensure quorum safety and maintain cluster integrity.

- Repository: [lni/dragonboat](https://github.com/lni/dragonboat)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Dragonboat treats membership changes as ordinary Raft log entries that are replicated and applied in the same way as client commands, requiring a specific sequence of adding non-voting nodes, promoting them, and removing old nodes to guarantee quorum safety.**

Dynamic membership changes in the `lni/dragonboat` Raft library allow you to add or remove nodes from a running cluster without downtime. Because Dragonboat implements the Raft protocol's single-server membership change approach, the library guarantees that updates to the cluster's replica set remain linearizable, ordered, and idempotent.

## The Membership Change Architecture

Dragonboat represents every membership update as a `pb.ConfigChange` entry in the Raft log. This design means that adding a node, removing a node, or changing node roles undergoes the same **consensus protocol** as regular client proposals. Only the leader can propose a configuration change; followers reject such requests with `ErrNotLeader`.

The internal flow follows a strict pipeline to ensure safety. In [`node.go`](https://github.com/lni/dragonboat/blob/main/node.go), the `requestConfigChange` method validates the target address (for all types except `RemoveNode`) before creating the configuration struct. The code explicitly checks for malformed addresses:

```go
if cct != pb.RemoveNode && !n.validateTarget(target) {
    return nil, ErrInvalidAddress
}

```

*Source: `node.go:32-44`*

## Step-by-Step Execution Flow

When you invoke `RequestAddNode`, `RequestAddNonVoting`, `RequestAddWitness`, or `RequestDeleteNode`, the library executes the following sequence:

### 1. Request Enqueuing and ID Assignment

The `pendingConfigChange` component assigns a unique request key and forwards the `ConfigChange` to the Raft layer. The struct carries the **change type**, **replica ID**, **Raft address**, and a caller-provided **order ID** that becomes the `ConfigChangeId` stored in the log.

*Source: `node.go:44-51`*

### 2. Raft Replication and Commit

The Raft leader appends the entry to its log and replicates it to a majority of the current quorum. The entry is considered committed only after surviving the standard Raft consensus checks. This happens within the core Raft engine implementation.

*Source: [`internal/raft/raft.go`](https://github.com/lni/dragonboat/blob/main/internal/raft/raft.go)*

### 3. State Machine Application

Once committed, `node.applyConfigChange` invokes `n.p.ApplyConfigChange(cc)`. The underlying state machine then calls `membership.handleConfigChange` in [`internal/rsm/membership.go`](https://github.com/lni/dragonboat/blob/main/internal/rsm/membership.go) to validate the change against the current membership snapshot and update the in-memory `Membership` protobuf.

*Source: `node.go:71-90`, `internal/rsm/membership.go:74-99`*

### 4. Registry Updates and Client Notification

For `Add*` changes, the node registry receives a new `(shardID, replicaID, address)` entry. For `RemoveNode`, the entry is deleted. Finally, `pendingConfigChange.apply` marks the request as applied or rejected, unblocking the caller's `RequestState` so it can check the result.

*Source: `node.go:75-85`, `node.go:94-105`*

## The Safe Replacement Procedure

To avoid temporary loss of quorum or data loss, Dragonboat recommends the **add-non-voting → promote → remove** sequence documented in [`docs/devops.md`](https://github.com/lni/dragonboat/blob/main/docs/devops.md):

1. **Add a non-voting replica** using `RequestAddNonVoting` with a fresh replica ID and address. This starts log replication without affecting the quorum requirements.
2. **Wait for catch-up** by blocking on the returned `RequestState` until the change commits and the new replica receives the log.
3. **Promote to full member** by issuing `RequestAddNode` with the same replica ID, converting the non-voting replica into a voting member.
4. **Remove the old replica** using `RequestDeleteNode` only after confirming the new member is active and the cluster maintains a healthy quorum.

The following Go example demonstrates this safe workflow:

```go
package main

import (
	"log"
	"time"

	"github.com/lni/dragonboat/v4/config"
	"github.com/lni/dragonboat/v4/nodehost"
)

func main() {
	// Create NodeHost (simplified configuration)
	nh, err := nodehost.NewNodeHost(config.NodeHostConfig{
		WALDir:         "wal-dir",
		NodeHostDir:    "nh-dir",
		RTTMillisecond: 200,
	})
	if err != nil {
		log.Fatalf("NodeHost creation failed: %v", err)
	}
	defer nh.Close()

	const shardID = 1
	
	// Get client for an existing replica (e.g., replica ID 1)
	nc, err := nh.GetNode(shardID, 1)
	if err != nil {
		log.Fatalf("cannot get client: %v", err)
	}

	// Step 1: Add non-voting replica (ID 4, address 10.0.0.4:63001)
	// Use monotonic order ID for idempotency
	orderID := uint64(time.Now().UnixNano())
	rs := nc.RequestAddNonVoting(4, "10.0.0.4:63001", orderID)
	if err := rs.Err(); err != nil {
		log.Fatalf("AddNonVoting failed: %v", err)
	}
	log.Println("Non-voting replica added")

	// Step 2: Promote to full member
	orderID = uint64(time.Now().UnixNano())
	rs = nc.RequestAddNode(4, "10.0.0.4:63001", orderID)
	if err := rs.Err(); err != nil {
		log.Fatalf("Promotion failed: %v", err)
	}
	log.Println("Replica promoted to full member")

	// Step 3: Safely remove old replica (ID 2)
	// Ensure quorum is maintained before executing
	orderID = uint64(time.Now().UnixNano())
	rs = nc.RequestDeleteNode(2, orderID)
	if err := rs.Err(); err != nil {
		log.Fatalf("Remove failed: %v", err)
	}
	log.Println("Old replica removed")

	// Step 4: Clean up persisted data
	if err := nh.RemoveNodeData(shardID, 2); err != nil {
		log.Fatalf("Failed to clean node data: %v", err)
	}
	log.Println("Removed node's data deleted")
}

```

## Idempotency and Ordering Guarantees

Dragonboat prevents duplicate or out-of-order membership changes through two mechanisms. First, the `ConfigChangeId` recorded in `Membership.ConfigChangeId` equals the Raft log index of the entry, ensuring that later changes see a consistent view of the cluster. Second, `membership.handleConfigChange` checks `upToDateCC` and `alreadyMember` flags to detect duplicate requests. If a client retries the same change using the same **order ID**, Raft treats it as a duplicate and ignores it, preventing accidental double-adds or double-removes.

## Cleaning Up Removed Node Data

After successfully removing a node from the cluster membership, you should reclaim disk space by deleting its on-disk log and snapshot data. The `NodeHost.RemoveNodeData` method is safe to call only after the node has been fully removed from the configuration via `requestRemoval`.

*Source: `nodehost.go:1309-1310`*

## Summary

- **Membership changes are Raft log entries**: Dragonboat treats `AddNode`, `RemoveNode`, and similar operations as regular consensus proposals, ensuring linearizable updates.
- **Use the safe replacement sequence**: Always add a non-voting replica first, promote it to a full member, and only then remove the old node to maintain quorum.
- **Leverage idempotency**: Supply a monotonic `orderID` (such as a timestamp) to make retries safe and prevent duplicate changes.
- **Clean up after removal**: Call `RemoveNodeData` after a node is removed to delete persistent state and free disk space.

## Frequently Asked Questions

### Why must I add a non-voting node before removing an old one?

Adding a non-voting node with `RequestAddNonVoting` allows the new replica to catch up on the log without affecting the cluster's quorum requirements. If you remove a node first, you risk losing quorum if another node fails during the transition. The recommended sequence ensures that a majority of nodes always remain available to accept writes.

### How does Dragonboat ensure membership changes are idempotent?

The library uses the caller-provided **order ID** (stored as `ConfigChangeId` in the Raft log) to detect duplicate requests. When `membership.handleConfigChange` processes a change, it compares the incoming ID against the current `ConfigChangeId` and checks `alreadyMember` flags. If the change was already applied, the request succeeds without modifying the state, making retries safe.

### What is the difference between AddNonVoting and AddWitness?

`RequestAddNonVoting` adds a replica that receives log entries and can serve read requests but does not vote in leader elections or count toward quorum. `RequestAddWitness` adds a lightweight node that votes in elections but does not store the full log—useful for establishing quorum in remote regions without the storage overhead. Both use the same underlying `pb.ConfigChange` mechanism but with different type constants.

### When is it safe to delete a removed node's persistent data?

You can safely call `NodeHost.RemoveNodeData` only after the node has been successfully removed from the cluster membership via `RequestDeleteNode` and the removal has been applied to the state machine. Attempting to delete data for an active member can cause the node to crash or corrupt the cluster state. Always verify the removal request completed by checking the `RequestState` error before cleanup.