# How r-nacos Uses the Raft Consensus Protocol for Distributed Data Storage in Cluster Mode

> Discover how r-nacos leverages the Raft consensus protocol for strongly consistent distributed data storage. Learn about its log replication, persistent storage, and gRPC transport for cluster synchronization. Read now!

- Repository: [Nacos Group/r-nacos](https://github.com/nacos-group/r-nacos)
- Tags: deep-dive
- Published: 2026-03-07

---

**r-nacos embeds the async-raft-ext crate to provide strongly consistent distributed data storage in cluster mode, using Raft log replication, persistent file-based storage, and gRPC-based RPC transport to keep configuration and naming metadata synchronized across all nodes.**

The r-nacos project implements a high-performance, Rust-based Nacos server that achieves fault-tolerant distributed state through the Raft consensus protocol. When operating in cluster mode, r-nacos leverages the **async-raft-ext** crate to replicate configuration, service naming, and metadata changes across multiple nodes, ensuring linearizable consistency and automatic failover.

## Raft Instance Construction and Configuration

When the service starts, the `build_raft` function in [`src/starter.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/starter.rs) constructs the Raft components using the `NacosRaft` type alias defined in [`src/raft/mod.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/mod.rs):

```rust
// src/starter.rs
let raft = build_raft(&sys_config, store.clone(), cluster_sender.clone()).await?;

```

The `build_raft` function configures the Raft instance with specific timing parameters and snapshot policies:

```rust
// src/starter.rs (lines 94-124)
let config = Config::build("rnacos raft".to_owned())
    .heartbeat_interval(1000)               // 1 second heartbeats
    .election_timeout_min(2500)
    .election_timeout_max(5000)
    .snapshot_policy(async_raft_ext::SnapshotPolicy::LogsSinceLast(
        sys_config.raft_snapshot_log_size,
    ))
    .snapshot_max_chunk_size(3 * 1024 * 1024)  // 3MB chunks
    .validate()
    .unwrap();

let network = Arc::new(RaftRouter::new(store.clone(), cluster_sender.clone()));
let raft = Arc::new(Raft::new(
    sys_config.raft_node_id,
    config,
    network,
    store.clone(),
));

```

These configuration parameters ensure the cluster remains responsive while preventing split-brain scenarios through configurable election timeouts.

## Persistent Storage and the State Machine

The `RaftStore` struct in [`src/raft/store/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/store/core.rs) implements the `RaftStorage` trait, providing durable persistence for the Raft log, hard state, and snapshots. It delegates actual I/O operations to the `FileStore` actors:

```rust
// src/raft/store/core.rs (excerpt)
#[async_trait]
impl RaftStorage<ClientRequest, ClientResponse> for RaftStore {
    async fn get_initial_state(&self) -> anyhow::Result<InitialState> {
        match self.send_store_msg(StoreRequest::GetInitialState).await? {
            StoreResponse::InitialState(v) => Ok(v),
            _ => Err(self.store_response_err()),
        }
    }
    
    // Additional methods: get_log_entries, append_entry_to_log, 
    // apply_entry_to_state_machine, do_log_compaction, etc.
}

```

The state machine processes application-level requests defined in [`src/raft/store/mod.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/store/mod.rs) through the `ClientRequest` enum, which includes variants such as `NodeAddr`, `Members`, `ConfigSet`, `NamingReq`, and `CacheReq`. The corresponding `ClientResponse` enum carries operation results back to callers.

When entries are committed, the `apply_entry_to_state_machine` method executes the requested operation against the local Nacos state (configuration, naming, or caching), ensuring all nodes eventually converge to identical states.

## Raft RPC Network Transport Layer

Inter-node communication uses gRPC through the `RaftRouter` implementation in [`src/raft/network/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/network/core.rs), which implements the `RaftNetwork` trait:

```rust
// src/raft/network/core.rs (lines 36-71)
#[async_trait]
impl RaftNetwork<ClientRequest> for RaftRouter {
    async fn append_entries(
        &self, 
        target: NodeId, 
        req: AppendEntriesRequest<ClientRequest>
    ) -> anyhow::Result<AppendEntriesResponse> {
        let payload = PayloadUtils::build_payload(
            RAFT_APPEND_REQUEST, 
            serde_json::to_string(&req).unwrap_or_default()
        );
        let resp = self.send_request(target, payload).await?;
        Ok(serde_json::from_slice(&resp.body.unwrap_or_default().value)?)
    }
    
    // install_snapshot and vote follow the same pattern
}

```

The router resolves target node addresses from the `FileStore` and serializes Raft RPCs into Nacos-specific payloads using `PayloadUtils`. The `RaftClusterRequestSender` handles the actual HTTP transport using an actix-based client defined in [`src/raft/network/factory.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/network/factory.rs).

HTTP endpoints exposing the Raft RPCs are defined in [`src/raft/network/raft.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/network/raft.rs):

```rust
// src/raft/network/raft.rs
pub async fn vote(
    app: Data<Arc<AppShareData>>, 
    req: Json<VoteRequest>
) -> actix_web::Result<impl Responder> {
    Ok(Json(app.raft.vote(req.0).await.unwrap()))
}

pub async fn append(
    app: Data<Arc<AppShareData>>, 
    req: Json<AppendEntriesRequest<ClientRequest>>
) -> actix_web::Result<impl Responder> {
    Ok(Json(app.raft.append_entries(req.0).await.unwrap()))
}

pub async fn snapshot(
    app: Data<Arc<AppShareData>>, 
    req: Json<InstallSnapshotRequest>
) -> actix_web::Result<impl Responder> {
    Ok(Json(app.raft.install_snapshot(req.0).await.unwrap()))
}

```

These handlers allow cluster nodes to exchange vote requests, log entries, and snapshots over the public Nacos HTTP API.

## Cluster Bootstrap and Membership Management

The [`src/starter.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/starter.rs) file handles cluster initialization through two distinct paths controlled by configuration flags:

```rust
// src/starter.rs (lines 124-128)
if sys_config.raft_auto_init {
    tokio::spawn(auto_init_raft(store, raft.clone(), sys_config.clone()));
} else if !sys_config.raft_join_addr.is_empty() {
    tokio::spawn(auto_join_raft(store, sys_config.clone(), cluster_sender));
}

```

### Auto-Initialization

When `raft_auto_init` is enabled, the node creates a single-node cluster:

```rust
// src/starter.rs (lines 44-57)
raft.initialize(members).await.ok();
raft.client_write(ClientWriteRequest::new(
    ClientRequest::NodeAddr { id, addr }
)).await.ok();
raft.client_write(ClientWriteRequest::new(
    ClientRequest::Members(vec![id])
)).await.ok();

```

This establishes the initial membership and registers the node's address in the Raft state machine.

### Auto-Join

New nodes join existing clusters by contacting a seed node:

```rust
// src/starter.rs (lines 90-99)
let req = RouterRequest::JoinNode { node_id, node_addr };
let payload = PayloadUtils::build_payload(
    RAFT_ROUTE_REQUEST, 
    serde_json::to_string(&req).unwrap_or_default()
);
cluster_sender.send_request(
    Arc::new(sys_config.raft_join_addr), 
    payload
).await?;

```

The leader receives the join request and updates cluster membership through the `join_node` function in [`src/raft/mod.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/mod.rs), ensuring consensus on the new topology before acknowledging the join.

## Application-Level Writes and Strong Consistency

All mutable operations flow through the Raft client API to guarantee linearizability. The `client_write` method in [`src/raft/mod.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/mod.rs) serves as the primary interface:

```rust
// Example from src/raft/mod.rs
raft.client_write(ClientWriteRequest::new(
    ClientRequest::Members(members)
))
.await
.unwrap();

```

The `ClientRequest` enum in [`src/raft/store/mod.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/store/mod.rs) defines all permissible state machine operations:

- **NodeAddr**: Register node network addresses
- **Members**: Update cluster membership
- **ConfigSet**: Store configuration data
- **NamingReq**: Register service instances
- **CacheReq**: Update caching layer metadata

When a leader receives a `client_write` request, it follows the standard Raft protocol:

1. **Append**: Writes the request to its local log as a new entry
2. **Replicate**: Sends `AppendEntries` RPCs to followers until a majority acknowledges
3. **Apply**: Commits the entry and invokes `apply_entry_to_state_machine` in `RaftStore`
4. **Respond**: Returns a `ClientResponse` (such as `Success` or detailed query results) to the caller

This workflow ensures that every configuration change or naming update is **durably persisted** on a majority of nodes before returning success, providing strong consistency guarantees even during network partitions or node failures.

## Log Compaction and Snapshotting

To prevent unbounded log growth, r-nacos implements Raft log compaction through the snapshot mechanism configured in `build_raft`. The `snapshot_policy` triggers compaction after a configurable number of entries:

```rust
// From src/starter.rs
.snapshot_policy(async_raft_ext::SnapshotPolicy::LogsSinceLast(
    sys_config.raft_snapshot_log_size,
))
.snapshot_max_chunk_size(3 * 1024 * 1024)  // 3MB chunks

```

When the policy threshold is reached, the `do_log_compaction` method in `RaftStore` creates a point-in-time snapshot of the current state machine. The `InstallSnapshot` RPC, implemented in `RaftRouter` and exposed via the `/raft/snapshot` endpoint, transmits these snapshots to lagging followers or new nodes joining the cluster, allowing them to catch up without replaying the entire log history.

## Summary

- r-nacos embeds the **async-raft-ext** crate to provide a complete Raft consensus implementation for cluster coordination.
- The **RaftStore** in [`src/raft/store/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/store/core.rs) implements persistent storage by forwarding operations to **FileStore** actors, ensuring durability of the replicated log and snapshots.
- **RaftRouter** in [`src/raft/network/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/network/core.rs) implements the network layer, serializing Raft RPCs into Nacos-specific gRPC payloads and transmitting them via the **RaftClusterRequestSender**.
- Cluster bootstrap supports **auto-initialization** for single-node clusters and **auto-join** for adding nodes to existing clusters, managed through [`src/starter.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/starter.rs).
- All mutable operations execute through the **client_write** API, providing linearizable consistency by replicating entries to a majority before applying them to the state machine.
- **Snapshotting** prevents unbounded log growth, with configurable policies triggering compaction and chunk-based snapshot transmission to followers.

## Frequently Asked Questions

### What Raft implementation does r-nacos use?

r-nacos uses the **async-raft-ext** crate, an extended asynchronous Raft consensus library. This implementation provides the core Raft algorithms for leader election, log replication, and membership changes, which r-nacos integrates with its own storage and network layers to create the `NacosRaft` type used throughout the application.

### How does r-nacos handle node failures and leader election?

When a node fails, the remaining cluster members detect the absence of heartbeats from the leader. After the configurable election timeout (minimum 2500ms, maximum 5000ms), healthy followers transition to candidate state and request votes. The `vote` handler in [`src/raft/network/raft.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/raft/network/raft.rs) processes these requests, and once a node receives a majority of votes, it becomes the new leader and resumes normal operations.

### What is the difference between auto-init and auto-join cluster modes?

**Auto-init** mode (`raft_auto_init: true`) initializes a new single-node cluster when the first server starts, creating the initial membership and registering its own address. **Auto-join** mode uses the `raft_join_addr` configuration to contact an existing cluster node, sending a `JoinNode` request to the leader, which then adds the new node to the Raft membership through consensus before the node begins processing requests.

### How are configuration changes replicated across the cluster?

All configuration changes flow through the `client_write` method, which serializes the operation as a `ClientRequest` variant (such as `ConfigSet`). The leader appends this to its Raft log, sends `AppendEntries` RPCs to followers via `RaftRouter`, and waits for acknowledgment from a majority. Once committed, the `apply_entry_to_state_machine` method in `RaftStore` executes the change locally, ensuring every node applies the same operation in the same order.