How r-nacos Uses the Raft Consensus Protocol for Distributed Data Storage in Cluster Mode
r-nacos embeds the async-raft-ext crate to provide strongly consistent distributed data storage in cluster mode, using Raft log replication, persistent file-based storage, and gRPC-based RPC transport to keep configuration and naming metadata synchronized across all nodes.
The r-nacos project implements a high-performance, Rust-based Nacos server that achieves fault-tolerant distributed state through the Raft consensus protocol. When operating in cluster mode, r-nacos leverages the async-raft-ext crate to replicate configuration, service naming, and metadata changes across multiple nodes, ensuring linearizable consistency and automatic failover.
Raft Instance Construction and Configuration
When the service starts, the build_raft function in src/starter.rs constructs the Raft components using the NacosRaft type alias defined in src/raft/mod.rs:
// src/starter.rs
let raft = build_raft(&sys_config, store.clone(), cluster_sender.clone()).await?;
The build_raft function configures the Raft instance with specific timing parameters and snapshot policies:
// src/starter.rs (lines 94-124)
let config = Config::build("rnacos raft".to_owned())
.heartbeat_interval(1000) // 1 second heartbeats
.election_timeout_min(2500)
.election_timeout_max(5000)
.snapshot_policy(async_raft_ext::SnapshotPolicy::LogsSinceLast(
sys_config.raft_snapshot_log_size,
))
.snapshot_max_chunk_size(3 * 1024 * 1024) // 3MB chunks
.validate()
.unwrap();
let network = Arc::new(RaftRouter::new(store.clone(), cluster_sender.clone()));
let raft = Arc::new(Raft::new(
sys_config.raft_node_id,
config,
network,
store.clone(),
));
These configuration parameters ensure the cluster remains responsive while preventing split-brain scenarios through configurable election timeouts.
Persistent Storage and the State Machine
The RaftStore struct in src/raft/store/core.rs implements the RaftStorage trait, providing durable persistence for the Raft log, hard state, and snapshots. It delegates actual I/O operations to the FileStore actors:
// src/raft/store/core.rs (excerpt)
#[async_trait]
impl RaftStorage<ClientRequest, ClientResponse> for RaftStore {
async fn get_initial_state(&self) -> anyhow::Result<InitialState> {
match self.send_store_msg(StoreRequest::GetInitialState).await? {
StoreResponse::InitialState(v) => Ok(v),
_ => Err(self.store_response_err()),
}
}
// Additional methods: get_log_entries, append_entry_to_log,
// apply_entry_to_state_machine, do_log_compaction, etc.
}
The state machine processes application-level requests defined in src/raft/store/mod.rs through the ClientRequest enum, which includes variants such as NodeAddr, Members, ConfigSet, NamingReq, and CacheReq. The corresponding ClientResponse enum carries operation results back to callers.
When entries are committed, the apply_entry_to_state_machine method executes the requested operation against the local Nacos state (configuration, naming, or caching), ensuring all nodes eventually converge to identical states.
Raft RPC Network Transport Layer
Inter-node communication uses gRPC through the RaftRouter implementation in src/raft/network/core.rs, which implements the RaftNetwork trait:
// src/raft/network/core.rs (lines 36-71)
#[async_trait]
impl RaftNetwork<ClientRequest> for RaftRouter {
async fn append_entries(
&self,
target: NodeId,
req: AppendEntriesRequest<ClientRequest>
) -> anyhow::Result<AppendEntriesResponse> {
let payload = PayloadUtils::build_payload(
RAFT_APPEND_REQUEST,
serde_json::to_string(&req).unwrap_or_default()
);
let resp = self.send_request(target, payload).await?;
Ok(serde_json::from_slice(&resp.body.unwrap_or_default().value)?)
}
// install_snapshot and vote follow the same pattern
}
The router resolves target node addresses from the FileStore and serializes Raft RPCs into Nacos-specific payloads using PayloadUtils. The RaftClusterRequestSender handles the actual HTTP transport using an actix-based client defined in src/raft/network/factory.rs.
HTTP endpoints exposing the Raft RPCs are defined in src/raft/network/raft.rs:
// src/raft/network/raft.rs
pub async fn vote(
app: Data<Arc<AppShareData>>,
req: Json<VoteRequest>
) -> actix_web::Result<impl Responder> {
Ok(Json(app.raft.vote(req.0).await.unwrap()))
}
pub async fn append(
app: Data<Arc<AppShareData>>,
req: Json<AppendEntriesRequest<ClientRequest>>
) -> actix_web::Result<impl Responder> {
Ok(Json(app.raft.append_entries(req.0).await.unwrap()))
}
pub async fn snapshot(
app: Data<Arc<AppShareData>>,
req: Json<InstallSnapshotRequest>
) -> actix_web::Result<impl Responder> {
Ok(Json(app.raft.install_snapshot(req.0).await.unwrap()))
}
These handlers allow cluster nodes to exchange vote requests, log entries, and snapshots over the public Nacos HTTP API.
Cluster Bootstrap and Membership Management
The src/starter.rs file handles cluster initialization through two distinct paths controlled by configuration flags:
// src/starter.rs (lines 124-128)
if sys_config.raft_auto_init {
tokio::spawn(auto_init_raft(store, raft.clone(), sys_config.clone()));
} else if !sys_config.raft_join_addr.is_empty() {
tokio::spawn(auto_join_raft(store, sys_config.clone(), cluster_sender));
}
Auto-Initialization
When raft_auto_init is enabled, the node creates a single-node cluster:
// src/starter.rs (lines 44-57)
raft.initialize(members).await.ok();
raft.client_write(ClientWriteRequest::new(
ClientRequest::NodeAddr { id, addr }
)).await.ok();
raft.client_write(ClientWriteRequest::new(
ClientRequest::Members(vec![id])
)).await.ok();
This establishes the initial membership and registers the node's address in the Raft state machine.
Auto-Join
New nodes join existing clusters by contacting a seed node:
// src/starter.rs (lines 90-99)
let req = RouterRequest::JoinNode { node_id, node_addr };
let payload = PayloadUtils::build_payload(
RAFT_ROUTE_REQUEST,
serde_json::to_string(&req).unwrap_or_default()
);
cluster_sender.send_request(
Arc::new(sys_config.raft_join_addr),
payload
).await?;
The leader receives the join request and updates cluster membership through the join_node function in src/raft/mod.rs, ensuring consensus on the new topology before acknowledging the join.
Application-Level Writes and Strong Consistency
All mutable operations flow through the Raft client API to guarantee linearizability. The client_write method in src/raft/mod.rs serves as the primary interface:
// Example from src/raft/mod.rs
raft.client_write(ClientWriteRequest::new(
ClientRequest::Members(members)
))
.await
.unwrap();
The ClientRequest enum in src/raft/store/mod.rs defines all permissible state machine operations:
- NodeAddr: Register node network addresses
- Members: Update cluster membership
- ConfigSet: Store configuration data
- NamingReq: Register service instances
- CacheReq: Update caching layer metadata
When a leader receives a client_write request, it follows the standard Raft protocol:
- Append: Writes the request to its local log as a new entry
- Replicate: Sends
AppendEntriesRPCs to followers until a majority acknowledges - Apply: Commits the entry and invokes
apply_entry_to_state_machineinRaftStore - Respond: Returns a
ClientResponse(such asSuccessor detailed query results) to the caller
This workflow ensures that every configuration change or naming update is durably persisted on a majority of nodes before returning success, providing strong consistency guarantees even during network partitions or node failures.
Log Compaction and Snapshotting
To prevent unbounded log growth, r-nacos implements Raft log compaction through the snapshot mechanism configured in build_raft. The snapshot_policy triggers compaction after a configurable number of entries:
// From src/starter.rs
.snapshot_policy(async_raft_ext::SnapshotPolicy::LogsSinceLast(
sys_config.raft_snapshot_log_size,
))
.snapshot_max_chunk_size(3 * 1024 * 1024) // 3MB chunks
When the policy threshold is reached, the do_log_compaction method in RaftStore creates a point-in-time snapshot of the current state machine. The InstallSnapshot RPC, implemented in RaftRouter and exposed via the /raft/snapshot endpoint, transmits these snapshots to lagging followers or new nodes joining the cluster, allowing them to catch up without replaying the entire log history.
Summary
- r-nacos embeds the async-raft-ext crate to provide a complete Raft consensus implementation for cluster coordination.
- The RaftStore in
src/raft/store/core.rsimplements persistent storage by forwarding operations to FileStore actors, ensuring durability of the replicated log and snapshots. - RaftRouter in
src/raft/network/core.rsimplements the network layer, serializing Raft RPCs into Nacos-specific gRPC payloads and transmitting them via the RaftClusterRequestSender. - Cluster bootstrap supports auto-initialization for single-node clusters and auto-join for adding nodes to existing clusters, managed through
src/starter.rs. - All mutable operations execute through the client_write API, providing linearizable consistency by replicating entries to a majority before applying them to the state machine.
- Snapshotting prevents unbounded log growth, with configurable policies triggering compaction and chunk-based snapshot transmission to followers.
Frequently Asked Questions
What Raft implementation does r-nacos use?
r-nacos uses the async-raft-ext crate, an extended asynchronous Raft consensus library. This implementation provides the core Raft algorithms for leader election, log replication, and membership changes, which r-nacos integrates with its own storage and network layers to create the NacosRaft type used throughout the application.
How does r-nacos handle node failures and leader election?
When a node fails, the remaining cluster members detect the absence of heartbeats from the leader. After the configurable election timeout (minimum 2500ms, maximum 5000ms), healthy followers transition to candidate state and request votes. The vote handler in src/raft/network/raft.rs processes these requests, and once a node receives a majority of votes, it becomes the new leader and resumes normal operations.
What is the difference between auto-init and auto-join cluster modes?
Auto-init mode (raft_auto_init: true) initializes a new single-node cluster when the first server starts, creating the initial membership and registering its own address. Auto-join mode uses the raft_join_addr configuration to contact an existing cluster node, sending a JoinNode request to the leader, which then adds the new node to the Raft membership through consensus before the node begins processing requests.
How are configuration changes replicated across the cluster?
All configuration changes flow through the client_write method, which serializes the operation as a ClientRequest variant (such as ConfigSet). The leader appends this to its Raft log, sends AppendEntries RPCs to followers via RaftRouter, and waits for acknowledgment from a majority. Once committed, the apply_entry_to_state_machine method in RaftStore executes the change locally, ensuring every node applies the same operation in the same order.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →