How r-nacos Implements Service Health Checking and Instance Heartbeat Mechanisms
r-nacos uses a dual-layer approach where HealthManager pings core subsystems every 2 seconds to verify node health, while NamingActor runs a separate timeout heartbeat to detect and remove stale service instances.
The nacos-group/r-nacos project is a Rust implementation of the Nacos naming and configuration service. It relies on an Actix-based actor system to coordinate service health checking and instance heartbeat mechanisms that maintain cluster consistency and client connectivity.
Service Health Checking Architecture
The node-level health system ensures that internal components—config, naming, cache, user, and raft stores—remain responsive. The HealthManager actor orchestrates this by firing periodic probes and aggregating responses.
HealthManager Actor and the Heartbeat Loop
Located in src/health/core.rs, the HealthManager initiates a self-rescheduling timer during actor injection. The heartbeat method drives the loop:
// src/health/core.rs
impl HealthManager {
// Starts the repeating heartbeat loop (2 s interval)
fn heartbeat(&mut self, ctx: &mut Context<Self>) {
ctx.run_later(Duration::from_millis(2000), |act, ctx| {
act.do_check(ctx).ok(); // fire health probes
act.heartbeat(ctx); // schedule next run
});
}
}
This pattern—run-later → execute logic → reschedule—ensures continuous monitoring without blocking the async runtime.
Component Health Probes and Ping-Pong Protocol
When do_check is invoked, it dispatches a HealthCheckRequest::Ping to every registered subsystem actor:
// src/health/core.rs
fn do_check(&mut self, ctx: &mut Context<Self>) -> anyhow::Result<()> {
let self_addr = ctx.address();
if let Some(config) = self.config_actor.as_ref() {
config.do_send(HealthCheckRequest::Ping(self_addr.clone()));
}
if let Some(naming) = self.naming_actor.as_ref() {
naming.do_send(HealthCheckRequest::Ping(self_addr.clone()));
}
// … cache, user, raft … (similar calls)
Ok(())
}
Each target actor implements Handler<HealthCheckRequest> (visible in src/config/core.rs, src/naming/core.rs, etc.). Upon receiving a Ping, the actor replies with a HealthBackRequest::Pong containing its HealthCheckType identifier, as shown in src/health/health_handler_impl.rs.
Health Check Aggregation and Status Reporting
The HealthManager aggregates responses in its handle method (lines 84-90 of src/health/core.rs). It updates the last_success timestamp of each HealthCheckItem. If any subsystem fails to respond within the expected window, the manager returns a CheckHealthResult::Failed via HealthManagerResponse (lines 185-191), allowing external monitoring tools to detect node degradation.
Instance Heartbeat Mechanisms
While HealthManager monitors internal subsystems, the Naming service monitors the liveness of registered service instances. The NamingActor detects stale instances, removes empty services, and notifies clients of changes.
NamingActor and the Timeout Heartbeat Loop
The NamingActor in src/naming/core.rs schedules a periodic instance_time_out_heartbeat:
// src/naming/core.rs
pub fn instance_time_out_heartbeat(&self, ctx: &mut actix::Context<Self>) {
ctx.run_later(Duration::from_millis(2000), |act, ctx| {
act.clear_empty_service(); // remove services without instances
act.clear_timeout_instance_metadata(); // purge expired metadata
act.trigger_perpetual_health_check(); // optional host probes
// Notify listeners with timeout events
let addr = ctx.address();
addr.do_send(NamingCmd::PeekListenerTimeout);
// Reschedule
act.instance_time_out_heartbeat(ctx);
});
}
This method is triggered once during NamingActor::inject (line 49) and repeats every 2 seconds.
Stale Instance Cleanup and Service Pruning
Within the heartbeat loop, clear_empty_service removes services whose instance_size reaches zero after a configured TTL. The clear_timeout_instance_metadata function purges metadata associated with expired instances. For perpetual instances (persistent service registrations), trigger_perpetual_health_check initiates NetSniffing probes to verify host reachability, ensuring that even non-ephemeral instances are validated.
Delayed Listener Notifications
Instance state changes are not immediately pushed to clients. Instead, the NamingActor sends NamingCmd::PeekListenerTimeout to itself, which triggers the DelayNotifyActor (located in src/naming/naming_delay_nofity.rs).
The DelayNotifyActor aggregates events and flushes them periodically:
// src/naming/naming_delay_nofity.rs
pub fn notify_heartbeat(&self, ctx: &mut actix::Context<Self>) {
ctx.run_later(Duration::from_millis(500), |act, ctx| {
let events = act.inner_delay_notify.timeout().unwrap_or_default();
let naming_addr = act.naming_addr.clone();
async move {
Self::fill_event_data_and_notify(naming_addr, events).await;
}
.into_actor(act)
.map(|_, act, ctx| act.notify_heartbeat(ctx))
.wait(ctx);
});
}
Running every 500 ms, this heartbeat fetches the latest ServiceInfo from NamingActor and pushes updates to clients via BiStreamManage, ensuring eventual consistency between the registry and consumers.
Additional Subsystem Heartbeats
Beyond the core health and naming systems, r-nacos employs the same run-later → reschedule pattern for other critical actors:
| Actor | Heartbeat Implementation | Source File |
|---|---|---|
| Raft | Configured with heartbeat_interval(1000) for leader election and log replication |
src/starter.rs (lines 306-307) |
| MCP SSE Manager | Sends create_heartbeat_message every 3 seconds to keep Server-Sent Events alive |
src/mcp/sse_manage.rs (lines 115-131) |
| gRPC Bi-Stream Manager | Calls time_out_heartbeat every 2 seconds to detect idle or dead client streams |
src/grpc/bistream_manage.rs (lines 163-177) |
| Cache Manager | Runs heartbeat method to maintain cache actor liveness, analogous to HealthManager |
src/cache/core.rs (lines 61-64) |
Each implementation ensures that the distributed system maintains consensus, detects network partitions, and cleans up stale resources.
Practical Usage Examples
Registering a Service Instance
When you register an instance, the naming system automatically applies heartbeat tracking. Ephemeral instances expire if not renewed, while perpetual instances undergo host probing:
use r_nacos::naming::model::Instance;
use r_nacos::naming::NamingCmd;
use std::sync::Arc;
// Build an Instance (ephemeral = false for permanent)
let instance = Instance {
namespace_id: Arc::new("public".to_string()),
group_name: Arc::new("DEFAULT_GROUP".to_string()),
service_name: Arc::new("demo-service".to_string()),
ip: "192.168.1.10".to_string(),
port: 8080,
weight: 1.0,
// … other fields …
..Default::default()
};
// Send the registration command to the NamingActor
naming_addr.do_send(NamingCmd::Update(instance, None));
After registration, NamingActor stores the instance and its last-beat timestamp. The instance_time_out_heartbeat loop automatically expires stale entries and removes empty services.
Querying Node Health Status
To programmatically verify that all internal subsystems are responsive, query the HealthManager:
use r_nacos::health::{HealthManagerRequest, HealthManagerResponse};
let health_addr: Addr<HealthManager> = /* obtained from AppShareData */;
let res = health_addr.send(HealthManagerRequest).await?;
match res {
HealthManagerResponse::StatusResult(CheckHealthResult::Success) => {
println!("All subsystems healthy");
}
HealthManagerResponse::StatusResult(failure) => {
println!("Health check failed: {:?}", failure);
}
}
This request triggers the internal heartbeat logic that pings the config, naming, cache, user, and raft actors. The response reflects the latest aggregation of HealthCheckItem timestamps.
Summary
- Service health checking relies on the
HealthManageractor insrc/health/core.rs, which runs a 2-secondheartbeatloop to ping all subsystems and aggregates responses into a cluster-wide health status. - Instance heartbeat is managed by
NamingActorinsrc/naming/core.rsviainstance_time_out_heartbeat, which cleans up stale instances, removes empty services, and triggers perpetual-instance probes every 2 seconds. - Client notifications are decoupled through
DelayNotifyActorinsrc/naming/naming_delay_nofity.rs, which aggregates service changes and pushes them to clients via a 500-ms heartbeat. - All subsystems use the same Actix pattern:
run_later→ execute logic → reschedule, ensuring non-blocking, continuous monitoring across the r-nacos cluster.
Frequently Asked Questions
How does r-nacos distinguish between ephemeral and perpetual instance heartbeats?
Ephemeral instances rely on client-sent updates to refresh their last-beat timestamp in NamingActor. If the client stops sending updates, the instance_time_out_heartbeat loop marks them as expired after the configured TTL. Perpetual instances (persistent) do not expire based on client silence; instead, trigger_perpetual_health_check initiates NetSniffing host probes to verify reachability, allowing the system to detect offline persistent instances.
What happens when the HealthManager detects a failed subsystem?
When a subsystem fails to respond to a HealthCheckRequest::Ping within the expected window, HealthManager updates the corresponding HealthCheckItem to reflect the failure. Subsequent queries to HealthManager return a HealthManagerResponse containing CheckHealthResult::Failed, which can trigger alerts or cause load balancers to stop routing traffic to the unhealthy node.
How does the delayed notification heartbeat improve performance?
Rather than pushing every instance change immediately to all clients—which could overwhelm the network with micro-updates—DelayNotifyActor buffers events in inner_delay_notify. Its 500-ms notify_heartbeat aggregates multiple changes into a single batch, fetches the latest ServiceInfo from NamingActor, and pushes the consolidated update via BiStreamManage. This reduces connection churn and lowers CPU usage on both the server and client sides.
Can the heartbeat intervals be configured?
The source code shows hardcoded intervals for critical loops: HealthManager and NamingActor use 2-second intervals (Duration::from_millis(2000)), while DelayNotifyActor uses 500 ms and the MCP SSE manager uses 3 seconds. While these values are compile-time constants in the analyzed version, the Raft subsystem does expose configuration via heartbeat_interval(1000) in src/starter.rs, suggesting that future versions may externalize these intervals via configuration files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →