How r-nacos Implements Service Health Checking and Instance Heartbeat Mechanisms

r-nacos uses a dual-layer approach where HealthManager pings core subsystems every 2 seconds to verify node health, while NamingActor runs a separate timeout heartbeat to detect and remove stale service instances.

The nacos-group/r-nacos project is a Rust implementation of the Nacos naming and configuration service. It relies on an Actix-based actor system to coordinate service health checking and instance heartbeat mechanisms that maintain cluster consistency and client connectivity.


Service Health Checking Architecture

The node-level health system ensures that internal components—config, naming, cache, user, and raft stores—remain responsive. The HealthManager actor orchestrates this by firing periodic probes and aggregating responses.

HealthManager Actor and the Heartbeat Loop

Located in src/health/core.rs, the HealthManager initiates a self-rescheduling timer during actor injection. The heartbeat method drives the loop:

// src/health/core.rs
impl HealthManager {
    // Starts the repeating heartbeat loop (2 s interval)
    fn heartbeat(&mut self, ctx: &mut Context<Self>) {
        ctx.run_later(Duration::from_millis(2000), |act, ctx| {
            act.do_check(ctx).ok();   // fire health probes
            act.heartbeat(ctx);       // schedule next run
        });
    }
}

This pattern—run-later → execute logic → reschedule—ensures continuous monitoring without blocking the async runtime.

Component Health Probes and Ping-Pong Protocol

When do_check is invoked, it dispatches a HealthCheckRequest::Ping to every registered subsystem actor:

// src/health/core.rs
fn do_check(&mut self, ctx: &mut Context<Self>) -> anyhow::Result<()> {
    let self_addr = ctx.address();
    if let Some(config) = self.config_actor.as_ref() {
        config.do_send(HealthCheckRequest::Ping(self_addr.clone()));
    }
    if let Some(naming) = self.naming_actor.as_ref() {
        naming.do_send(HealthCheckRequest::Ping(self_addr.clone()));
    }
    // … cache, user, raft … (similar calls)
    Ok(())
}

Each target actor implements Handler<HealthCheckRequest> (visible in src/config/core.rs, src/naming/core.rs, etc.). Upon receiving a Ping, the actor replies with a HealthBackRequest::Pong containing its HealthCheckType identifier, as shown in src/health/health_handler_impl.rs.

Health Check Aggregation and Status Reporting

The HealthManager aggregates responses in its handle method (lines 84-90 of src/health/core.rs). It updates the last_success timestamp of each HealthCheckItem. If any subsystem fails to respond within the expected window, the manager returns a CheckHealthResult::Failed via HealthManagerResponse (lines 185-191), allowing external monitoring tools to detect node degradation.


Instance Heartbeat Mechanisms

While HealthManager monitors internal subsystems, the Naming service monitors the liveness of registered service instances. The NamingActor detects stale instances, removes empty services, and notifies clients of changes.

NamingActor and the Timeout Heartbeat Loop

The NamingActor in src/naming/core.rs schedules a periodic instance_time_out_heartbeat:

// src/naming/core.rs
pub fn instance_time_out_heartbeat(&self, ctx: &mut actix::Context<Self>) {
    ctx.run_later(Duration::from_millis(2000), |act, ctx| {
        act.clear_empty_service();                 // remove services without instances
        act.clear_timeout_instance_metadata();     // purge expired metadata
        act.trigger_perpetual_health_check();      // optional host probes
        // Notify listeners with timeout events
        let addr = ctx.address();
        addr.do_send(NamingCmd::PeekListenerTimeout);
        // Reschedule
        act.instance_time_out_heartbeat(ctx);
    });
}

This method is triggered once during NamingActor::inject (line 49) and repeats every 2 seconds.

Stale Instance Cleanup and Service Pruning

Within the heartbeat loop, clear_empty_service removes services whose instance_size reaches zero after a configured TTL. The clear_timeout_instance_metadata function purges metadata associated with expired instances. For perpetual instances (persistent service registrations), trigger_perpetual_health_check initiates NetSniffing probes to verify host reachability, ensuring that even non-ephemeral instances are validated.

Delayed Listener Notifications

Instance state changes are not immediately pushed to clients. Instead, the NamingActor sends NamingCmd::PeekListenerTimeout to itself, which triggers the DelayNotifyActor (located in src/naming/naming_delay_nofity.rs).

The DelayNotifyActor aggregates events and flushes them periodically:

// src/naming/naming_delay_nofity.rs
pub fn notify_heartbeat(&self, ctx: &mut actix::Context<Self>) {
    ctx.run_later(Duration::from_millis(500), |act, ctx| {
        let events = act.inner_delay_notify.timeout().unwrap_or_default();
        let naming_addr = act.naming_addr.clone();
        async move {
            Self::fill_event_data_and_notify(naming_addr, events).await;
        }
        .into_actor(act)
        .map(|_, act, ctx| act.notify_heartbeat(ctx))
        .wait(ctx);
    });
}

Running every 500 ms, this heartbeat fetches the latest ServiceInfo from NamingActor and pushes updates to clients via BiStreamManage, ensuring eventual consistency between the registry and consumers.


Additional Subsystem Heartbeats

Beyond the core health and naming systems, r-nacos employs the same run-later → reschedule pattern for other critical actors:

Actor Heartbeat Implementation Source File
Raft Configured with heartbeat_interval(1000) for leader election and log replication src/starter.rs (lines 306-307)
MCP SSE Manager Sends create_heartbeat_message every 3 seconds to keep Server-Sent Events alive src/mcp/sse_manage.rs (lines 115-131)
gRPC Bi-Stream Manager Calls time_out_heartbeat every 2 seconds to detect idle or dead client streams src/grpc/bistream_manage.rs (lines 163-177)
Cache Manager Runs heartbeat method to maintain cache actor liveness, analogous to HealthManager src/cache/core.rs (lines 61-64)

Each implementation ensures that the distributed system maintains consensus, detects network partitions, and cleans up stale resources.


Practical Usage Examples

Registering a Service Instance

When you register an instance, the naming system automatically applies heartbeat tracking. Ephemeral instances expire if not renewed, while perpetual instances undergo host probing:

use r_nacos::naming::model::Instance;
use r_nacos::naming::NamingCmd;
use std::sync::Arc;

// Build an Instance (ephemeral = false for permanent)
let instance = Instance {
    namespace_id: Arc::new("public".to_string()),
    group_name: Arc::new("DEFAULT_GROUP".to_string()),
    service_name: Arc::new("demo-service".to_string()),
    ip: "192.168.1.10".to_string(),
    port: 8080,
    weight: 1.0,
    // … other fields …
    ..Default::default()
};

// Send the registration command to the NamingActor
naming_addr.do_send(NamingCmd::Update(instance, None));

After registration, NamingActor stores the instance and its last-beat timestamp. The instance_time_out_heartbeat loop automatically expires stale entries and removes empty services.

Querying Node Health Status

To programmatically verify that all internal subsystems are responsive, query the HealthManager:

use r_nacos::health::{HealthManagerRequest, HealthManagerResponse};

let health_addr: Addr<HealthManager> = /* obtained from AppShareData */;
let res = health_addr.send(HealthManagerRequest).await?;
match res {
    HealthManagerResponse::StatusResult(CheckHealthResult::Success) => {
        println!("All subsystems healthy");
    }
    HealthManagerResponse::StatusResult(failure) => {
        println!("Health check failed: {:?}", failure);
    }
}

This request triggers the internal heartbeat logic that pings the config, naming, cache, user, and raft actors. The response reflects the latest aggregation of HealthCheckItem timestamps.


Summary

  • Service health checking relies on the HealthManager actor in src/health/core.rs, which runs a 2-second heartbeat loop to ping all subsystems and aggregates responses into a cluster-wide health status.
  • Instance heartbeat is managed by NamingActor in src/naming/core.rs via instance_time_out_heartbeat, which cleans up stale instances, removes empty services, and triggers perpetual-instance probes every 2 seconds.
  • Client notifications are decoupled through DelayNotifyActor in src/naming/naming_delay_nofity.rs, which aggregates service changes and pushes them to clients via a 500-ms heartbeat.
  • All subsystems use the same Actix pattern: run_later → execute logic → reschedule, ensuring non-blocking, continuous monitoring across the r-nacos cluster.

Frequently Asked Questions

How does r-nacos distinguish between ephemeral and perpetual instance heartbeats?

Ephemeral instances rely on client-sent updates to refresh their last-beat timestamp in NamingActor. If the client stops sending updates, the instance_time_out_heartbeat loop marks them as expired after the configured TTL. Perpetual instances (persistent) do not expire based on client silence; instead, trigger_perpetual_health_check initiates NetSniffing host probes to verify reachability, allowing the system to detect offline persistent instances.

What happens when the HealthManager detects a failed subsystem?

When a subsystem fails to respond to a HealthCheckRequest::Ping within the expected window, HealthManager updates the corresponding HealthCheckItem to reflect the failure. Subsequent queries to HealthManager return a HealthManagerResponse containing CheckHealthResult::Failed, which can trigger alerts or cause load balancers to stop routing traffic to the unhealthy node.

How does the delayed notification heartbeat improve performance?

Rather than pushing every instance change immediately to all clients—which could overwhelm the network with micro-updates—DelayNotifyActor buffers events in inner_delay_notify. Its 500-ms notify_heartbeat aggregates multiple changes into a single batch, fetches the latest ServiceInfo from NamingActor, and pushes the consolidated update via BiStreamManage. This reduces connection churn and lowers CPU usage on both the server and client sides.

Can the heartbeat intervals be configured?

The source code shows hardcoded intervals for critical loops: HealthManager and NamingActor use 2-second intervals (Duration::from_millis(2000)), while DelayNotifyActor uses 500 ms and the MCP SSE manager uses 3 seconds. While these values are compile-time constants in the analyzed version, the Raft subsystem does expose configuration via heartbeat_interval(1000) in src/starter.rs, suggesting that future versions may externalize these intervals via configuration files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →