# How r-nacos Implements Service Health Checking and Instance Heartbeat Mechanisms

> Discover how r-nacos employs dual-layer service health checking with node pings and instance heartbeat to ensure reliable service discovery and instance availability.

- Repository: [Nacos Group/r-nacos](https://github.com/nacos-group/r-nacos)
- Tags: internals
- Published: 2026-03-07

---

**r-nacos uses a dual-layer approach where `HealthManager` pings core subsystems every 2 seconds to verify node health, while `NamingActor` runs a separate timeout heartbeat to detect and remove stale service instances.**

The nacos-group/r-nacos project is a Rust implementation of the Nacos naming and configuration service. It relies on an Actix-based actor system to coordinate **service health checking** and **instance heartbeat mechanisms** that maintain cluster consistency and client connectivity.

---

## Service Health Checking Architecture

The node-level health system ensures that internal components—config, naming, cache, user, and raft stores—remain responsive. The `HealthManager` actor orchestrates this by firing periodic probes and aggregating responses.

### HealthManager Actor and the Heartbeat Loop

Located in [`src/health/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/health/core.rs), the `HealthManager` initiates a self-rescheduling timer during actor injection. The `heartbeat` method drives the loop:

```rust
// src/health/core.rs
impl HealthManager {
    // Starts the repeating heartbeat loop (2 s interval)
    fn heartbeat(&mut self, ctx: &mut Context<Self>) {
        ctx.run_later(Duration::from_millis(2000), |act, ctx| {
            act.do_check(ctx).ok();   // fire health probes
            act.heartbeat(ctx);       // schedule next run
        });
    }
}

```

This pattern—**run-later → execute logic → reschedule**—ensures continuous monitoring without blocking the async runtime.

### Component Health Probes and Ping-Pong Protocol

When `do_check` is invoked, it dispatches a `HealthCheckRequest::Ping` to every registered subsystem actor:

```rust
// src/health/core.rs
fn do_check(&mut self, ctx: &mut Context<Self>) -> anyhow::Result<()> {
    let self_addr = ctx.address();
    if let Some(config) = self.config_actor.as_ref() {
        config.do_send(HealthCheckRequest::Ping(self_addr.clone()));
    }
    if let Some(naming) = self.naming_actor.as_ref() {
        naming.do_send(HealthCheckRequest::Ping(self_addr.clone()));
    }
    // … cache, user, raft … (similar calls)
    Ok(())
}

```

Each target actor implements `Handler<HealthCheckRequest>` (visible in [`src/config/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/config/core.rs), [`src/naming/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/naming/core.rs), etc.). Upon receiving a `Ping`, the actor replies with a `HealthBackRequest::Pong` containing its `HealthCheckType` identifier, as shown in [`src/health/health_handler_impl.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/health/health_handler_impl.rs).

### Health Check Aggregation and Status Reporting

The `HealthManager` aggregates responses in its `handle` method (lines 84-90 of [`src/health/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/health/core.rs)). It updates the `last_success` timestamp of each `HealthCheckItem`. If any subsystem fails to respond within the expected window, the manager returns a `CheckHealthResult::Failed` via `HealthManagerResponse` (lines 185-191), allowing external monitoring tools to detect node degradation.

---

## Instance Heartbeat Mechanisms

While `HealthManager` monitors internal subsystems, the **Naming service** monitors the liveness of registered service instances. The `NamingActor` detects stale instances, removes empty services, and notifies clients of changes.

### NamingActor and the Timeout Heartbeat Loop

The `NamingActor` in [`src/naming/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/naming/core.rs) schedules a periodic `instance_time_out_heartbeat`:

```rust
// src/naming/core.rs
pub fn instance_time_out_heartbeat(&self, ctx: &mut actix::Context<Self>) {
    ctx.run_later(Duration::from_millis(2000), |act, ctx| {
        act.clear_empty_service();                 // remove services without instances
        act.clear_timeout_instance_metadata();     // purge expired metadata
        act.trigger_perpetual_health_check();      // optional host probes
        // Notify listeners with timeout events
        let addr = ctx.address();
        addr.do_send(NamingCmd::PeekListenerTimeout);
        // Reschedule
        act.instance_time_out_heartbeat(ctx);
    });
}

```

This method is triggered once during `NamingActor::inject` (line 49) and repeats every 2 seconds.

### Stale Instance Cleanup and Service Pruning

Within the heartbeat loop, `clear_empty_service` removes services whose `instance_size` reaches zero after a configured TTL. The `clear_timeout_instance_metadata` function purges metadata associated with expired instances. For **perpetual instances** (persistent service registrations), `trigger_perpetual_health_check` initiates `NetSniffing` probes to verify host reachability, ensuring that even non-ephemeral instances are validated.

### Delayed Listener Notifications

Instance state changes are not immediately pushed to clients. Instead, the `NamingActor` sends `NamingCmd::PeekListenerTimeout` to itself, which triggers the `DelayNotifyActor` (located in [`src/naming/naming_delay_nofity.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/naming/naming_delay_nofity.rs)).

The `DelayNotifyActor` aggregates events and flushes them periodically:

```rust
// src/naming/naming_delay_nofity.rs
pub fn notify_heartbeat(&self, ctx: &mut actix::Context<Self>) {
    ctx.run_later(Duration::from_millis(500), |act, ctx| {
        let events = act.inner_delay_notify.timeout().unwrap_or_default();
        let naming_addr = act.naming_addr.clone();
        async move {
            Self::fill_event_data_and_notify(naming_addr, events).await;
        }
        .into_actor(act)
        .map(|_, act, ctx| act.notify_heartbeat(ctx))
        .wait(ctx);
    });
}

```

Running every 500 ms, this heartbeat fetches the latest `ServiceInfo` from `NamingActor` and pushes updates to clients via `BiStreamManage`, ensuring eventual consistency between the registry and consumers.

---

## Additional Subsystem Heartbeats

Beyond the core health and naming systems, r-nacos employs the same **run-later → reschedule** pattern for other critical actors:

| Actor | Heartbeat Implementation | Source File |
|-------|--------------------------|-------------|
| **Raft** | Configured with `heartbeat_interval(1000)` for leader election and log replication | [`src/starter.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/starter.rs) (lines 306-307) |
| **MCP SSE Manager** | Sends `create_heartbeat_message` every 3 seconds to keep Server-Sent Events alive | [`src/mcp/sse_manage.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/mcp/sse_manage.rs) (lines 115-131) |
| **gRPC Bi-Stream Manager** | Calls `time_out_heartbeat` every 2 seconds to detect idle or dead client streams | [`src/grpc/bistream_manage.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/grpc/bistream_manage.rs) (lines 163-177) |
| **Cache Manager** | Runs `heartbeat` method to maintain cache actor liveness, analogous to `HealthManager` | [`src/cache/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/cache/core.rs) (lines 61-64) |

Each implementation ensures that the distributed system maintains consensus, detects network partitions, and cleans up stale resources.

---

## Practical Usage Examples

### Registering a Service Instance

When you register an instance, the naming system automatically applies heartbeat tracking. Ephemeral instances expire if not renewed, while perpetual instances undergo host probing:

```rust
use r_nacos::naming::model::Instance;
use r_nacos::naming::NamingCmd;
use std::sync::Arc;

// Build an Instance (ephemeral = false for permanent)
let instance = Instance {
    namespace_id: Arc::new("public".to_string()),
    group_name: Arc::new("DEFAULT_GROUP".to_string()),
    service_name: Arc::new("demo-service".to_string()),
    ip: "192.168.1.10".to_string(),
    port: 8080,
    weight: 1.0,
    // … other fields …
    ..Default::default()
};

// Send the registration command to the NamingActor
naming_addr.do_send(NamingCmd::Update(instance, None));

```

After registration, `NamingActor` stores the instance and its **last-beat timestamp**. The `instance_time_out_heartbeat` loop automatically expires stale entries and removes empty services.

### Querying Node Health Status

To programmatically verify that all internal subsystems are responsive, query the `HealthManager`:

```rust
use r_nacos::health::{HealthManagerRequest, HealthManagerResponse};

let health_addr: Addr<HealthManager> = /* obtained from AppShareData */;
let res = health_addr.send(HealthManagerRequest).await?;
match res {
    HealthManagerResponse::StatusResult(CheckHealthResult::Success) => {
        println!("All subsystems healthy");
    }
    HealthManagerResponse::StatusResult(failure) => {
        println!("Health check failed: {:?}", failure);
    }
}

```

This request triggers the internal heartbeat logic that pings the config, naming, cache, user, and raft actors. The response reflects the latest aggregation of `HealthCheckItem` timestamps.

---

## Summary

- **Service health checking** relies on the `HealthManager` actor in [`src/health/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/health/core.rs), which runs a 2-second `heartbeat` loop to ping all subsystems and aggregates responses into a cluster-wide health status.
- **Instance heartbeat** is managed by `NamingActor` in [`src/naming/core.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/naming/core.rs) via `instance_time_out_heartbeat`, which cleans up stale instances, removes empty services, and triggers perpetual-instance probes every 2 seconds.
- **Client notifications** are decoupled through `DelayNotifyActor` in [`src/naming/naming_delay_nofity.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/naming/naming_delay_nofity.rs), which aggregates service changes and pushes them to clients via a 500-ms heartbeat.
- All subsystems use the same Actix pattern: `run_later` → execute logic → reschedule, ensuring non-blocking, continuous monitoring across the r-nacos cluster.

---

## Frequently Asked Questions

### How does r-nacos distinguish between ephemeral and perpetual instance heartbeats?

Ephemeral instances rely on client-sent updates to refresh their last-beat timestamp in `NamingActor`. If the client stops sending updates, the `instance_time_out_heartbeat` loop marks them as expired after the configured TTL. Perpetual instances (persistent) do not expire based on client silence; instead, `trigger_perpetual_health_check` initiates `NetSniffing` host probes to verify reachability, allowing the system to detect offline persistent instances.

### What happens when the HealthManager detects a failed subsystem?

When a subsystem fails to respond to a `HealthCheckRequest::Ping` within the expected window, `HealthManager` updates the corresponding `HealthCheckItem` to reflect the failure. Subsequent queries to `HealthManager` return a `HealthManagerResponse` containing `CheckHealthResult::Failed`, which can trigger alerts or cause load balancers to stop routing traffic to the unhealthy node.

### How does the delayed notification heartbeat improve performance?

Rather than pushing every instance change immediately to all clients—which could overwhelm the network with micro-updates—`DelayNotifyActor` buffers events in `inner_delay_notify`. Its 500-ms `notify_heartbeat` aggregates multiple changes into a single batch, fetches the latest `ServiceInfo` from `NamingActor`, and pushes the consolidated update via `BiStreamManage`. This reduces connection churn and lowers CPU usage on both the server and client sides.

### Can the heartbeat intervals be configured?

The source code shows hardcoded intervals for critical loops: `HealthManager` and `NamingActor` use 2-second intervals (`Duration::from_millis(2000)`), while `DelayNotifyActor` uses 500 ms and the MCP SSE manager uses 3 seconds. While these values are compile-time constants in the analyzed version, the `Raft` subsystem does expose configuration via `heartbeat_interval(1000)` in [`src/starter.rs`](https://github.com/nacos-group/r-nacos/blob/main/src/starter.rs), suggesting that future versions may externalize these intervals via configuration files.