How iroh Handles Connection Closed Events and Graceful Shutdown

iroh implements a fully-async, token-driven shutdown flow using CancellationToken from tokio-util that guarantees every background task, connection, and runtime resource finishes cleanly before the process exits.

The n0-computer/iroh repository provides a modern, async networking stack that requires careful coordination when shutting down to prevent resource leaks. Understanding how iroh handles connection closed events and graceful shutdown is essential for building reliable applications that manage network resources correctly. This article examines the token-driven architecture that enables clean termination of the router, protocol handlers, and underlying runtime.

The Token-Driven Shutdown Architecture

At the core of iroh's graceful shutdown mechanism is a hierarchical token system based on tokio_util::sync::CancellationToken. This architecture ensures that when a shutdown is initiated, every component—from the top-level router down to individual transport actors—receives the signal and can perform necessary cleanup before exiting.

Router and the Shutdown Entry Point

The Router serves as the top-level object that owns the network stack, protocol handlers, and runtime. When your application needs to shut down, calling router.shutdown().await triggers the entire shutdown cascade.

In iroh/src/protocol.rs, the shutdown method first checks Router::is_shutdown (line 415) to prevent double-shutdown scenarios, then cancels the internal shutdown token (line 424). This single action propagates through the entire system, notifying all dependent tasks to begin their cleanup procedures.

ShutdownState and Cancellation Tokens

The ShutdownState struct manages two distinct cancellation tokens: at_close_start (signaling that the endpoint is beginning to close) and at_endpoint_closed (signaling the endpoint has fully closed).

As implemented in iroh/src/socket.rs (line 347), the Socket struct stores this ShutdownState. When the router initiates shutdown, it cancels at_close_start (line 1140), which propagates to all tasks holding a child token. Every long-living task, such as the relay actor in iroh/src/socket/transports/relay/actor.rs (line 424), holds a child token and observes the parent's cancelled() future to trigger its own exit logic.

ProtocolHandler Integration

Each protocol in iroh (relay, DNS, etc.) implements a ProtocolHandler trait with a dedicated shutdown method. During the shutdown sequence, the router invokes these handlers concurrently (lines 398-404 in protocol.rs), giving each protocol a chance to close its own connections and cancel internal tokens before the runtime stops accepting new tasks.

The Graceful Shutdown Sequence Step-by-Step

When your application calls router.shutdown().await, iroh executes a precise sequence to ensure no resources are left dangling:

  1. Router validation: The system checks is_shutdown to verify this is the first shutdown request.
  2. Token cancellation: The router-wide shutdown token is cancelled, triggering ShutdownState::at_close_start.
  3. Task propagation: The cancellation propagates to every task holding a child token, including transport actors and protocol handlers.
  4. Protocol cleanup: Each ProtocolHandler receives a shutdown request, closes its connections, and cancels transport-specific tokens.
  5. Socket drainage: The Socket task observes its shutdown_token cancelled (as seen in the select! branch at line 1507 of socket.rs), stops accepting new packets, and drains remaining inbound/outbound work.
  6. Grace period: The system waits 100ms (implemented in socket.rs lines 1173-1181) to allow tasks to finish before forcing runtime shutdown.
  7. Runtime termination: The router awaits self.runtime.shutdown().await (line 1192), ensuring the Tokio runtime stops accepting new tasks and drains pending work.
  8. Final state: The closed flag on ShutdownState is set (line 1194) to indicate complete termination.

This flow works uniformly across all transports (Relay, UDP, QUIC, etc.) because it relies on cooperative cancellation rather than abrupt aborts.

Key Implementation Details

Socket and Endpoint Lifecycle

The Socket struct in iroh/src/socket.rs serves as the entry point for inbound and outbound connections. When a socket is dropped, it first cancels its shutdown token (line 1225). The per-endpoint task loop observes this cancellation via a select! branch matching _ = shutdown_token.cancelled() (line 1507), allowing it to exit cleanly after processing any remaining packets.

Runtime Drain and Grace Period

After all protocol handlers have been asked to shut down, the router performs a final cleanup of the Runtime. The implementation in socket.rs includes a 100ms grace period (lines 1173-1181) that waits for tasks to finish naturally. If this timeout fires, a warning is logged but the shutdown proceeds, preventing deadlocked processes from hanging indefinitely.

Practical Code Examples

Initiating Graceful Shutdown

The primary interface for graceful shutdown is the Router method:

use iroh::protocol::Router;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    // Build and start the router
    let router = Router::builder()
        .listen("tcp://0.0.0.0:12345")?
        .build()
        .await?;

    // ... use the router (e.g. start a client, serve connections) ...

    // When the program receives SIGINT or the test ends:
    router.shutdown().await?;
    // All connections are closed, background tasks have been cancelled,
    // and the runtime is cleanly shut down.
    Ok(())
}

Observing Shutdown in Custom Tasks

To ensure your background tasks participate in graceful shutdown, observe the cancellation token:

use tokio::select;
use tokio_util::sync::CancellationToken;
use std::time::Duration;

async fn my_background_task(shutdown: CancellationToken) {
    loop {
        select! {
            // Normal work
            _ = do_some_io().fuse() => { /* … */ }

            // React to graceful shutdown
            _ = shutdown.cancelled() => {
                // Perform any last-minute cleanup here
                log::info!("my_background_task: shutting down");
                break;
            }
        }
    }
}

Embedding Tasks in the Router

When spawning custom tasks within the iroh ecosystem, clone the router's shutdown token to ensure proper coordination:

let shutdown = router.shutdown_token.clone(); // the token is exposed via `Router::shutdown_token`
tokio::spawn(my_background_task(shutdown));

When router.shutdown().await is called, the token is cancelled, causing my_background_task to exit after finishing its current iteration.

Summary

  • iroh uses a hierarchical CancellationToken system from tokio-util to coordinate shutdown across all async components.
  • The Router in iroh/src/protocol.rs initiates shutdown by cancelling its internal token, which propagates through ShutdownState in iroh/src/socket.rs.
  • Every ProtocolHandler receives a concurrent shutdown request, allowing protocols to close connections cleanly before the runtime stops.
  • A 100ms grace period ensures tasks have time to finish, after which the Tokio runtime is drained and stopped.
  • Custom tasks can participate in graceful shutdown by cloning the router's shutdown_token and observing its cancelled() future.

Frequently Asked Questions

How does iroh prevent resource leaks during shutdown?

iroh prevents resource leaks by implementing a cooperative cancellation model using CancellationToken. Rather than aborting tasks abruptly, the system signals them to stop via token cancellation, allowing each component—from the Socket in iroh/src/socket.rs to individual transport actors like those in iroh/src/socket/transports/relay/actor.rs—to perform necessary cleanup before exiting. The router then waits for all protocol handlers to complete and drains the runtime before exiting.

What happens if a task doesn't finish within the 100ms grace period?

If a task doesn't finish within the 100ms grace period implemented in iroh/src/socket.rs (lines 1173-1181), the system logs a warning but proceeds with the runtime shutdown anyway. This prevents the process from hanging indefinitely due to a misbehaving task while still allowing most well-behaved tasks time to complete their cleanup operations.

Can I implement custom shutdown logic for my protocol?

Yes, you can implement custom shutdown logic by implementing the ProtocolHandler trait for your protocol. The router calls each handler's shutdown method concurrently (as seen in iroh/src/protocol.rs lines 398-404), giving you a chance to close connections, cancel internal tokens, and perform any necessary cleanup before the system shuts down the runtime.

How do I observe shutdown events in my own async tasks?

To observe shutdown events in your own async tasks, clone the router's shutdown_token (exposed via Router::shutdown_token) and use tokio::select! to watch for the shutdown.cancelled() future. When the router initiates shutdown, this future resolves, allowing your task to break out of its work loop and perform cleanup before the system completes the shutdown sequence.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →