How Multi-Repo Daemon Health Checks and Auto-Restart Work in Code Review Graph

The WatchDaemon in code-review-graph uses a dedicated health-checker thread to monitor child watcher processes every 30 seconds, automatically restarting any that have crashed or exited, while a separate config watcher handles live reloading of watch.toml changes.

The multi-repo watch daemon (WatchDaemon) is the core process orchestration component of the tirth8205/code-review-graph repository. It manages a fleet of per-repository watcher subprocesses, ensuring continuous code analysis even when individual workers fail. Understanding how its health check and auto-restart mechanisms function is essential for operating a reliable multi-repo code review pipeline.

Health-Checker Thread Architecture

The daemon's resilience stems from three cooperating components: a health-checker thread, a config watcher, and PID persistence. These work together to maintain process availability and configuration consistency.

Component Overview

Component Purpose Key Location
Health-checker thread Detects dead watchers and restarts them WatchDaemon._health_loop
Config watcher Reacts to watch.toml modifications ConfigWatcher.start
State persistence Enables cross-process status queries WatchDaemon._save_state

Starting the Health-Checker

When WatchDaemon.start() is invoked, it spawns the health-checker as a daemon thread. In daemon.py, the start_health_checker method initializes the monitoring infrastructure:

def start_health_checker(self) -> None:
    self._health_stop = threading.Event()
    self._health_thread = threading.Thread(
        target=self._health_loop,
        daemon=True,
        name="health-checker",
    )
    self._health_thread.start()
    logger.info(
        "Health checker started (interval=%ds)",
        _HEALTH_CHECK_INTERVAL,
    )

The _HEALTH_CHECK_INTERVAL constant defaults to 30 seconds (defined at line 100). Using a threading.Event for _health_stop ensures the thread can be interrupted promptly during daemon shutdown rather than blocking on a fixed sleep.

The Health-Check Loop

The _health_loop method implements an interruptible polling pattern:

def _health_loop(self) -> None:
    while not self._health_stop.is_set():
        self._health_stop.wait(_HEALTH_CHECK_INTERVAL)
        if self._health_stop.is_set():
            break
        self._check_health()

This design provides two advantages:

  • Responsive shutdown: wait(timeout) returns immediately when set() is called, avoiding fixed delays during termination
  • Regular health checks: Falls through to _check_health() every 30 seconds when running normally

Detecting and Restarting Dead Watchers

The core detection logic resides in _check_health at line 104 of daemon.py:

def _check_health(self) -> None:
    restarted = False
    with self._lock:
        for alias, repo in list(self._current_repos.items()):
            proc = self._children.get(alias)
            if proc is None or proc.poll() is not None:   # <-- dead process detection

                logger.warning("Watcher for '%s' is dead — restarting", alias)
                self._children.pop(alias, None)
                self._start_watcher(repo)                 # <-- auto-restart

                restarted = True
    if restarted:
        self._save_state()

Process liveness detection relies on subprocess.Popen.poll(), which returns None for running processes and an exit code integer for terminated ones. The implementation:

  1. Acquires _lock to coordinate with config changes
  2. Iterates over all tracked repositories
  3. Removes stale entries from _children
  4. Spawns fresh watchers via _start_watcher(repo)
  5. Persists updated PIDs to daemon-state.json if any restarts occurred

The _save_state call ensures that external CLI commands like crg status can accurately report current process states even when executed from different shell sessions.

Auto-Restart on Configuration Changes

Beyond crash recovery, the daemon handles live configuration reloading. The ConfigWatcher component (started at line 55) monitors watch.toml for modifications using platform-specific file system notifications or polling fallback.

When changes are detected, _on_config_change (line 152) triggers:

def _on_config_change(self) -> None:
    new_config = load_config()
    self.reconcile(new_config)  # Computes diff and starts/stops watchers accordingly

The reconcile method computes the symmetric difference between desired and current repository sets:

  • Added repos: Immediately launch new watcher processes
  • Removed repos: Terminate associated child processes cleanly
  • Modified repos: Restart watchers to pick up new paths or settings

This mechanism operates independently yet compatibly with the health-checker—newly added repositories become subject to health monitoring automatically, while removals clean up both running processes and monitoring entries.

Cross-Process Liveness and State Recovery

The daemon maintains durable state across potential crashes through two persistence mechanisms:

  • daemon.pid: Contains the main daemon's PID for singleton enforcement and crg daemon status queries
  • daemon-state.json: Maps repository aliases to active watcher PIDs

The helper function pid_alive abstracts OS-specific liveness checks:

  • POSIX: os.kill(pid, 0) — sends null signal to test existence
  • Windows: OpenProcess with PROCESS_QUERY_INFORMATION access

This abstraction guarantees reliable stale PID detection when the daemon restarts after unexpected termination.

Practical Examples

Running the Daemon in Foreground Mode

from code_review_graph.daemon import WatchDaemon

daemon = WatchDaemon()
daemon.start()            # spawns watchers and health-checker

daemon.run_forever()      # blocks until interrupted

The daemon reads ~/.code-review-graph/watch.toml, launches a subprocess per configured repository, and begins health monitoring immediately.

Simulating Watcher Failure and Auto-Restart

import os, signal, time
from code_review_graph.daemon import WatchDaemon

daemon = WatchDaemon()
daemon.start()
time.sleep(2)  # Allow watcher spawn

# Force-kill the first watcher child

child_pid = next(iter(daemon._children.values())).pid
print(f"Killing watcher {child_pid}")
os.kill(child_pid, signal.SIGKILL)

# Wait for health-checker detection (>30s)

time.sleep(35)

new_pid = daemon._children[list(daemon._children)[0]].pid
assert new_pid != child_pid, "Watcher was not restarted"
print(f"Auto-restarted: new PID {new_pid}")

Live Configuration Update

from code_review_graph.daemon import add_repo_to_config

# Modify configuration on disk

add_repo_to_config("/path/to/new/repo", alias="new-project")

# ConfigWatcher detects change within seconds and calls reconcile,

# starting a watcher for "new-project" without daemon restart

Daemonization and Platform Behavior

For production deployment, daemon.daemonize() performs a double-fork on Unix systems:

daemon = WatchDaemon()
daemon.daemonize()  # Detaches from terminal, writes PID file

daemon.start()
daemon.run_forever()

On Windows, the method logs a warning and runs in foreground mode since fork() semantics are unavailable. The health-checker and config watcher function identically in both modes.

Test Coverage

The implementation is exercised in tests/test_daemon.py, which validates:

  • PID file creation and cleanup
  • Health-checker thread lifecycle
  • Config reload and reconcile operations
  • Signal handling for graceful shutdown
  • State file consistency across restarts

Additional thread-spawning patterns appear in tests/test_incremental.py, demonstrating daemon=True thread configuration and health-check assertions.

Summary

  • Health-checker thread polls every 30 seconds via _health_loop, detecting crashed watchers with proc.poll()
  • Auto-restart immediately spawns replacement processes and updates daemon-state.json for external visibility
  • Config watcher enables live reloading of watch.toml without daemon restart through reconcile diffing
  • PID persistence via JSON state file and platform-abstracted pid_alive ensures accurate cross-process status reporting
  • Thread-safety through _lock coordination prevents races between health checks and configuration changes

Frequently Asked Questions

How often does the daemon check watcher health?

The health-checker thread runs every 30 seconds by default, as defined by _HEALTH_CHECK_INTERVAL in daemon.py. This interval balances responsiveness with system overhead for typical code review workloads.

What happens if the daemon itself crashes?

Child watcher processes become orphaned but continue running. When the daemon restarts, it reads daemon-state.json and uses pid_alive to determine which watchers are still active. New watchers are started only for missing or stale entries, preserving continuity where possible.

Can I change the health-check interval?

The interval is currently hardcoded as _HEALTH_CHECK_INTERVAL = 30 at line 100 of daemon.py. Modifying this requires editing the source or submitting a pull request to make it configurable via watch.toml.

Does auto-restart preserve watcher state?

Individual watcher processes are stateless by design—they re-scan repository state on startup. The daemon maintains no in-memory state about watcher progress, so restarts are safe and idempotent. Any incremental tracking relies on persistent storage outside the daemon (e.g., commit hashes in a database).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →