# How Multi-Repo Daemon Health Checks and Auto-Restart Work in Code Review Graph

> Learn how multi-repo daemon health checks and auto-restart work in Code Review Graph. Discover how the WatchDaemon monitors and restarts child processes automatically.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: how-to-guide
- Published: 2026-08-16

---

**The `WatchDaemon` in `code-review-graph` uses a dedicated health-checker thread to monitor child watcher processes every 30 seconds, automatically restarting any that have crashed or exited, while a separate config watcher handles live reloading of [`watch.toml`](https://github.com/tirth8205/code-review-graph/blob/main/watch.toml) changes.**

The **multi-repo watch daemon** (`WatchDaemon`) is the core process orchestration component of the [`tirth8205/code-review-graph`](https://github.com/tirth8205/code-review-graph) repository. It manages a fleet of per-repository watcher subprocesses, ensuring continuous code analysis even when individual workers fail. Understanding how its health check and auto-restart mechanisms function is essential for operating a reliable multi-repo code review pipeline.

## Health-Checker Thread Architecture

The daemon's resilience stems from three cooperating components: a **health-checker thread**, a **config watcher**, and **PID persistence**. These work together to maintain process availability and configuration consistency.

### Component Overview

| Component | Purpose | Key Location |
|-----------|---------|--------------|
| Health-checker thread | Detects dead watchers and restarts them | [`WatchDaemon._health_loop`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/daemon.py#L97) |
| Config watcher | Reacts to [`watch.toml`](https://github.com/tirth8205/code-review-graph/blob/main/watch.toml) modifications | [`ConfigWatcher.start`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/daemon.py#L55) |
| State persistence | Enables cross-process status queries | [`WatchDaemon._save_state`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/daemon.py#L122) |

## Starting the Health-Checker

When `WatchDaemon.start()` is invoked, it spawns the health-checker as a daemon thread. In [`daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/daemon.py), the `start_health_checker` method initializes the monitoring infrastructure:

```python
def start_health_checker(self) -> None:
    self._health_stop = threading.Event()
    self._health_thread = threading.Thread(
        target=self._health_loop,
        daemon=True,
        name="health-checker",
    )
    self._health_thread.start()
    logger.info(
        "Health checker started (interval=%ds)",
        _HEALTH_CHECK_INTERVAL,
    )

```

The `_HEALTH_CHECK_INTERVAL` constant defaults to **30 seconds** (defined at line 100). Using a `threading.Event` for `_health_stop` ensures the thread can be interrupted promptly during daemon shutdown rather than blocking on a fixed sleep.

## The Health-Check Loop

The `_health_loop` method implements an interruptible polling pattern:

```python
def _health_loop(self) -> None:
    while not self._health_stop.is_set():
        self._health_stop.wait(_HEALTH_CHECK_INTERVAL)
        if self._health_stop.is_set():
            break
        self._check_health()

```

This design provides two advantages:

- **Responsive shutdown**: `wait(timeout)` returns immediately when `set()` is called, avoiding fixed delays during termination
- **Regular health checks**: Falls through to `_check_health()` every 30 seconds when running normally

## Detecting and Restarting Dead Watchers

The core detection logic resides in `_check_health` at line 104 of [`daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/daemon.py):

```python
def _check_health(self) -> None:
    restarted = False
    with self._lock:
        for alias, repo in list(self._current_repos.items()):
            proc = self._children.get(alias)
            if proc is None or proc.poll() is not None:   # <-- dead process detection

                logger.warning("Watcher for '%s' is dead — restarting", alias)
                self._children.pop(alias, None)
                self._start_watcher(repo)                 # <-- auto-restart

                restarted = True
    if restarted:
        self._save_state()

```

**Process liveness detection** relies on `subprocess.Popen.poll()`, which returns `None` for running processes and an exit code integer for terminated ones. The implementation:

1. Acquires `_lock` to coordinate with config changes
2. Iterates over all tracked repositories
3. Removes stale entries from `_children`
4. Spawns fresh watchers via `_start_watcher(repo)`
5. Persists updated PIDs to [`daemon-state.json`](https://github.com/tirth8205/code-review-graph/blob/main/daemon-state.json) if any restarts occurred

The `_save_state` call ensures that external CLI commands like `crg status` can accurately report current process states even when executed from different shell sessions.

## Auto-Restart on Configuration Changes

Beyond crash recovery, the daemon handles **live configuration reloading**. The `ConfigWatcher` component (started at line 55) monitors [`watch.toml`](https://github.com/tirth8205/code-review-graph/blob/main/watch.toml) for modifications using platform-specific file system notifications or polling fallback.

When changes are detected, `_on_config_change` (line 152) triggers:

```python
def _on_config_change(self) -> None:
    new_config = load_config()
    self.reconcile(new_config)  # Computes diff and starts/stops watchers accordingly

```

The `reconcile` method computes the symmetric difference between desired and current repository sets:

- **Added repos**: Immediately launch new watcher processes
- **Removed repos**: Terminate associated child processes cleanly
- **Modified repos**: Restart watchers to pick up new paths or settings

This mechanism operates **independently yet compatibly** with the health-checker—newly added repositories become subject to health monitoring automatically, while removals clean up both running processes and monitoring entries.

## Cross-Process Liveness and State Recovery

The daemon maintains durable state across potential crashes through two persistence mechanisms:

- **`daemon.pid`**: Contains the main daemon's PID for singleton enforcement and `crg daemon status` queries
- **[`daemon-state.json`](https://github.com/tirth8205/code-review-graph/blob/main/daemon-state.json)**: Maps repository aliases to active watcher PIDs

The helper function `pid_alive` abstracts OS-specific liveness checks:

- **POSIX**: `os.kill(pid, 0)` — sends null signal to test existence
- **Windows**: `OpenProcess` with `PROCESS_QUERY_INFORMATION` access

This abstraction guarantees reliable **stale PID detection** when the daemon restarts after unexpected termination.

## Practical Examples

### Running the Daemon in Foreground Mode

```python
from code_review_graph.daemon import WatchDaemon

daemon = WatchDaemon()
daemon.start()            # spawns watchers and health-checker

daemon.run_forever()      # blocks until interrupted

```

The daemon reads `~/.code-review-graph/watch.toml`, launches a subprocess per configured repository, and begins health monitoring immediately.

### Simulating Watcher Failure and Auto-Restart

```python
import os, signal, time
from code_review_graph.daemon import WatchDaemon

daemon = WatchDaemon()
daemon.start()
time.sleep(2)  # Allow watcher spawn

# Force-kill the first watcher child

child_pid = next(iter(daemon._children.values())).pid
print(f"Killing watcher {child_pid}")
os.kill(child_pid, signal.SIGKILL)

# Wait for health-checker detection (>30s)

time.sleep(35)

new_pid = daemon._children[list(daemon._children)[0]].pid
assert new_pid != child_pid, "Watcher was not restarted"
print(f"Auto-restarted: new PID {new_pid}")

```

### Live Configuration Update

```python
from code_review_graph.daemon import add_repo_to_config

# Modify configuration on disk

add_repo_to_config("/path/to/new/repo", alias="new-project")

# ConfigWatcher detects change within seconds and calls reconcile,

# starting a watcher for "new-project" without daemon restart

```

## Daemonization and Platform Behavior

For production deployment, `daemon.daemonize()` performs a **double-fork on Unix systems**:

```python
daemon = WatchDaemon()
daemon.daemonize()  # Detaches from terminal, writes PID file

daemon.start()
daemon.run_forever()

```

On Windows, the method logs a warning and runs in foreground mode since `fork()` semantics are unavailable. The health-checker and config watcher function identically in both modes.

## Test Coverage

The implementation is exercised in [`tests/test_daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_daemon.py), which validates:

- PID file creation and cleanup
- Health-checker thread lifecycle
- Config reload and reconcile operations
- Signal handling for graceful shutdown
- State file consistency across restarts

Additional thread-spawning patterns appear in [`tests/test_incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_incremental.py), demonstrating `daemon=True` thread configuration and health-check assertions.

## Summary

- **Health-checker thread** polls every 30 seconds via `_health_loop`, detecting crashed watchers with `proc.poll()`
- **Auto-restart** immediately spawns replacement processes and updates [`daemon-state.json`](https://github.com/tirth8205/code-review-graph/blob/main/daemon-state.json) for external visibility
- **Config watcher** enables live reloading of [`watch.toml`](https://github.com/tirth8205/code-review-graph/blob/main/watch.toml) without daemon restart through `reconcile` diffing
- **PID persistence** via JSON state file and platform-abstracted `pid_alive` ensures accurate cross-process status reporting
- **Thread-safety** through `_lock` coordination prevents races between health checks and configuration changes

## Frequently Asked Questions

### How often does the daemon check watcher health?

The health-checker thread runs every **30 seconds** by default, as defined by `_HEALTH_CHECK_INTERVAL` in [`daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/daemon.py). This interval balances responsiveness with system overhead for typical code review workloads.

### What happens if the daemon itself crashes?

Child watcher processes become orphaned but continue running. When the daemon restarts, it reads [`daemon-state.json`](https://github.com/tirth8205/code-review-graph/blob/main/daemon-state.json) and uses `pid_alive` to determine which watchers are still active. New watchers are started only for missing or stale entries, preserving continuity where possible.

### Can I change the health-check interval?

The interval is currently hardcoded as `_HEALTH_CHECK_INTERVAL = 30` at line 100 of [`daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/daemon.py). Modifying this requires editing the source or submitting a pull request to make it configurable via [`watch.toml`](https://github.com/tirth8205/code-review-graph/blob/main/watch.toml).

### Does auto-restart preserve watcher state?

Individual watcher processes are **stateless by design**—they re-scan repository state on startup. The daemon maintains no in-memory state about watcher progress, so restarts are safe and idempotent. Any incremental tracking relies on persistent storage outside the daemon (e.g., commit hashes in a database).