How Multi-Repo Daemon Health Checks and Auto-Restart Work in Code Review Graph
The WatchDaemon in code-review-graph uses a dedicated health-checker thread to monitor child watcher processes every 30 seconds, automatically restarting any that have crashed or exited, while a separate config watcher handles live reloading of watch.toml changes.
The multi-repo watch daemon (WatchDaemon) is the core process orchestration component of the tirth8205/code-review-graph repository. It manages a fleet of per-repository watcher subprocesses, ensuring continuous code analysis even when individual workers fail. Understanding how its health check and auto-restart mechanisms function is essential for operating a reliable multi-repo code review pipeline.
Health-Checker Thread Architecture
The daemon's resilience stems from three cooperating components: a health-checker thread, a config watcher, and PID persistence. These work together to maintain process availability and configuration consistency.
Component Overview
| Component | Purpose | Key Location |
|---|---|---|
| Health-checker thread | Detects dead watchers and restarts them | WatchDaemon._health_loop |
| Config watcher | Reacts to watch.toml modifications |
ConfigWatcher.start |
| State persistence | Enables cross-process status queries | WatchDaemon._save_state |
Starting the Health-Checker
When WatchDaemon.start() is invoked, it spawns the health-checker as a daemon thread. In daemon.py, the start_health_checker method initializes the monitoring infrastructure:
def start_health_checker(self) -> None:
self._health_stop = threading.Event()
self._health_thread = threading.Thread(
target=self._health_loop,
daemon=True,
name="health-checker",
)
self._health_thread.start()
logger.info(
"Health checker started (interval=%ds)",
_HEALTH_CHECK_INTERVAL,
)
The _HEALTH_CHECK_INTERVAL constant defaults to 30 seconds (defined at line 100). Using a threading.Event for _health_stop ensures the thread can be interrupted promptly during daemon shutdown rather than blocking on a fixed sleep.
The Health-Check Loop
The _health_loop method implements an interruptible polling pattern:
def _health_loop(self) -> None:
while not self._health_stop.is_set():
self._health_stop.wait(_HEALTH_CHECK_INTERVAL)
if self._health_stop.is_set():
break
self._check_health()
This design provides two advantages:
- Responsive shutdown:
wait(timeout)returns immediately whenset()is called, avoiding fixed delays during termination - Regular health checks: Falls through to
_check_health()every 30 seconds when running normally
Detecting and Restarting Dead Watchers
The core detection logic resides in _check_health at line 104 of daemon.py:
def _check_health(self) -> None:
restarted = False
with self._lock:
for alias, repo in list(self._current_repos.items()):
proc = self._children.get(alias)
if proc is None or proc.poll() is not None: # <-- dead process detection
logger.warning("Watcher for '%s' is dead — restarting", alias)
self._children.pop(alias, None)
self._start_watcher(repo) # <-- auto-restart
restarted = True
if restarted:
self._save_state()
Process liveness detection relies on subprocess.Popen.poll(), which returns None for running processes and an exit code integer for terminated ones. The implementation:
- Acquires
_lockto coordinate with config changes - Iterates over all tracked repositories
- Removes stale entries from
_children - Spawns fresh watchers via
_start_watcher(repo) - Persists updated PIDs to
daemon-state.jsonif any restarts occurred
The _save_state call ensures that external CLI commands like crg status can accurately report current process states even when executed from different shell sessions.
Auto-Restart on Configuration Changes
Beyond crash recovery, the daemon handles live configuration reloading. The ConfigWatcher component (started at line 55) monitors watch.toml for modifications using platform-specific file system notifications or polling fallback.
When changes are detected, _on_config_change (line 152) triggers:
def _on_config_change(self) -> None:
new_config = load_config()
self.reconcile(new_config) # Computes diff and starts/stops watchers accordingly
The reconcile method computes the symmetric difference between desired and current repository sets:
- Added repos: Immediately launch new watcher processes
- Removed repos: Terminate associated child processes cleanly
- Modified repos: Restart watchers to pick up new paths or settings
This mechanism operates independently yet compatibly with the health-checker—newly added repositories become subject to health monitoring automatically, while removals clean up both running processes and monitoring entries.
Cross-Process Liveness and State Recovery
The daemon maintains durable state across potential crashes through two persistence mechanisms:
daemon.pid: Contains the main daemon's PID for singleton enforcement andcrg daemon statusqueriesdaemon-state.json: Maps repository aliases to active watcher PIDs
The helper function pid_alive abstracts OS-specific liveness checks:
- POSIX:
os.kill(pid, 0)— sends null signal to test existence - Windows:
OpenProcesswithPROCESS_QUERY_INFORMATIONaccess
This abstraction guarantees reliable stale PID detection when the daemon restarts after unexpected termination.
Practical Examples
Running the Daemon in Foreground Mode
from code_review_graph.daemon import WatchDaemon
daemon = WatchDaemon()
daemon.start() # spawns watchers and health-checker
daemon.run_forever() # blocks until interrupted
The daemon reads ~/.code-review-graph/watch.toml, launches a subprocess per configured repository, and begins health monitoring immediately.
Simulating Watcher Failure and Auto-Restart
import os, signal, time
from code_review_graph.daemon import WatchDaemon
daemon = WatchDaemon()
daemon.start()
time.sleep(2) # Allow watcher spawn
# Force-kill the first watcher child
child_pid = next(iter(daemon._children.values())).pid
print(f"Killing watcher {child_pid}")
os.kill(child_pid, signal.SIGKILL)
# Wait for health-checker detection (>30s)
time.sleep(35)
new_pid = daemon._children[list(daemon._children)[0]].pid
assert new_pid != child_pid, "Watcher was not restarted"
print(f"Auto-restarted: new PID {new_pid}")
Live Configuration Update
from code_review_graph.daemon import add_repo_to_config
# Modify configuration on disk
add_repo_to_config("/path/to/new/repo", alias="new-project")
# ConfigWatcher detects change within seconds and calls reconcile,
# starting a watcher for "new-project" without daemon restart
Daemonization and Platform Behavior
For production deployment, daemon.daemonize() performs a double-fork on Unix systems:
daemon = WatchDaemon()
daemon.daemonize() # Detaches from terminal, writes PID file
daemon.start()
daemon.run_forever()
On Windows, the method logs a warning and runs in foreground mode since fork() semantics are unavailable. The health-checker and config watcher function identically in both modes.
Test Coverage
The implementation is exercised in tests/test_daemon.py, which validates:
- PID file creation and cleanup
- Health-checker thread lifecycle
- Config reload and reconcile operations
- Signal handling for graceful shutdown
- State file consistency across restarts
Additional thread-spawning patterns appear in tests/test_incremental.py, demonstrating daemon=True thread configuration and health-check assertions.
Summary
- Health-checker thread polls every 30 seconds via
_health_loop, detecting crashed watchers withproc.poll() - Auto-restart immediately spawns replacement processes and updates
daemon-state.jsonfor external visibility - Config watcher enables live reloading of
watch.tomlwithout daemon restart throughreconcilediffing - PID persistence via JSON state file and platform-abstracted
pid_aliveensures accurate cross-process status reporting - Thread-safety through
_lockcoordination prevents races between health checks and configuration changes
Frequently Asked Questions
How often does the daemon check watcher health?
The health-checker thread runs every 30 seconds by default, as defined by _HEALTH_CHECK_INTERVAL in daemon.py. This interval balances responsiveness with system overhead for typical code review workloads.
What happens if the daemon itself crashes?
Child watcher processes become orphaned but continue running. When the daemon restarts, it reads daemon-state.json and uses pid_alive to determine which watchers are still active. New watchers are started only for missing or stale entries, preserving continuity where possible.
Can I change the health-check interval?
The interval is currently hardcoded as _HEALTH_CHECK_INTERVAL = 30 at line 100 of daemon.py. Modifying this requires editing the source or submitting a pull request to make it configurable via watch.toml.
Does auto-restart preserve watcher state?
Individual watcher processes are stateless by design—they re-scan repository state on startup. The daemon maintains no in-memory state about watcher progress, so restarts are safe and idempotent. Any incremental tracking relies on persistent storage outside the daemon (e.g., commit hashes in a database).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →