How the Multi‑Repo Daemon Manages Multiple Projects in code‑review‑graph

The multi‑repo daemon in code_review_graph/daemon.py manages multiple projects by spawning a separate child process for each repository, persisting state to JSON, health‑checking every 30 seconds, and automatically reconciling configuration changes.

The multi‑repo daemon is the core architectural component that enables code-review-graph to monitor many Git or SVN repositories simultaneously. Instead of running a single monolithic process, the daemon delegates each repository to its own isolated watcher, ensuring failure isolation and independent update cycles. This article examines the daemon's layered design, from configuration management to process supervision, based on the source implementation in tirth8205/code-review-graph.

Configuration and Repository Representation

The daemon's behavior is driven by a TOML configuration file that declaratively specifies which repositories to watch.

DaemonConfig and WatchRepo Dataclasses

Two dataclasses define the configuration schema. In code_review_graph/daemon.py, the WatchRepo class normalizes each repository entry into a structured object containing the repository path and an optional alias:

@dataclass
class WatchRepo:
    path: str
    alias: Optional[str] = None

The DaemonConfig dataclass aggregates these entries alongside global settings: session name, polling interval, and log location. The default configuration lives at ~/.code-review-graph/watch.toml, returned by default_config_path() (lines 43–46). Utility functions load_config() and save_config() handle serialization.

Live Configuration Watching

The ConfigWatcher class monitors watch.toml for edits without requiring a daemon restart. When the optional watchdog package is unavailable, it transparently falls back to file-system polling (ConfigWatcher.start(), lines 82–88). Upon detecting a change, the watcher triggers WatchDaemon._on_config_change(), which reloads the configuration and invokes reconciliation logic.

Process Management Architecture

The daemon's supervision model relies on child process isolation—each repository runs its own code-review-graph watch process.

Spawning Watchers with _start_watcher()

The private method WatchDaemon._start_watcher() (lines 382–425) creates a subprocess.Popen instance for each repository. Each child receives:

  • A dedicated log file for isolated troubleshooting
  • The repository path and alias as arguments
  • Independent stdout/stderr streams

This design prevents a crash or hang in one repository's graph builder from affecting others.

PID Tracking and State Persistence

To enable cross-process status queries, the daemon persists child PIDs to daemon-state.json via _save_state(). The default state path is provided by default_state_path() (lines 53–56). When the CLI runs code_review_graph/cli.py status commands from a separate process, it reads this JSON file and validates each PID with pid_alive() rather than requiring IPC to the daemon.

PID File for Singleton Guarantees

The daemon ensures only one instance per user through write_pid(), read_pid(), and is_daemon_running() (lines 80–101). The PID file daemon.pid is written on startup; subsequent invocations detect the existing process and exit or attach accordingly.

Health Checking and Automatic Recovery

A background thread maintains daemon reliability through periodic supervision.

The _health_loop() Mechanism

The constant _HEALTH_CHECK_INTERVAL = 30 configures a 30-second check cycle. The _health_loop() method (lines 744–761) delegates to _check_health(), which:

  1. Iterates over all tracked PIDs in daemon-state.json
  2. Verifies process liveness with pid_alive()
  3. Automatically restarts dead watchers via _start_watcher()
  4. Rewrites the state file with updated PIDs

This self-healing behavior ensures repository monitoring survives transient failures without manual intervention.

Dynamic Reconciliation of Repository Sets

When configuration changes occur, the daemon must align running processes with the desired state.

The reconcile() Method

WatchDaemon.reconcile() (lines 224–292) computes a set diff between:

  • Desired state: Repositories listed in the current DaemonConfig
  • Actual state: PIDs currently tracked in memory and on disk

The method then applies the minimal set of operations:

  • Start new repositories not currently watched
  • Stop removed repositories (terminates their child processes)
  • Restart repositories whose configuration (e.g., alias) has changed

This diff-based approach avoids unnecessary process churn during minor edits.

Daemonization and Platform Handling

The WatchDaemon.daemonize() method (lines 724–782) handles platform-specific background execution.

Unix Double-Fork

On Unix systems, the daemon:

  1. First fork: Detaches from the controlling terminal
  2. Second fork: Prevents reacquisition of a controlling terminal
  3. Redirects stdin, stdout, stderr to daemon.log
  4. Changes working directory to /

Windows Foreground with Warning

On Windows, where true daemonization is unsupported, daemonize() issues a warning and continues in the foreground. Production deployments on Windows typically use service wrappers or scheduled tasks.

Complete Workflow Example

The following demonstrates starting the daemon, adding a repository, and querying status programmatically:

from code_review_graph.daemon import WatchDaemon, add_repo_to_config, is_daemon_running

# Start as background service (Unix)

daemon = WatchDaemon()
daemon.daemonize()
daemon.start()           # Spawns watcher per repo in watch.toml

daemon.run_forever()     # Blocks until signal termination

# Add repository while daemon runs

cfg = add_repo_to_config(
    repo_path="/home/dev/projects/new-service",
    alias="newservice",
)

# Daemon detects change via ConfigWatcher and reconciles within polling interval

# Query status from separate CLI process

if is_daemon_running():
    daemon = WatchDaemon()
    status = daemon.status()  # Reads daemon-state.json, checks PIDs

    for repo, info in status.items():
        print(f"{repo}: {'alive' if info['alive'] else 'dead'}")

Summary

  • Configuration-driven: Repositories are declared in ~/.code-review-graph/watch.toml and mapped to WatchRepo objects
  • Process-per-repo: WatchDaemon._start_watcher() isolates each repository in its own subprocess with dedicated logging
  • State persistence: Child PIDs are written to daemon-state.json enabling cross-process status queries
  • Health supervision: _health_loop() checks every 30 seconds and restarts failed watchers automatically
  • Live reconciliation: reconcile() adapts running processes to configuration changes without full restarts
  • Singleton enforcement: PID file prevents duplicate daemon instances
  • Platform portability: Double-fork on Unix, foreground with warning on Windows

Frequently Asked Questions

How does the daemon handle configuration file changes while running?

The ConfigWatcher class monitors watch.toml using watchdog when available, falling back to polling otherwise. When a change is detected, _on_config_change() reloads the configuration and invokes reconcile() to start, stop, or restart child processes as needed. This ensures the running state matches the declared configuration without requiring a manual daemon restart.

What happens if a repository watcher process crashes?

The health checker thread running _health_loop() detects crashed processes during its 30-second cycle. Dead PIDs trigger automatic restart via _start_watcher(), and the updated process information is persisted to daemon-state.json. This self-healing design requires no user intervention for transient failures.

Can I query daemon status from a different terminal or script?

Yes. The is_daemon_running() function checks the PID file, and WatchDaemon.status() reads daemon-state.json to validate each child PID with pid_alive(). This state-file architecture allows any process with filesystem access to determine daemon health without requiring active communication with the daemon itself.

Why does the daemon spawn separate processes instead of using threads?

Process isolation provides failure containment—a crash, infinite loop, or memory leak in one repository's graph builder cannot affect others. Additionally, Python's Global Interpreter Lock (GIL) limits true parallelism for CPU-bound graph construction; separate processes bypass this constraint and enable independent I/O scheduling per repository.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →