How the Multi‑Repo Daemon Manages Multiple Projects in code‑review‑graph
The multi‑repo daemon in code_review_graph/daemon.py manages multiple projects by spawning a separate child process for each repository, persisting state to JSON, health‑checking every 30 seconds, and automatically reconciling configuration changes.
The multi‑repo daemon is the core architectural component that enables code-review-graph to monitor many Git or SVN repositories simultaneously. Instead of running a single monolithic process, the daemon delegates each repository to its own isolated watcher, ensuring failure isolation and independent update cycles. This article examines the daemon's layered design, from configuration management to process supervision, based on the source implementation in tirth8205/code-review-graph.
Configuration and Repository Representation
The daemon's behavior is driven by a TOML configuration file that declaratively specifies which repositories to watch.
DaemonConfig and WatchRepo Dataclasses
Two dataclasses define the configuration schema. In code_review_graph/daemon.py, the WatchRepo class normalizes each repository entry into a structured object containing the repository path and an optional alias:
@dataclass
class WatchRepo:
path: str
alias: Optional[str] = None
The DaemonConfig dataclass aggregates these entries alongside global settings: session name, polling interval, and log location. The default configuration lives at ~/.code-review-graph/watch.toml, returned by default_config_path() (lines 43–46). Utility functions load_config() and save_config() handle serialization.
Live Configuration Watching
The ConfigWatcher class monitors watch.toml for edits without requiring a daemon restart. When the optional watchdog package is unavailable, it transparently falls back to file-system polling (ConfigWatcher.start(), lines 82–88). Upon detecting a change, the watcher triggers WatchDaemon._on_config_change(), which reloads the configuration and invokes reconciliation logic.
Process Management Architecture
The daemon's supervision model relies on child process isolation—each repository runs its own code-review-graph watch process.
Spawning Watchers with _start_watcher()
The private method WatchDaemon._start_watcher() (lines 382–425) creates a subprocess.Popen instance for each repository. Each child receives:
- A dedicated log file for isolated troubleshooting
- The repository path and alias as arguments
- Independent stdout/stderr streams
This design prevents a crash or hang in one repository's graph builder from affecting others.
PID Tracking and State Persistence
To enable cross-process status queries, the daemon persists child PIDs to daemon-state.json via _save_state(). The default state path is provided by default_state_path() (lines 53–56). When the CLI runs code_review_graph/cli.py status commands from a separate process, it reads this JSON file and validates each PID with pid_alive() rather than requiring IPC to the daemon.
PID File for Singleton Guarantees
The daemon ensures only one instance per user through write_pid(), read_pid(), and is_daemon_running() (lines 80–101). The PID file daemon.pid is written on startup; subsequent invocations detect the existing process and exit or attach accordingly.
Health Checking and Automatic Recovery
A background thread maintains daemon reliability through periodic supervision.
The _health_loop() Mechanism
The constant _HEALTH_CHECK_INTERVAL = 30 configures a 30-second check cycle. The _health_loop() method (lines 744–761) delegates to _check_health(), which:
- Iterates over all tracked PIDs in
daemon-state.json - Verifies process liveness with
pid_alive() - Automatically restarts dead watchers via
_start_watcher() - Rewrites the state file with updated PIDs
This self-healing behavior ensures repository monitoring survives transient failures without manual intervention.
Dynamic Reconciliation of Repository Sets
When configuration changes occur, the daemon must align running processes with the desired state.
The reconcile() Method
WatchDaemon.reconcile() (lines 224–292) computes a set diff between:
- Desired state: Repositories listed in the current
DaemonConfig - Actual state: PIDs currently tracked in memory and on disk
The method then applies the minimal set of operations:
- Start new repositories not currently watched
- Stop removed repositories (terminates their child processes)
- Restart repositories whose configuration (e.g., alias) has changed
This diff-based approach avoids unnecessary process churn during minor edits.
Daemonization and Platform Handling
The WatchDaemon.daemonize() method (lines 724–782) handles platform-specific background execution.
Unix Double-Fork
On Unix systems, the daemon:
- First fork: Detaches from the controlling terminal
- Second fork: Prevents reacquisition of a controlling terminal
- Redirects stdin, stdout, stderr to
daemon.log - Changes working directory to
/
Windows Foreground with Warning
On Windows, where true daemonization is unsupported, daemonize() issues a warning and continues in the foreground. Production deployments on Windows typically use service wrappers or scheduled tasks.
Complete Workflow Example
The following demonstrates starting the daemon, adding a repository, and querying status programmatically:
from code_review_graph.daemon import WatchDaemon, add_repo_to_config, is_daemon_running
# Start as background service (Unix)
daemon = WatchDaemon()
daemon.daemonize()
daemon.start() # Spawns watcher per repo in watch.toml
daemon.run_forever() # Blocks until signal termination
# Add repository while daemon runs
cfg = add_repo_to_config(
repo_path="/home/dev/projects/new-service",
alias="newservice",
)
# Daemon detects change via ConfigWatcher and reconciles within polling interval
# Query status from separate CLI process
if is_daemon_running():
daemon = WatchDaemon()
status = daemon.status() # Reads daemon-state.json, checks PIDs
for repo, info in status.items():
print(f"{repo}: {'alive' if info['alive'] else 'dead'}")
Summary
- Configuration-driven: Repositories are declared in
~/.code-review-graph/watch.tomland mapped toWatchRepoobjects - Process-per-repo:
WatchDaemon._start_watcher()isolates each repository in its own subprocess with dedicated logging - State persistence: Child PIDs are written to
daemon-state.jsonenabling cross-process status queries - Health supervision:
_health_loop()checks every 30 seconds and restarts failed watchers automatically - Live reconciliation:
reconcile()adapts running processes to configuration changes without full restarts - Singleton enforcement: PID file prevents duplicate daemon instances
- Platform portability: Double-fork on Unix, foreground with warning on Windows
Frequently Asked Questions
How does the daemon handle configuration file changes while running?
The ConfigWatcher class monitors watch.toml using watchdog when available, falling back to polling otherwise. When a change is detected, _on_config_change() reloads the configuration and invokes reconcile() to start, stop, or restart child processes as needed. This ensures the running state matches the declared configuration without requiring a manual daemon restart.
What happens if a repository watcher process crashes?
The health checker thread running _health_loop() detects crashed processes during its 30-second cycle. Dead PIDs trigger automatic restart via _start_watcher(), and the updated process information is persisted to daemon-state.json. This self-healing design requires no user intervention for transient failures.
Can I query daemon status from a different terminal or script?
Yes. The is_daemon_running() function checks the PID file, and WatchDaemon.status() reads daemon-state.json to validate each child PID with pid_alive(). This state-file architecture allows any process with filesystem access to determine daemon health without requiring active communication with the daemon itself.
Why does the daemon spawn separate processes instead of using threads?
Process isolation provides failure containment—a crash, infinite loop, or memory leak in one repository's graph builder cannot affect others. Additionally, Python's Global Interpreter Lock (GIL) limits true parallelism for CPU-bound graph construction; separate processes bypass this constraint and enable independent I/O scheduling per repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →