How the code-review-graph File-Watcher Daemon Detects and Processes File Changes in Real-Time

The code-review-graph file-watcher daemon uses a two-stage pipeline: a filesystem observer (watchdog or polling fallback) generates raw events, which are batched, filtered, and processed by incremental_update() to refresh only affected graph nodes.

The code-review-graph open-source tool provides real-time code analysis through its watch subcommand and WatchDaemon orchestrator. Understanding how this file-watcher daemon detects and processes file changes reveals a robust architecture designed for incremental updates with minimal overhead.


File-System Observer Setup: watchdog vs. Polling Fallback

The watcher process in code_review_graph/daemon.py implements dual strategies for cross-platform filesystem monitoring.

Primary: watchdog Integration

When the watchdog library is available, ConfigWatcher.start() (lines 55–80 in code_review_graph/daemon.py) instantiates a watchdog.observers.Observer and registers a custom FileSystemEventHandler:


# From daemon.py - watcher initialization with watchdog

from watchdog.observers import Observer
from watchdog.events import FileSystemEventHandler

observer = Observer()
handler = _create_watch_handler(repo_root, store, on_files_updated)
observer.schedule(handler, str(repo_root), recursive=True)
observer.start()

Fallback: Polling Thread

If watchdog import fails, the code starts a polling thread that checks file modification times every poll_interval seconds. This ensures functionality on restricted environments without native filesystem event APIs.


Event Handling and Batching Pipeline

Raw filesystem events undergo normalization and filtering before triggering graph updates.

The Event Handler Factory

_create_watch_handler in code_review_graph/incremental.py (around line 1540) constructs handlers for create, modify, delete, and move operations:

Stage Function Purpose
Path normalization Handler methods Convert watchdog paths to repo-relative strings
Ignore filtering _should_ignore() Skip patterns from .gitignore and explicit exclusions
Binary detection _is_binary() Exclude non-source files from parsing
Temporal batching Internal queue Group events within 1-second window

Batching and Dispatch Logic

The handler aggregates rapid-fire changes into atomic batches:


# Conceptual flow from incremental.py

changed_files = []  # Collected within batch_window (default 1s)

def on_any_event(event):
    if _should_ignore(event.src_path) or _is_binary(event.src_path):
        return
    changed_files.append(normalize_path(event.src_path))
    

# When batch_window expires:

incremental_update(repo_root, store, changed_files=batched_files)
if on_files_updated:
    on_files_updated(batched_files)

Incremental Graph Update: incremental_update()

The core optimization lies in partial reprocessing rather than full rebuilds.

Dependency-Aware Reparsing

incremental_update() performs these operations:

  1. Parse changed files – Extract AST nodes and edges
  2. Resolve dependents – Identify files importing or referencing changed modules
  3. Update graph store – Write to SQLite backend with ACID guarantees
  4. Trigger callbacks – Execute optional post-processing hooks
from code_review_graph.incremental import watch, incremental_update
from code_review_graph.store import GraphStore
from pathlib import Path

repo_root = Path("/path/to/my/repo")
store = GraphStore(db_path="/tmp/graph.db")

# Direct API usage

watch(
    repo_root,
    store,
    poll_interval=2.0,  # Only used in fallback mode

    on_files_updated=lambda files: print(f"Processed: {files}"),
)

Daemon Orchestration: WatchDaemon Process Management

Per-repository isolation and fault tolerance come from WatchDaemon in code_review_graph/daemon.py.

Architecture Overview

  • One process per repo: code-review-graph watch runs as isolated child
  • PID persistence: State files enable crash recovery
  • Health monitoring: _health_loop (lines 998–1006) polls child status
  • Auto-restart: Dead watchers respawn automatically

CLI Control Interface


# Start daemon managing watchers for multiple repos

code-review-graph daemon start --repo-list /etc/crg/repos.conf

# Manual single-repo watcher (useful for development)

code-review-graph watch --repo /path/to/my/repo \
    --on-files-updated "./notify_team.sh"

# Daemon lifecycle management

code-review-graph daemon status
code-review-graph daemon stop
code-review-graph daemon restart

The _check_health method validates child process liveness and restarts any terminated watcher within seconds, ensuring continuous monitoring.


File Structure and Source References

File Lines Responsibility
code_review_graph/incremental.py ~1600 watch(), incremental_update(), event handlers, dependency resolution
code_review_graph/daemon.py ~1000+ WatchDaemon, ConfigWatcher, process management, health loops
code_review_graph/daemon_cli.py ~200 CLI entry points, argument parsing, daemonize support
tests/test_incremental.py ~400 Batching behavior, ignore rules, callback verification

Direct implementation links:


Summary

  • Detection layer uses watchdog with polling fallback for universal filesystem event capture
  • Processing layer batches events, filters irrelevant paths, and invokes incremental_update() for dependency-aware reparsing
  • Storage layer commits partial updates to SQLite without full graph rebuilds
  • Orchestration layer maintains isolated per-repo processes with automatic recovery via WatchDaemon._health_loop

The architecture balances real-time responsiveness with resource efficiency through batched incremental updates rather than continuous full reprocessing.


Frequently Asked Questions

What happens if the watchdog library is not installed?

The watcher falls back to a polling thread that checks file modification timestamps every poll_interval seconds (default 2.0). This poll-based mode is slower but works on any platform without native filesystem event APIs.

How does the daemon handle watcher process crashes?

WatchDaemon._health_loop monitors child process PIDs at regular intervals. When _check_health detects a dead watcher, it automatically spawns a replacement process and logs the recovery event. State persistence ensures the new process resumes from the correct repository configuration.

Can I run the watcher without the full daemon?

Yes. The code-review-graph watch command launches a standalone watcher for a single repository. This bypasses WatchDaemon entirely and is suitable for development workflows or CI pipelines where persistent process management is unnecessary.

What types of file changes trigger graph updates?

The handler processes create, modify, delete, and move events. However, changes are filtered through _should_ignore() (respecting .gitignore patterns) and _is_binary() (excluding non-text files) before invoking incremental_update(). Only source files that pass both filters trigger reparsing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →