# How the code-review-graph File-Watcher Daemon Detects and Processes File Changes in Real-Time

> Discover how the code-review-graph file watcher daemon uses a two-stage pipeline with watchdog or polling to detect and process file changes in real-time, efficiently updating only affected nodes.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: internals
- Published: 2026-08-15

---

**The `code-review-graph` file-watcher daemon uses a two-stage pipeline: a filesystem observer (`watchdog` or polling fallback) generates raw events, which are batched, filtered, and processed by `incremental_update()` to refresh only affected graph nodes.**

The `code-review-graph` open-source tool provides real-time code analysis through its `watch` subcommand and `WatchDaemon` orchestrator. Understanding how this file-watcher daemon detects and processes file changes reveals a robust architecture designed for incremental updates with minimal overhead.

---

## File-System Observer Setup: watchdog vs. Polling Fallback

The watcher process in [`code_review_graph/daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/daemon.py) implements dual strategies for cross-platform filesystem monitoring.

### Primary: watchdog Integration

When the **watchdog** library is available, `ConfigWatcher.start()` (lines 55–80 in [`code_review_graph/daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/daemon.py)) instantiates a `watchdog.observers.Observer` and registers a custom `FileSystemEventHandler`:

```python

# From daemon.py - watcher initialization with watchdog

from watchdog.observers import Observer
from watchdog.events import FileSystemEventHandler

observer = Observer()
handler = _create_watch_handler(repo_root, store, on_files_updated)
observer.schedule(handler, str(repo_root), recursive=True)
observer.start()

```

### Fallback: Polling Thread

If `watchdog` import fails, the code starts a polling thread that checks file modification times every `poll_interval` seconds. This ensures functionality on restricted environments without native filesystem event APIs.

---

## Event Handling and Batching Pipeline

Raw filesystem events undergo normalization and filtering before triggering graph updates.

### The Event Handler Factory

`_create_watch_handler` in [`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py) (around line 1540) constructs handlers for create, modify, delete, and move operations:

| Stage | Function | Purpose |
|-------|----------|---------|
| Path normalization | Handler methods | Convert watchdog paths to repo-relative strings |
| Ignore filtering | `_should_ignore()` | Skip patterns from `.gitignore` and explicit exclusions |
| Binary detection | `_is_binary()` | Exclude non-source files from parsing |
| Temporal batching | Internal queue | Group events within 1-second window |

### Batching and Dispatch Logic

The handler aggregates rapid-fire changes into atomic batches:

```python

# Conceptual flow from incremental.py

changed_files = []  # Collected within batch_window (default 1s)

def on_any_event(event):
    if _should_ignore(event.src_path) or _is_binary(event.src_path):
        return
    changed_files.append(normalize_path(event.src_path))
    

# When batch_window expires:

incremental_update(repo_root, store, changed_files=batched_files)
if on_files_updated:
    on_files_updated(batched_files)

```

---

## Incremental Graph Update: `incremental_update()`

The core optimization lies in partial reprocessing rather than full rebuilds.

### Dependency-Aware Reparsing

`incremental_update()` performs these operations:

1. **Parse changed files** – Extract AST nodes and edges
2. **Resolve dependents** – Identify files importing or referencing changed modules
3. **Update graph store** – Write to SQLite backend with ACID guarantees
4. **Trigger callbacks** – Execute optional post-processing hooks

```python
from code_review_graph.incremental import watch, incremental_update
from code_review_graph.store import GraphStore
from pathlib import Path

repo_root = Path("/path/to/my/repo")
store = GraphStore(db_path="/tmp/graph.db")

# Direct API usage

watch(
    repo_root,
    store,
    poll_interval=2.0,  # Only used in fallback mode

    on_files_updated=lambda files: print(f"Processed: {files}"),
)

```

---

## Daemon Orchestration: WatchDaemon Process Management

Per-repository isolation and fault tolerance come from `WatchDaemon` in [`code_review_graph/daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/daemon.py).

### Architecture Overview

- **One process per repo**: `code-review-graph watch` runs as isolated child
- **PID persistence**: State files enable crash recovery
- **Health monitoring**: `_health_loop` (lines 998–1006) polls child status
- **Auto-restart**: Dead watchers respawn automatically

### CLI Control Interface

```bash

# Start daemon managing watchers for multiple repos

code-review-graph daemon start --repo-list /etc/crg/repos.conf

# Manual single-repo watcher (useful for development)

code-review-graph watch --repo /path/to/my/repo \
    --on-files-updated "./notify_team.sh"

# Daemon lifecycle management

code-review-graph daemon status
code-review-graph daemon stop
code-review-graph daemon restart

```

The `_check_health` method validates child process liveness and restarts any terminated watcher within seconds, ensuring continuous monitoring.

---

## File Structure and Source References

| File | Lines | Responsibility |
|------|-------|----------------|
| [`code_review_graph/incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py) | ~1600 | `watch()`, `incremental_update()`, event handlers, dependency resolution |
| [`code_review_graph/daemon.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/daemon.py) | ~1000+ | `WatchDaemon`, `ConfigWatcher`, process management, health loops |
| [`code_review_graph/daemon_cli.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/daemon_cli.py) | ~200 | CLI entry points, argument parsing, daemonize support |
| [`tests/test_incremental.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_incremental.py) | ~400 | Batching behavior, ignore rules, callback verification |

Direct implementation links:

- [daemon.py#L55](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/daemon.py#L55) – Observer initialization logic
- [incremental.py#L1540](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/incremental.py#L1540) – `watch()` function and handler setup
- [daemon.py#L998](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/daemon.py#L998) – Health check and restart mechanism

---

## Summary

- **Detection layer** uses `watchdog` with polling fallback for universal filesystem event capture
- **Processing layer** batches events, filters irrelevant paths, and invokes `incremental_update()` for dependency-aware reparsing
- **Storage layer** commits partial updates to SQLite without full graph rebuilds
- **Orchestration layer** maintains isolated per-repo processes with automatic recovery via `WatchDaemon._health_loop`

The architecture balances real-time responsiveness with resource efficiency through batched incremental updates rather than continuous full reprocessing.

---

## Frequently Asked Questions

### What happens if the watchdog library is not installed?

The watcher falls back to a polling thread that checks file modification timestamps every `poll_interval` seconds (default 2.0). This poll-based mode is slower but works on any platform without native filesystem event APIs.

### How does the daemon handle watcher process crashes?

`WatchDaemon._health_loop` monitors child process PIDs at regular intervals. When `_check_health` detects a dead watcher, it automatically spawns a replacement process and logs the recovery event. State persistence ensures the new process resumes from the correct repository configuration.

### Can I run the watcher without the full daemon?

Yes. The `code-review-graph watch` command launches a standalone watcher for a single repository. This bypasses `WatchDaemon` entirely and is suitable for development workflows or CI pipelines where persistent process management is unnecessary.

### What types of file changes trigger graph updates?

The handler processes create, modify, delete, and move events. However, changes are filtered through `_should_ignore()` (respecting `.gitignore` patterns) and `_is_binary()` (excluding non-text files) before invoking `incremental_update()`. Only source files that pass both filters trigger reparsing.