How Graphify's File Watch Mode Detects and Processes File Changes

Graphify's file watch mode uses a three-stage pipeline—detecting changes via filesystem events, queuing work through a persistent pending-changes file when rebuilds are active, and performing incremental or full rebuilds while holding a per-repository advisory lock to prevent race conditions.

The file watch mode in Graphify serves as the real-time engine that keeps your knowledge graph synchronized with code changes. Implemented in the Graphify-Labs/graphify repository, this system resides primarily in graphify/watch.py and orchestrates filesystem monitoring, work scheduling, and graph reconstruction without requiring manual intervention.

Three-Stage Architecture of Graphify's Watch Mode

Graphify's watch mode operates through a coordinated sequence of detection, queuing, and rebuilding phases. Each stage uses specific functions and persistence mechanisms to ensure no changes are lost during rapid editing sessions or concurrent operations.

Change Detection via Watchdog Events

The entry point watch() (exposed as graphify.watch.watch) initializes a watchdog observer when the optional watchdog package is installed. This observer monitors the watch directory for filesystem events including create, modify, delete, and move operations.

When an event fires, the handler invokes _changed_path_candidates() (lines 56-78 in graphify/watch.py) to normalize the affected path. This function performs critical transformations:

  • Converts relative paths to absolute forms
  • Resolves symbolic links to their targets
  • Generates both lexical and resolved path variants

This normalization supports diverse input sources, from Git hooks supplying relative paths to the watchdog providing absolute filesystem paths.

Work Queuing During Active Rebuilds

If a rebuild is already in progress, the system must prevent change loss while maintaining data integrity. Graphify implements a per-repo advisory lock (_rebuild_lock) using fcntl.flock (lines 31-45) to ensure exclusive rebuild access.

When the lock is held, incoming changes are not processed immediately. Instead, _queue_pending() (lines 18-38) appends the normalized paths to a .pending_changes file in the repository root. This file write is atomic on POSIX systems because each path is written in a separate write() call, guaranteeing that rapid successive edits are captured even if the process receives them faster than rebuilds can execute.

Incremental and Full Rebuild Execution

Once the advisory lock becomes available, the rebuild phase begins. The system first drains any queued work through _drain_pending() (lines 40-70), which reads the .pending_changes file and clears it atomically. These drained paths are merged with current event paths via _merge_changed_paths() (lines 110-128).

The core rebuild logic resides in _rebuild_code() (lines 66-188), which executes several specialized steps:

  1. Directory stabilization via _stabilize_rebuild_cwd() to ensure consistent working directory context
  2. Corpus detection using detect() while re-applying persisted --exclude patterns through _read_build_excludes()
  3. Selective extraction - if changed_paths is provided, only those files are processed; otherwise the full corpus is analyzed
  4. Graph reconciliation through _reconcile_existing_graph(), which merges new AST data with the existing graph structure
  5. Post-processing including clustering, scoring, and report generation
  6. Artifact writing for graph.json, GRAPH_REPORT.md, optional graph.html, and label files

Deletion detection occurs via _add_deleted_source() inside _rebuild_code(), which evicts removed files from the graph during incremental updates.

Key Mechanisms Ensuring Reliability

Several architectural decisions make Graphify's watch mode robust for production use.

Advisory Locking with fcntl: The _rebuild_lock() implementation prevents concurrent rebuild processes from corrupting the graph structure or writing conflicting artifacts. Only one process holds the lock at any time, and subsequent invocations wait or queue their changes.

Pending Changes Persistence: The .pending_changes file acts as a durable queue. Unlike in-memory queues that crash with the process, this file survives restarts, ensuring that changes made during a temporary watchdog outage are still processed on recovery.

Path Normalization Strategy: _changed_path_candidates() handles the complexity of filesystem path variations. By returning both absolute and resolved forms, it accommodates tools that reference files through symlinks while maintaining canonical identifiers for the graph database.

Incremental vs. Full Rebuild Logic: The _rebuild_code() function accepts an optional changed_paths parameter. When provided, the system bypasses full corpus detection and targets only the specified files, significantly reducing computation time for large repositories during active development.

Running Graphify in Watch Mode

To start monitoring a directory for changes, use the watch command with the target path:


# Watch the current directory (default behavior)

graphify watch .

# Run in background with debounce for CI environments

nohup graphify watch . --debounce 0.5 &

When you save a Python file within the watched tree, the watchdog handler triggers watch() → _rebuild_code(changed_paths=[Path("my_module.py")]). The system acquires the advisory lock, drains any pending changes, and re-extracts only my_module.py. The resulting AST is merged with the existing graph, and updated artifacts are written to disk.

Summary

  • Graphify's file watch mode is implemented in graphify/watch.py and uses the watchdog library for filesystem monitoring.
  • Change detection normalizes paths through _changed_path_candidates() to handle symlinks and relative path variants.
  • Work queuing persists changes to .pending_changes when _rebuild_lock is held, using atomic writes for POSIX safety.
  • Rebuild logic in _rebuild_code() supports both incremental updates (processing only changed_paths) and full corpus rebuilds.
  • Dependency modules include graphify/detect.py for file scanning, graphify/extract.py for AST parsing, and graphify/build.py for graph construction.

Frequently Asked Questions

How does Graphify prevent duplicate rebuilds when multiple files change simultaneously?

Graphify uses a per-repository advisory lock implemented via fcntl.flock in _rebuild_lock(). When a rebuild is active, subsequent change events are not lost but are instead appended to the .pending_changes file via _queue_pending(). Once the current rebuild completes and releases the lock, the next iteration drains these pending changes through _drain_pending() and processes them in a single batch, preventing redundant rebuild cycles.

What happens to file changes if the Graphify watch process crashes?

Changes detected during a crash are preserved in the .pending_changes file. Because Graphify writes each path to this file using atomic write() calls before attempting to acquire the rebuild lock, the queue survives process restarts. When the watch mode restarts and acquires the advisory lock, _drain_pending() reads and clears this file, ensuring no filesystem events are lost due to temporary outages.

Does Graphify support incremental rebuilds or does it reprocess the entire codebase?

Graphify supports incremental rebuilds through the changed_paths parameter in _rebuild_code(). When the watchdog detects specific file modifications, it passes these paths to the rebuild function, which bypasses full corpus detection and extracts AST data only for the changed files. The new data is then merged with the existing graph via _reconcile_existing_graph(). Full rebuilds occur only when changed_paths is empty or during initial setup.

Which file extensions does Graphify monitor in watch mode?

The watch mode delegates file filtering to the detection logic in graphify/detect.py. While the watchdog observer receives all filesystem events, the _changed_path_candidates() function and subsequent _rebuild_code() logic utilize the detection module to identify supported file extensions. During the rebuild phase, _read_build_excludes() also re-applies any persisted --exclude patterns to filter the final candidate list before extraction.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →