How Graphify Integrates with Git for Auto-Rebuild: Complete Implementation Guide
Graphify integrates with Git for auto-rebuild by using a post-commit hook that triggers a background watcher process, which employs advisory file locking and a pending-changes queue to ensure atomic, incremental, and lossless knowledge graph updates synchronized to specific commit hashes.
Graphify is an open-source knowledge graph generator that integrates with Git for auto-rebuild to keep documentation and code intelligence synchronized with your repository. According to the Graphify-Labs/graphify source code, the implementation combines Git hooks with advisory file locking and a pending-changes queue to ensure that every commit triggers an accurate, incremental update without dropping changes during rapid commit sequences.
Git Hook Integration: The Auto-Rebuild Trigger
The auto-rebuild pipeline starts with a Git hook. When developers run git commit, a post-commit hook invokes graphify watch in the background. This non-blocking approach allows the commit to complete immediately while Graphify handles graph regeneration asynchronously.
The watcher entry point is implemented in graphify/watch.py, which handles the orchestration of lock acquisition, change queueing, and the rebuild process.
Advisory Locking and Concurrency Control
To handle concurrent commits, Graphify implements an advisory per-repo lock. The _rebuild_lock function in graphify/watch.py creates a file lock on <out-dir>/.rebuild.lock using flock. The lock file stores the owning PID, allowing external pollers to check whether a rebuild is in progress.
This locking mechanism prevents race conditions when multiple commits occur in rapid succession, ensuring that only one rebuild process runs at a time while others wait or queue.
Change Queueing for Lossless Updates
When a hook cannot obtain the rebuild lock, Graphify queues changes instead of dropping them. The watcher appends changed paths to <out-dir>/.pending_changes using _queue_pending. When the lock holder finishes, it drains this queue using _drain_pending and merges changes via _merge_changed_paths.
This design guarantees atomic, idempotent, and lossless graph updates. Even when many commits happen in rapid succession, no commit is silently dropped because the system persists pending changes to disk.
Incremental vs. Full Rebuild Logic
The watcher distinguishes between incremental and full rebuilds in the _rebuild_code function. If the hook supplies a list of changed paths, Graphify only re-extracts those files using the changed_paths parameter. If the lock is acquired after queuing, the system merges queued changes with new changes and processes them in a single pass.
This optimization minimizes rebuild time for large repositories by avoiding unnecessary re-processing of unchanged files.
Commit Fingerprinting and Traceability
Every build embeds the exact Git revision for traceability. The _git_head() function in graphify/watch.py runs git rev-parse HEAD to capture the current commit hash, storing it as built_at_commit in both graph.json and GRAPH_REPORT.md. The report generation is handled by graphify/report.py.
This ties each knowledge graph version to a specific code state, providing complete auditability of which repository revision produced each graph artifact.
Core Implementation Files
The Git integration spans several key modules:
graphify/watch.py– Core watcher implementation containing_queue_pending,_drain_pending,_rebuild_lock,_git_head, and the main rebuild routine_rebuild_code.graphify/detect.py– Detects source files and respects.gitignoreand custom excludes; used by the watcher to identify valid repository paths.graphify/build.py– Converts extracted AST/semantic data into a graph structure after the watcher collects changed files.graphify/paths.py– Defines the output directory (GRAPHIFY_OUT) used for lock files, pending queues, and generated artifacts.graphify/report.py– GeneratesGRAPH_REPORT.mdand embeds the commit hash obtained via_git_head().
Setting Up the Post-Commit Hook
To enable auto-rebuild, place the following executable script in .git/hooks/post-commit within your repository:
#!/usr/bin/env python3
import subprocess
import sys
from pathlib import Path
# 1️⃣ Ensure we are inside a Git repo
try:
subprocess.run(["git", "rev-parse", "--show-toplevel"],
check=True, stdout=subprocess.DEVNULL)
except subprocess.CalledProcessError:
sys.exit(0) # not a repo → nothing to do
# 2️⃣ Run Graphify's watcher in the background (non-blocking)
watch_cmd = [
"graphify", "watch",
"--repo", "." # watch the current repo root
]
# The watch command itself handles queueing, locking and incremental rebuild.
subprocess.Popen(watch_cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
The script simply starts graphify watch; all heavy lifting—including change-set queuing, lock handling, and incremental extraction—is performed by the library code in graphify/watch.py.
Summary
- Graphify integrates with Git for auto-rebuild through a post-commit hook that triggers
graphify watchas a background process. - Advisory file locking via
_rebuild_lockon<out-dir>/.rebuild.lockprevents concurrent rebuild conflicts. - Lossless queueing using
<out-dir>/.pending_changesensures no commits are dropped during rapid commit sequences. - Incremental rebuilds process only changed files via
_rebuild_code(..., changed_paths=...)when possible, falling back to full rebuilds when necessary. - Commit fingerprinting embeds the Git HEAD hash into
graph.jsonandGRAPH_REPORT.mdfor full traceability.
Frequently Asked Questions
How does Graphify prevent lost commits during rapid-fire git commits?
When a post-commit hook cannot acquire the rebuild lock, it appends changed paths to <out-dir>/.pending_changes using _queue_pending instead of exiting silently. The current lock holder drains this file using _drain_pending and merges queued changes via _merge_changed_paths before processing, ensuring every commit triggers a graph update even when they occur in quick succession.
What happens if the Graphify rebuild process crashes while holding the lock?
The advisory lock in graphify/watch.py stores the owning PID in <out-dir>/.rebuild.lock. External pollers can detect stale locks by checking if the PID still exists. When a new watcher starts after a crash, it can identify and clear orphaned locks, allowing the next rebuild to proceed normally without indefinite blocking.
Can Graphify perform incremental rebuilds on specific files only?
Yes. When the hook supplies a list of changed paths, the _rebuild_code function in graphify/watch.py accepts a changed_paths parameter to limit extraction to only those files. If the system must drain a pending queue, it merges those paths with new changes to create a minimal, efficient rebuild set rather than reprocessing the entire repository.
Where does Graphify store the commit hash for traceability?
The _git_head() function captures the current commit hash and stores it as built_at_commit in both graph.json (the main graph output) and GRAPH_REPORT.md (the human-readable report). This is implemented in graphify/watch.py and graphify/report.py respectively, providing a permanent link between the knowledge graph and the exact Git revision built.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →