# How the Real-Time Updates Feature Works in Code-Graph-RAG: Architecture and Implementation

> Discover how Code-Graph-RAG's real-time updates feature synchronizes its knowledge graph with live repository changes using file-system watchers and an incremental pipeline for efficient updates.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: architecture
- Published: 2026-08-19

---

**Code-Graph-RAG keeps its knowledge graph synchronized with live repository changes through a file-system watcher that applies hybrid debouncing and an incremental five-step pipeline to update only affected files without full re-indexing.**

The Code-Graph-RAG project (`vitali87/code-graph-rag`) maintains a dynamic knowledge graph of your codebase that stays current as developers edit source files. Instead of expensive full-repository rescans, the real-time updates feature captures file-system events and applies surgical graph mutations. This article examines the exact implementation from the source code, covering the watchdog-based observer in [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py) and the transactional update logic in [`graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_updater.py).

## File-System Watcher and Event Handling

The entry point for real-time synchronization lives in [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py), which sets up a **watchdog** observer on the repository root. The `CodeChangeEventHandler` class (lines [62‑74]) processes every `created`, `modified`, and `deleted` event by first filtering irrelevant paths (ignoring specific patterns and suffixes) and then feeding valid events into the debounce pipeline.

When the handler detects a relevant change, it stores the event in an internal `pending_events` dictionary and records the `first_event_time` timestamp. This data structure enables the hybrid debounce mechanism that balances responsiveness with efficiency.

## Hybrid Debounce Mechanics

To prevent the graph from flooding with redundant updates during rapid save sequences, the system implements a **hybrid debounce** strategy (configured in lines [66‑78] and implemented in lines [71‑84]).

When a change arrives, the handler checks if the elapsed time since `first_event_time` exceeds `max_wait_seconds`. If so, it forces immediate processing via `_schedule_immediate_processing`. Otherwise, it starts a timer with a delay calculated as `min(debounce_seconds, remaining_wait)`. The system cancels any existing timer for the same path before scheduling a new one, preventing duplicate processing. This guarantees that an update is processed **no later than `max_wait_seconds`** while still batching rapid edits.

Setting `--debounce 0` disables this behavior entirely and restores legacy immediate processing (see lines [88‑90] of the docstring).

## Incremental Graph Building Pipeline

Once the debouncer fires, the watcher creates a **`GraphUpdater`** instance (line [93] of [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py)) and delegates to `_process_change` (lines [124‑150]). This method runs inside a **single lock (`_update_lock`)** to serialize access to shared parser state (lines [99‑104]), ensuring thread-safe mutations.

The update executes a **five-step pipeline** for each affected file:

1. **Delete old graph data.** The handler executes Cypher queries `CYPHER_DELETE_MODULE` and `CYPHER_DELETE_FILE` to remove stale nodes and edges associated with the previous file state (lines [56‑68]).

2. **Clear in-memory caches.** The system evicts cached AST representations, import maps, and Rust path caches to prevent stale semantic data from corrupting the new graph state (lines [74‑84]).

3. **Re-parse the file.** For source files, the parser regenerates the AST and creates a generic `File` node. For non-source files, it updates the generic node metadata (lines [90‑107]).

4. **Reset and recompute `CALLS` edges.** The pipeline deletes old call edges, resets resolution caches, and invokes `_process_function_calls` to rebuild the call graph based on the new AST (lines [142‑148]).

5. **Flush the batch to the database.** The accumulated mutations are committed in a single transaction (line [150]).

## Language-Specific Frontend Consistency

The updater maintains semantic consistency for polyglot repositories by conditionally re-running language-specific frontends. When processing changes, `GraphUpdater.run` invokes `_run_go_frontend` and `_run_csharp_frontend` only when the changed file belongs to those languages. This ensures that Go `IMPLEMENTS` relationships and C# partial type groups remain accurate without wasting cycles on unrelated file types.

## Running the Real-Time Watcher

The CLI entry point resides in the `main` function of [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py) (lines [332‑363]). You can tune the debounce behavior to match your development workflow:

```bash
python -m codebase_rag.realtime_updater /path/to/repo \
    --host 127.0.0.1 --port 7687 \
    --debounce 5      # wait 5s after the last edit

    --max-wait 30     # force an update after 30s of continuous edits

```

The `--debounce` parameter controls the idle time required to trigger an update, while `--max-wait` sets an absolute ceiling to ensure changes never wait indefinitely during active editing sessions.

## Summary

- **File-system events** are captured by a watchdog observer in [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py) and filtered through the `CodeChangeEventHandler`.
- **Hybrid debouncing** batches rapid changes using `pending_events` and `max_wait_seconds` to balance latency with database load.
- **Thread safety** is enforced by `_update_lock`, which serializes access to parser state during incremental updates.
- **Five-step pipeline** deletes stale data, clears caches, re-parses files, rebuilds `CALLS` edges, and flushes transactions atomically.
- **Language-specific frontends** (Go, C#) run only when necessary to maintain semantic relationships without overhead.
- **CLI configuration** allows tuning via `--debounce` and `--max-wait` flags, with `0` disabling debounce for immediate processing.

## Frequently Asked Questions

### What triggers a real-time update in Code-Graph-RAG?

Any file-system event—`created`, `modified`, or `deleted`—within the watched repository triggers the `CodeChangeEventHandler`. The handler filters out irrelevant paths and file patterns before queuing the change for processing.

### How does the debounce mechanism prevent redundant updates?

The hybrid debouncer stores events in `pending_events` and waits for a pause in activity (`debounce_seconds`) or a hard timeout (`max_wait_seconds`), whichever comes first. It cancels duplicate timers for the same path, ensuring that rapid saves result in a single graph update rather than one per keystroke.

### Is the update process thread-safe?

Yes. The `_process_change` method acquires `_update_lock` before accessing shared parser state and cache structures. This lock serializes concurrent file changes, preventing race conditions during AST parsing and edge computation.

### Can I disable debouncing for immediate updates?

Yes. Setting the CLI argument `--debounce 0` disables the timer-based debouncing and forces the system to process file changes immediately, matching the legacy behavior prior to the hybrid debounce implementation.