# How to Update a Knowledge Graph for a Repository Using code-graph-rag

> Learn how to update the knowledge graph for a repository using code-graph-rag. The GraphUpdater reingest method efficiently re-parses only changed files for incremental updates.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-06

---

**The `GraphUpdater.reingest` method performs a scoped, incremental update of the knowledge graph by re-parsing only changed files and their dependents while preserving all untouched graph state.**

The **code-graph-rag** project maintains a structural and semantic representation of your codebase in a graph database. When source files change, you need an efficient way to synchronize the graph without rebuilding from scratch. This guide explains how to update a knowledge graph for a repository using the incremental re-ingestion pipeline implemented in `vitali87/code-graph-rag`.

## Understanding the GraphUpdater.reingest Architecture

The `reingest` method in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) orchestrates a 12-step pipeline that guarantees atomic, consistent updates. If any step fails before graph mutation, a `ReingestAborted` exception rolls back all changes—no partial writes ever reach the database.

### Step-by-Step Re-ingestion Flow

- **Mark incremental build** — Sets `self._is_full_build = False` in `GraphUpdater.__init__` (line 4495) to distinguish from full rebuilds
- **Split paths into three groups** — The `_reingest_split` method (line 4524) categorizes paths as **present** (on disk), **gone** (deleted), or **skipped** (matches ignore rules)
- **Hydrate from existing graph** — `_hydrate_for_reingest` (line 4538) loads current packages and folders for context
- **Load hash cache** — `_load_hash_cache` (line 777) reads `HASH_CACHE_FILENAME` to detect actual file changes efficiently
- **Resolve stem-flux** — `_reingest_flux_survivors` (line 4690) handles cases like header files being added/removed alongside implementations
- **Seed module Q-N map** — `_seed_module_qns_from_graph` (line 4770) preserves qualified names for unchanged files to prevent name theft from new siblings
- **Compute dependents** — `_reingest_dependents` (line 4820) finds one-hop inbound edges to determine minimal re-parse set
- **Capture inbound edges** — `_capture_inbound_edges` (line 4850) snapshots the subgraph that will be rewritten
- **Delete and re-parse** — `_reingest_delete` (line 4960) removes stale nodes; `_reingest_reparse` (line 4975) generates fresh definitions, imports, and calls
- **Resolve calls** — `_reingest_resolve` (line 5000) links calls only for re-parsed files, preserving untouched inbound edges
- **Update hash cache** — `_update_hashes` (line 5030) atomically writes new file hashes
- **Emit ReingestReport** — Returns elapsed time plus counts of reparsed, affected, removed, and skipped files

## How to Update a Knowledge Graph Programmatically

Use `GraphUpdater.reingest` directly when integrating into build scripts, CI pipelines, or custom tooling.

```python
from pathlib import Path
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.services import IngestorProtocol

# 1. Create your ingestor (e.g., Neo4jIngestor)

ingestor: IngestorProtocol = ...

# 2. Instantiate the updater with parsers and queries for your languages

updater = GraphUpdater(
    ingestor=ingestor,
    repo_path=Path("/path/to/your/repo"),
    parsers={"python": python_parser, "cpp": cpp_parser},
    queries={"python": python_queries, "cpp": cpp_queries},
)

# 3. Re-ingest specific changed files

changed = [Path("src/module/foo.py"), Path("src/module/bar.py")]
report = updater.reingest(changed)

# 4. Inspect the results

print("Re-parsed:", report.reparsed)
print("Affected (dependents):", report.affected)
print("Removed:", report.removed)
print("Skipped (ignored):", report.skipped)
print("Elapsed ms:", round(report.elapsed_ms, 1))

```

The `reingest` method accepts `Path` objects that can be absolute or relative to `repo_path`. It automatically expands the input set to include dependents—files that import from or call into your changed files—ensuring the graph remains internally consistent.

## How to Update a Knowledge Graph via CLI

For manual updates or shell integrations, use the `cgr` CLI wrapper defined in [`codebase_rag/graph_cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_cli.py):

```bash
cgr reingest --repo /path/to/repo \
    src/module/foo.py src/module/bar.py

```

The CLI constructs the same `GraphUpdater` instance internally, making it equivalent to the programmatic approach. Pass multiple paths as positional arguments; the tool handles path normalization and ignore-rule filtering automatically.

## Key Source Files for Knowledge Graph Updates

Understanding these files helps when debugging or extending re-ingestion behavior:

- **[`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py)** — Contains `GraphUpdater` class with `reingest` entry point and all 12 pipeline steps
- **[`codebase_rag/constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants.py)** — Defines `HASH_CACHE_FILENAME` and log message keys used throughout re-ingestion
- **[`codebase_rag/services.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services.py)** — Abstract `IngestorProtocol` and query interfaces that `GraphUpdater` reads and writes
- **[`codebase_rag/graph_cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_cli.py)** — Command-line frontend that parses arguments and invokes `reingest`
- **[`codebase_rag/utils/path_utils.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/path_utils.py)** — Path normalization, ignore-rule evaluation, and stem-based qualified name utilities

## Safety Guarantees in reingest

The `reingest` implementation provides two critical consistency guarantees:

1. **Atomic failure handling** — If `_hydrate_for_reingest`, `_capture_inbound_edges`, or any pre-mutation step fails, `ReingestAborted` prevents any graph changes
2. **Minimal re-parse scope** — Only changed files, their dependents, and stem-flux survivors are processed; untouched files retain all inbound and outbound edges exactly as before

This design makes `reingest` safe to run frequently—even on every commit—without degrading graph performance or integrity.

## Summary

- Call **`GraphUpdater.reingest`** with changed file paths to perform incremental knowledge graph updates
- The method automatically computes **dependents** and handles **stem-flux** without manual intervention
- Use the **`cgr reingest`** CLI for command-line workflows
- All changes are **atomic**: failures abort with `ReingestAborted` before any graph mutation occurs
- Core implementation lives in **[`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py)** with supporting utilities in [`constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/constants.py), [`services.py`](https://github.com/vitali87/code-graph-rag/blob/main/services.py), and [`path_utils.py`](https://github.com/vitali87/code-graph-rag/blob/main/path_utils.py)

## Frequently Asked Questions

### What happens if a file is deleted from the repository?

The `_reingest_split` method detects deletions and routes them to `_reingest_delete`, which removes stale nodes from the graph. The hash cache is updated to remove the deleted file's entry, ensuring subsequent runs don't attempt to process it.

### How does reingest avoid re-parsing unchanged files?

The system maintains a file hash cache in `HASH_CACHE_FILENAME` (defined in [`constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/constants.py)). During `_load_hash_cache`, current file hashes are compared against cached values. Only files with mismatched hashes—or those with changed dependents—enter the re-parse set.

### Can reingest handle multiple programming languages in one call?

Yes. The `GraphUpdater` constructor accepts `parsers` and `queries` dictionaries mapping language identifiers to tree-sitter parsers and compiled query objects. The `reingest` method dispatches each file to its matching parser based on extension or path rules.

### What is "stem-flux" and why does it matter?

Stem-flux occurs when related files (like [`foo.h`](https://github.com/vitali87/code-graph-rag/blob/main/foo.h) and [`foo.cpp`](https://github.com/vitali87/code-graph-rag/blob/main/foo.cpp)) change together. The `_reingest_flux_survivors` step identifies these relationships to prevent qualified name collisions and ensure that adding or removing companion files doesn't corrupt the module namespace of existing code.