How to Update a Knowledge Graph for a Repository Using code-graph-rag
The GraphUpdater.reingest method performs a scoped, incremental update of the knowledge graph by re-parsing only changed files and their dependents while preserving all untouched graph state.
The code-graph-rag project maintains a structural and semantic representation of your codebase in a graph database. When source files change, you need an efficient way to synchronize the graph without rebuilding from scratch. This guide explains how to update a knowledge graph for a repository using the incremental re-ingestion pipeline implemented in vitali87/code-graph-rag.
Understanding the GraphUpdater.reingest Architecture
The reingest method in codebase_rag/graph_updater.py orchestrates a 12-step pipeline that guarantees atomic, consistent updates. If any step fails before graph mutation, a ReingestAborted exception rolls back all changes—no partial writes ever reach the database.
Step-by-Step Re-ingestion Flow
- Mark incremental build — Sets
self._is_full_build = FalseinGraphUpdater.__init__(line 4495) to distinguish from full rebuilds - Split paths into three groups — The
_reingest_splitmethod (line 4524) categorizes paths as present (on disk), gone (deleted), or skipped (matches ignore rules) - Hydrate from existing graph —
_hydrate_for_reingest(line 4538) loads current packages and folders for context - Load hash cache —
_load_hash_cache(line 777) readsHASH_CACHE_FILENAMEto detect actual file changes efficiently - Resolve stem-flux —
_reingest_flux_survivors(line 4690) handles cases like header files being added/removed alongside implementations - Seed module Q-N map —
_seed_module_qns_from_graph(line 4770) preserves qualified names for unchanged files to prevent name theft from new siblings - Compute dependents —
_reingest_dependents(line 4820) finds one-hop inbound edges to determine minimal re-parse set - Capture inbound edges —
_capture_inbound_edges(line 4850) snapshots the subgraph that will be rewritten - Delete and re-parse —
_reingest_delete(line 4960) removes stale nodes;_reingest_reparse(line 4975) generates fresh definitions, imports, and calls - Resolve calls —
_reingest_resolve(line 5000) links calls only for re-parsed files, preserving untouched inbound edges - Update hash cache —
_update_hashes(line 5030) atomically writes new file hashes - Emit ReingestReport — Returns elapsed time plus counts of reparsed, affected, removed, and skipped files
How to Update a Knowledge Graph Programmatically
Use GraphUpdater.reingest directly when integrating into build scripts, CI pipelines, or custom tooling.
from pathlib import Path
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.services import IngestorProtocol
# 1. Create your ingestor (e.g., Neo4jIngestor)
ingestor: IngestorProtocol = ...
# 2. Instantiate the updater with parsers and queries for your languages
updater = GraphUpdater(
ingestor=ingestor,
repo_path=Path("/path/to/your/repo"),
parsers={"python": python_parser, "cpp": cpp_parser},
queries={"python": python_queries, "cpp": cpp_queries},
)
# 3. Re-ingest specific changed files
changed = [Path("src/module/foo.py"), Path("src/module/bar.py")]
report = updater.reingest(changed)
# 4. Inspect the results
print("Re-parsed:", report.reparsed)
print("Affected (dependents):", report.affected)
print("Removed:", report.removed)
print("Skipped (ignored):", report.skipped)
print("Elapsed ms:", round(report.elapsed_ms, 1))
The reingest method accepts Path objects that can be absolute or relative to repo_path. It automatically expands the input set to include dependents—files that import from or call into your changed files—ensuring the graph remains internally consistent.
How to Update a Knowledge Graph via CLI
For manual updates or shell integrations, use the cgr CLI wrapper defined in codebase_rag/graph_cli.py:
cgr reingest --repo /path/to/repo \
src/module/foo.py src/module/bar.py
The CLI constructs the same GraphUpdater instance internally, making it equivalent to the programmatic approach. Pass multiple paths as positional arguments; the tool handles path normalization and ignore-rule filtering automatically.
Key Source Files for Knowledge Graph Updates
Understanding these files helps when debugging or extending re-ingestion behavior:
codebase_rag/graph_updater.py— ContainsGraphUpdaterclass withreingestentry point and all 12 pipeline stepscodebase_rag/constants.py— DefinesHASH_CACHE_FILENAMEand log message keys used throughout re-ingestioncodebase_rag/services.py— AbstractIngestorProtocoland query interfaces thatGraphUpdaterreads and writescodebase_rag/graph_cli.py— Command-line frontend that parses arguments and invokesreingestcodebase_rag/utils/path_utils.py— Path normalization, ignore-rule evaluation, and stem-based qualified name utilities
Safety Guarantees in reingest
The reingest implementation provides two critical consistency guarantees:
- Atomic failure handling — If
_hydrate_for_reingest,_capture_inbound_edges, or any pre-mutation step fails,ReingestAbortedprevents any graph changes - Minimal re-parse scope — Only changed files, their dependents, and stem-flux survivors are processed; untouched files retain all inbound and outbound edges exactly as before
This design makes reingest safe to run frequently—even on every commit—without degrading graph performance or integrity.
Summary
- Call
GraphUpdater.reingestwith changed file paths to perform incremental knowledge graph updates - The method automatically computes dependents and handles stem-flux without manual intervention
- Use the
cgr reingestCLI for command-line workflows - All changes are atomic: failures abort with
ReingestAbortedbefore any graph mutation occurs - Core implementation lives in
codebase_rag/graph_updater.pywith supporting utilities inconstants.py,services.py, andpath_utils.py
Frequently Asked Questions
What happens if a file is deleted from the repository?
The _reingest_split method detects deletions and routes them to _reingest_delete, which removes stale nodes from the graph. The hash cache is updated to remove the deleted file's entry, ensuring subsequent runs don't attempt to process it.
How does reingest avoid re-parsing unchanged files?
The system maintains a file hash cache in HASH_CACHE_FILENAME (defined in constants.py). During _load_hash_cache, current file hashes are compared against cached values. Only files with mismatched hashes—or those with changed dependents—enter the re-parse set.
Can reingest handle multiple programming languages in one call?
Yes. The GraphUpdater constructor accepts parsers and queries dictionaries mapping language identifiers to tree-sitter parsers and compiled query objects. The reingest method dispatches each file to its matching parser based on extension or path rules.
What is "stem-flux" and why does it matter?
Stem-flux occurs when related files (like foo.h and foo.cpp) change together. The _reingest_flux_survivors step identifies these relationships to prevent qualified name collisions and ensure that adding or removing companion files doesn't corrupt the module namespace of existing code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →