How to Customize Graph Generation in code-graph-rag: 4 Methods Explained
You can customize graph generation in code-graph-rag through capture selection (filtering node labels and relationships), parser configuration, CLI flags, or by subclassing GraphUpdater to inject post-processing hooks.
The code-graph-rag repository transforms source code into a queryable knowledge graph stored in Memgraph. While the default pipeline extracts standard relationships like CALLS and READS_FROM, the architecture exposes multiple extension points in codebase_rag/graph_updater.py, codebase_rag/capture.py, and codebase_rag/cli.py that allow precise control over which nodes are created, which relationships are captured, and how the graph is post-processed.
The Graph Generation Pipeline
Understanding the internal flow helps identify where to inject customizations. The GraphUpdater.run() method in codebase_rag/graph_updater.py orchestrates five distinct stages:
- File discovery –
_collect_eligible_fileswalks the repository, applyingshould_skip_pathand user-provided exclusion patterns. - Parsing – Language-specific frontends (Tree-sitter-based or hybrid parsers) process files based on the configuration in
load_parsers(). - Capture – The
CaptureSelectiondataclass determines which relationships are emitted (e.g.,CALLS,READS_FROM,WRITES_TO) based on soft-dependency rules defined in_SOFT_DEPENDENCIES. - Ingestion – Nodes and relationships buffer in
_CapturingIngestorand flush to Memgraph viaFilteringIngestor, which respects the active capture selection. - Optional embeddings – If enabled,
_generate_semantic_embeddingspushes source snippets to Qdrant for vector search.
Each stage offers distinct customization mechanisms that do not require modifying core library code.
Method 1: Controlling Capture Selection
The most common customization involves filtering which relationship types and node labels enter the graph. The CaptureSelection dataclass in codebase_rag/capture.py manages this through three interfaces:
Environment Variable: Set CGR_CAPTURE before running the CLI.
import os
os.environ["CGR_CAPTURE"] = "calls,io"
CLI Flag: Use --capture with specification tokens (+ to add, - to remove).
python -m codebase_rag.cli sync ./my_repo my_project --capture +calls,-reads_from
Programmatic: Instantiate CaptureSelection directly and pass it to GraphUpdater.
from pathlib import Path
from codebase_rag.capture import CaptureSelection, resolve_capture, split_spec
from codebase_rag.graph_updater import GraphUpdater
capture = resolve_capture(split_spec("calls,io"))
updater = GraphUpdater(
ingestor=memgraph_ingestor,
repo_path=Path("./my_repo"),
capture=capture,
# ... other args
)
The resolve_capture function processes tokens like io, none, or +CALLS to build the final enabled set, which FilteringIngestor uses to drop unwanted relationships before database insertion.
Method 2: Configuring Parsers and Queries
To modify how source files are parsed or to support additional languages, customize the parser pipeline in codebase_rag/graph_updater.py:
- Factory modification: Edit
codebase_rag/parsers/factory.pyto add or removeProcessorFactorypipelines that map file extensions to Tree-sitter parsers. - Custom queries: Provide custom
.scm(Scheme) query files and pass them via thequeriesargument when instantiatingGraphUpdater. These queries define how Tree-sitter extracts AST nodes for relationship mapping.
The load_parsers() method in GraphUpdater initializes these pipelines, allowing you to inject custom language frontends without altering the core ingestion logic.
Method 3: CLI Pipeline Options
The command-line interface in codebase_rag/cli.py exposes high-level controls through _run_graph_sync and _capture_selection. These flags modify pipeline behavior without code changes:
--exclude <patterns>: Skip specific files or folders during the discovery phase.--clean: Wipe the existing Memgraph database before ingestion for a fresh load.--skip-embeddings: Disable the optional Qdrant embedding step for faster syncs when vector search is not required.--batch-size: Tune the Memgraph transaction size to optimize ingestion performance on large codebases.
Example usage combining multiple options:
python -m codebase_rag.cli sync ./large_repo project_name \
--exclude "tests/,vendor/" \
--clean \
--capture calls,definitions \
--batch-size 1000
Method 4: Extending GraphUpdater with Subclasses
For advanced customizations, subclass GraphUpdater to override private helper methods called during the run() lifecycle, such as _emit_pending_endpoints or _prune_orphan_nodes. Evaluation scripts in evals/cgr_graph.py and usage examples in examples/graph_export_example.py demonstrate patterns for extracting and exporting custom graph subsets.
The following example injects synthetic DEPENDS_ON relationships after standard ingestion:
from pathlib import Path
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.capture import resolve_capture, split_spec
from codebase_rag.services import FilteringIngestor
class CustomGraphUpdater(GraphUpdater):
def _post_process(self):
# Link every Function node to a global Config node
self._sink.ensure_node_batch("Config", {"name": "global", "id": 1})
for fn in self._sink.nodes:
self._sink.ensure_relationship_batch(
from_spec=("Function", "id", fn[1]["id"]),
rel_type="DEPENDS_ON",
to_spec=("Config", "id", 1),
properties=None,
)
self._sink.flush_all()
# Instantiate with custom configuration
capture = resolve_capture(split_spec("calls,io"))
updater = CustomGraphUpdater(
ingestor=memgraph_ingestor,
repo_path=Path("./my_repo"),
parsers=parsers,
queries=queries,
capture=capture,
)
updater.run()
This pattern allows arbitrary graph mutations—adding synthetic nodes, computing derived relationships, or cleaning orphans—while reusing the standard parsing and ingestion infrastructure.
Summary
- Capture selection controls which relationships and nodes enter the graph via
CGR_CAPTURE,--captureflags, or theCaptureSelectiondataclass incodebase_rag/capture.py. - Parser configuration customizes language support through
ProcessorFactoryincodebase_rag/parsers/factory.pyand custom.scmquery files passed toGraphUpdater. - CLI options in
codebase_rag/cli.pyprovide toggles for exclusions, clean runs, embedding generation, and batch sizing without code modification. - Subclassing
GraphUpdaterenables post-processing hooks and custom relationship injection by overriding methods like_post_processor_prune_orphan_nodes.
Frequently Asked Questions
How do I exclude test files from graph generation?
Use the --exclude flag in the CLI with glob patterns, or set exclusion rules programmatically by overriding _collect_eligible_files in a GraphUpdater subclass. The CLI approach requires no code changes: python -m codebase_rag.cli sync ./repo proj --exclude "tests/,*_test.py".
Can I add custom relationship types beyond CALLS and READS_FROM?
Yes. While the capture selection filters existing types, you can emit custom relationships by subclassing GraphUpdater and calling self._sink.ensure_relationship_batch() with your custom rel_type string in a post-processing hook. Ensure your Memgraph schema supports the new relationship type.
What is the difference between --capture none and --skip-embeddings?
--capture none (or CGR_CAPTURE=none) disables all relationship and node capture in the graph generation phase, resulting in an empty graph. --skip-embeddings only disables the optional vectorization step that sends code snippets to Qdrant; the structural graph in Memgraph is still built normally.
Where is the graph schema defined?
The schema is implicit in the CaptureSelection logic within codebase_rag/capture.py and the ingestion methods in codebase_rag/services.py. Node labels and relationship types are determined by the enabled capture tokens and the Tree-sitter query results, not by a static DDL file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →