How to Customize Graph Generation in code-graph-rag: 4 Methods Explained

You can customize graph generation in code-graph-rag through capture selection (filtering node labels and relationships), parser configuration, CLI flags, or by subclassing GraphUpdater to inject post-processing hooks.

The code-graph-rag repository transforms source code into a queryable knowledge graph stored in Memgraph. While the default pipeline extracts standard relationships like CALLS and READS_FROM, the architecture exposes multiple extension points in codebase_rag/graph_updater.py, codebase_rag/capture.py, and codebase_rag/cli.py that allow precise control over which nodes are created, which relationships are captured, and how the graph is post-processed.

The Graph Generation Pipeline

Understanding the internal flow helps identify where to inject customizations. The GraphUpdater.run() method in codebase_rag/graph_updater.py orchestrates five distinct stages:

  1. File discovery – _collect_eligible_files walks the repository, applying should_skip_path and user-provided exclusion patterns.
  2. Parsing – Language-specific frontends (Tree-sitter-based or hybrid parsers) process files based on the configuration in load_parsers().
  3. Capture – The CaptureSelection dataclass determines which relationships are emitted (e.g., CALLS, READS_FROM, WRITES_TO) based on soft-dependency rules defined in _SOFT_DEPENDENCIES.
  4. Ingestion – Nodes and relationships buffer in _CapturingIngestor and flush to Memgraph via FilteringIngestor, which respects the active capture selection.
  5. Optional embeddings – If enabled, _generate_semantic_embeddings pushes source snippets to Qdrant for vector search.

Each stage offers distinct customization mechanisms that do not require modifying core library code.

Method 1: Controlling Capture Selection

The most common customization involves filtering which relationship types and node labels enter the graph. The CaptureSelection dataclass in codebase_rag/capture.py manages this through three interfaces:

Environment Variable: Set CGR_CAPTURE before running the CLI.

import os
os.environ["CGR_CAPTURE"] = "calls,io"

CLI Flag: Use --capture with specification tokens (+ to add, - to remove).

python -m codebase_rag.cli sync ./my_repo my_project --capture +calls,-reads_from

Programmatic: Instantiate CaptureSelection directly and pass it to GraphUpdater.

from pathlib import Path
from codebase_rag.capture import CaptureSelection, resolve_capture, split_spec
from codebase_rag.graph_updater import GraphUpdater

capture = resolve_capture(split_spec("calls,io"))
updater = GraphUpdater(
    ingestor=memgraph_ingestor,
    repo_path=Path("./my_repo"),
    capture=capture,
    # ... other args

)

The resolve_capture function processes tokens like io, none, or +CALLS to build the final enabled set, which FilteringIngestor uses to drop unwanted relationships before database insertion.

Method 2: Configuring Parsers and Queries

To modify how source files are parsed or to support additional languages, customize the parser pipeline in codebase_rag/graph_updater.py:

  • Factory modification: Edit codebase_rag/parsers/factory.py to add or remove ProcessorFactory pipelines that map file extensions to Tree-sitter parsers.
  • Custom queries: Provide custom .scm (Scheme) query files and pass them via the queries argument when instantiating GraphUpdater. These queries define how Tree-sitter extracts AST nodes for relationship mapping.

The load_parsers() method in GraphUpdater initializes these pipelines, allowing you to inject custom language frontends without altering the core ingestion logic.

Method 3: CLI Pipeline Options

The command-line interface in codebase_rag/cli.py exposes high-level controls through _run_graph_sync and _capture_selection. These flags modify pipeline behavior without code changes:

  • --exclude <patterns>: Skip specific files or folders during the discovery phase.
  • --clean: Wipe the existing Memgraph database before ingestion for a fresh load.
  • --skip-embeddings: Disable the optional Qdrant embedding step for faster syncs when vector search is not required.
  • --batch-size: Tune the Memgraph transaction size to optimize ingestion performance on large codebases.

Example usage combining multiple options:

python -m codebase_rag.cli sync ./large_repo project_name \
  --exclude "tests/,vendor/" \
  --clean \
  --capture calls,definitions \
  --batch-size 1000

Method 4: Extending GraphUpdater with Subclasses

For advanced customizations, subclass GraphUpdater to override private helper methods called during the run() lifecycle, such as _emit_pending_endpoints or _prune_orphan_nodes. Evaluation scripts in evals/cgr_graph.py and usage examples in examples/graph_export_example.py demonstrate patterns for extracting and exporting custom graph subsets.

The following example injects synthetic DEPENDS_ON relationships after standard ingestion:

from pathlib import Path
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.capture import resolve_capture, split_spec
from codebase_rag.services import FilteringIngestor

class CustomGraphUpdater(GraphUpdater):
    def _post_process(self):
        # Link every Function node to a global Config node

        self._sink.ensure_node_batch("Config", {"name": "global", "id": 1})
        for fn in self._sink.nodes:
            self._sink.ensure_relationship_batch(
                from_spec=("Function", "id", fn[1]["id"]),
                rel_type="DEPENDS_ON",
                to_spec=("Config", "id", 1),
                properties=None,
            )
        self._sink.flush_all()

# Instantiate with custom configuration

capture = resolve_capture(split_spec("calls,io"))
updater = CustomGraphUpdater(
    ingestor=memgraph_ingestor,
    repo_path=Path("./my_repo"),
    parsers=parsers,
    queries=queries,
    capture=capture,
)
updater.run()

This pattern allows arbitrary graph mutations—adding synthetic nodes, computing derived relationships, or cleaning orphans—while reusing the standard parsing and ingestion infrastructure.

Summary

  • Capture selection controls which relationships and nodes enter the graph via CGR_CAPTURE, --capture flags, or the CaptureSelection dataclass in codebase_rag/capture.py.
  • Parser configuration customizes language support through ProcessorFactory in codebase_rag/parsers/factory.py and custom .scm query files passed to GraphUpdater.
  • CLI options in codebase_rag/cli.py provide toggles for exclusions, clean runs, embedding generation, and batch sizing without code modification.
  • Subclassing GraphUpdater enables post-processing hooks and custom relationship injection by overriding methods like _post_process or _prune_orphan_nodes.

Frequently Asked Questions

How do I exclude test files from graph generation?

Use the --exclude flag in the CLI with glob patterns, or set exclusion rules programmatically by overriding _collect_eligible_files in a GraphUpdater subclass. The CLI approach requires no code changes: python -m codebase_rag.cli sync ./repo proj --exclude "tests/,*_test.py".

Can I add custom relationship types beyond CALLS and READS_FROM?

Yes. While the capture selection filters existing types, you can emit custom relationships by subclassing GraphUpdater and calling self._sink.ensure_relationship_batch() with your custom rel_type string in a post-processing hook. Ensure your Memgraph schema supports the new relationship type.

What is the difference between --capture none and --skip-embeddings?

--capture none (or CGR_CAPTURE=none) disables all relationship and node capture in the graph generation phase, resulting in an empty graph. --skip-embeddings only disables the optional vectorization step that sends code snippets to Qdrant; the structural graph in Memgraph is still built normally.

Where is the graph schema defined?

The schema is implicit in the CaptureSelection logic within codebase_rag/capture.py and the ingestion methods in codebase_rag/services.py. Node labels and relationship types are determined by the enabled capture tokens and the Tree-sitter query results, not by a static DDL file.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →