How to Generate a Graph for a Specific Part of a Codebase with Code-Graph-RAG

TLDR: Load a full code-graph JSON with GraphLoader, query nodes by labels or file paths, collect connected relationships, and export a filtered sub-graph using standard Python collections.

The vitali87/code-graph-rag repository provides a Python toolkit for extracting focused subgraphs from large codebases. When you need to generate a graph for a specific part of the codebase—such as a single package, module, or class hierarchy—the GraphLoader class offers indexed lookup capabilities that make sub-graph extraction both efficient and memory-safe.

Understanding the GraphLoader Architecture

The GraphLoader class in codebase_rag/graph_loader.py serves as the primary interface for graph manipulation. When you invoke load_graph(path_to_json), the loader constructs several in-memory indexes that enable O(1) lookups without mutating the original graph data.

In-Memory Indexing Strategy

The loader builds four distinct lookup structures during initialization:

  • _nodes_by_id: Maps node_id strings to GraphNode objects for direct retrieval
  • _nodes_by_label: Indexes nodes by their NodeLabel classification (e.g., FUNCTION, CLASS, MODULE)
  • _outgoing_rels / _incoming_rels: Bidirectional relationship indexes mapping node_id to lists of GraphRelationship objects
  • _property_indexes: Lazy-built indexes that map property names to value-to-node mappings, constructed on first access via _build_property_index

These indexes allow you to efficiently query massive code graphs without loading the entire structure into a graph database.

Querying the Graph for Specific Code Regions

To generate a graph for a specific part of the codebase, you filter nodes using the public API methods exposed by GraphLoader.

By Node Labels

Use loader.find_nodes_by_label(NodeLabel.FUNCTION) to retrieve all function nodes, or substitute NodeLabel.CLASS and NodeLabel.MODULE for other scopes. This leverages the _nodes_by_label index for immediate results.

By File Path Properties

Query nodes by their source file location using loader.find_node_by_property("path", "src/utils.py"). The "path" property is defined in codebase_rag/constants/graph.py as the default key storing relative file paths. This method triggers _build_property_index on first call, then caches subsequent lookups.

By Relationships

Once you identify seed nodes, traverse connections using loader.get_relationships_for_node(node_id). For directional analysis, use loader.get_outgoing_relationships(node_id) or loader.get_incoming_relationships(node_id), which reference the _outgoing_rels and _incoming_rels indexes respectively.

Step-by-Step: Extracting a Sub-Graph

Follow this five-step workflow to isolate a specific portion of your codebase:

  1. Load the full graph using load_graph(path_to_exported_json)
  2. Select seed nodes by filtering on properties like file path or labels
  3. Collect transitive relationships by checking edges that touch your seed node IDs
  4. Build a new GraphData structure containing only the filtered nodes and relationships
  5. Serialize to JSON for consumption by downstream tools or Neo4j imports

Complete Implementation Example

The following script extracts all nodes under the my_pkg/ directory and writes a self-contained sub-graph to my_pkg_graph.json:

#!/usr/bin/env python3
import json
from pathlib import Path
from codebase_rag.graph_loader import load_graph, GraphLoader
from codebase_rag.constants import NodeLabel, KEY_PROPERTIES, KEY_LABELS, KEY_RELATIONSHIPS, KEY_NODES

# ----------------------------------------------------------------------

# 1️⃣ Load the full exported graph

# ----------------------------------------------------------------------

GRAPH_FILE = Path("exported_codegraph.json")          # <-- replace with your file

graph: GraphLoader = load_graph(str(GRAPH_FILE))

# ----------------------------------------------------------------------

# 2️⃣ Choose seed nodes (all nodes whose file path starts with "my_pkg/")

# ----------------------------------------------------------------------

seed_nodes = [
    node for node in graph.nodes
    if str(node.properties.get("path", "")).startswith("my_pkg/")
]

# ----------------------------------------------------------------------

# 3️⃣ Collect relationships that touch any of the seed nodes

# ----------------------------------------------------------------------

seed_ids = {node.node_id for node in seed_nodes}
relevant_rels = [
    rel for rel in graph.relationships
    if rel.from_id in seed_ids or rel.to_id in seed_ids
]

# ----------------------------------------------------------------------

# 4️⃣ Build the sub‑graph payload (same shape as the original JSON)

# ----------------------------------------------------------------------

subgraph_payload = {
    "metadata": graph.metadata,
    "nodes": [
        {
            "id": n.node_id,
            "labels": n.labels,
            "properties": n.properties,
        }
        for n in seed_nodes
    ],
    "relationships": [
        {
            "from": r.from_id,
            "to": r.to_id,
            "type": r.type,
            "properties": r.properties,
        }
        for r in relevant_rels
    ],
}

# ----------------------------------------------------------------------

# 5️⃣ Write the sub‑graph to disk

# ----------------------------------------------------------------------

OUT_FILE = Path("my_pkg_graph.json")
OUT_FILE.write_text(json.dumps(subgraph_payload, indent=2))
print(f"✅ Sub‑graph written to {OUT_FILE}")

Key implementation details:

  • Access the complete node and relationship collections via graph.nodes and graph.relationships
  • The "path" property comes from the constants defined in codebase_rag/constants/graph.py
  • The output JSON schema matches the original export format, ensuring compatibility with the CLI and Neo4j import tools

Core API and File Reference

These source files implement the sub-graph generation capabilities:

Summary

  • The GraphLoader class in codebase_rag/graph_loader.py builds fast in-memory indexes for O(1) node and relationship lookups
  • Filter nodes using find_nodes_by_label() for type-based extraction or find_node_by_property() for path-based filtering
  • Collect relationships by checking intersection with seed node IDs against graph.relationships
  • Export sub-graphs by constructing a new dictionary with metadata, nodes, and relationships keys
  • The resulting JSON maintains schema compatibility with the broader code-graph-rag toolchain

Frequently Asked Questions

How does GraphLoader handle large codebases without running out of memory?

GraphLoader keeps the full graph in memory but uses dictionary-based indexes (_nodes_by_id, _nodes_by_label) that store references rather than copies. Property indexes build lazily via _build_property_index, so you only pay memory costs for indexes you actively query. For extremely large graphs, filter early using file path properties rather than loading all nodes into Python lists.

Can I export a sub-graph that includes relationships two or three hops away from my seed nodes?

Yes. After collecting your initial seed_ids, iterate through graph.relationships to find nodes connected to your seeds, add those node IDs to your collection, and repeat for the desired hop depth. The relationship indexes in _outgoing_rels and _incoming_rels make these traversal steps efficient even for multi-hop collections.

What is the difference between load_graph() and the CLI graph-loader command?

load_graph() in codebase_rag/graph_loader.py is the Python API that returns a GraphLoader instance for programmatic manipulation. The graph-loader command in codebase_rag/cli.py wraps this function for interactive terminal use, printing summaries and allowing quick inspection without writing Python scripts.

Does the sub-graph JSON maintain compatibility with Neo4j import tools?

Yes. The output schema follows the same structure as the original export: top-level keys for metadata, nodes, and relationships, with each node containing id, labels, and properties. This format aligns with the Neo4j import specifications used by the broader code-graph-rag pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →