How to Generate a Graph for a Specific Part of a Codebase with Code-Graph-RAG
TLDR: Load a full code-graph JSON with GraphLoader, query nodes by labels or file paths, collect connected relationships, and export a filtered sub-graph using standard Python collections.
The vitali87/code-graph-rag repository provides a Python toolkit for extracting focused subgraphs from large codebases. When you need to generate a graph for a specific part of the codebase—such as a single package, module, or class hierarchy—the GraphLoader class offers indexed lookup capabilities that make sub-graph extraction both efficient and memory-safe.
Understanding the GraphLoader Architecture
The GraphLoader class in codebase_rag/graph_loader.py serves as the primary interface for graph manipulation. When you invoke load_graph(path_to_json), the loader constructs several in-memory indexes that enable O(1) lookups without mutating the original graph data.
In-Memory Indexing Strategy
The loader builds four distinct lookup structures during initialization:
_nodes_by_id: Mapsnode_idstrings toGraphNodeobjects for direct retrieval_nodes_by_label: Indexes nodes by theirNodeLabelclassification (e.g., FUNCTION, CLASS, MODULE)_outgoing_rels/_incoming_rels: Bidirectional relationship indexes mappingnode_idto lists ofGraphRelationshipobjects_property_indexes: Lazy-built indexes that map property names to value-to-node mappings, constructed on first access via_build_property_index
These indexes allow you to efficiently query massive code graphs without loading the entire structure into a graph database.
Querying the Graph for Specific Code Regions
To generate a graph for a specific part of the codebase, you filter nodes using the public API methods exposed by GraphLoader.
By Node Labels
Use loader.find_nodes_by_label(NodeLabel.FUNCTION) to retrieve all function nodes, or substitute NodeLabel.CLASS and NodeLabel.MODULE for other scopes. This leverages the _nodes_by_label index for immediate results.
By File Path Properties
Query nodes by their source file location using loader.find_node_by_property("path", "src/utils.py"). The "path" property is defined in codebase_rag/constants/graph.py as the default key storing relative file paths. This method triggers _build_property_index on first call, then caches subsequent lookups.
By Relationships
Once you identify seed nodes, traverse connections using loader.get_relationships_for_node(node_id). For directional analysis, use loader.get_outgoing_relationships(node_id) or loader.get_incoming_relationships(node_id), which reference the _outgoing_rels and _incoming_rels indexes respectively.
Step-by-Step: Extracting a Sub-Graph
Follow this five-step workflow to isolate a specific portion of your codebase:
- Load the full graph using
load_graph(path_to_exported_json) - Select seed nodes by filtering on properties like file path or labels
- Collect transitive relationships by checking edges that touch your seed node IDs
- Build a new
GraphDatastructure containing only the filtered nodes and relationships - Serialize to JSON for consumption by downstream tools or Neo4j imports
Complete Implementation Example
The following script extracts all nodes under the my_pkg/ directory and writes a self-contained sub-graph to my_pkg_graph.json:
#!/usr/bin/env python3
import json
from pathlib import Path
from codebase_rag.graph_loader import load_graph, GraphLoader
from codebase_rag.constants import NodeLabel, KEY_PROPERTIES, KEY_LABELS, KEY_RELATIONSHIPS, KEY_NODES
# ----------------------------------------------------------------------
# 1️⃣ Load the full exported graph
# ----------------------------------------------------------------------
GRAPH_FILE = Path("exported_codegraph.json") # <-- replace with your file
graph: GraphLoader = load_graph(str(GRAPH_FILE))
# ----------------------------------------------------------------------
# 2️⃣ Choose seed nodes (all nodes whose file path starts with "my_pkg/")
# ----------------------------------------------------------------------
seed_nodes = [
node for node in graph.nodes
if str(node.properties.get("path", "")).startswith("my_pkg/")
]
# ----------------------------------------------------------------------
# 3️⃣ Collect relationships that touch any of the seed nodes
# ----------------------------------------------------------------------
seed_ids = {node.node_id for node in seed_nodes}
relevant_rels = [
rel for rel in graph.relationships
if rel.from_id in seed_ids or rel.to_id in seed_ids
]
# ----------------------------------------------------------------------
# 4️⃣ Build the sub‑graph payload (same shape as the original JSON)
# ----------------------------------------------------------------------
subgraph_payload = {
"metadata": graph.metadata,
"nodes": [
{
"id": n.node_id,
"labels": n.labels,
"properties": n.properties,
}
for n in seed_nodes
],
"relationships": [
{
"from": r.from_id,
"to": r.to_id,
"type": r.type,
"properties": r.properties,
}
for r in relevant_rels
],
}
# ----------------------------------------------------------------------
# 5️⃣ Write the sub‑graph to disk
# ----------------------------------------------------------------------
OUT_FILE = Path("my_pkg_graph.json")
OUT_FILE.write_text(json.dumps(subgraph_payload, indent=2))
print(f"✅ Sub‑graph written to {OUT_FILE}")
Key implementation details:
- Access the complete node and relationship collections via
graph.nodesandgraph.relationships - The
"path"property comes from the constants defined incodebase_rag/constants/graph.py - The output JSON schema matches the original export format, ensuring compatibility with the CLI and Neo4j import tools
Core API and File Reference
These source files implement the sub-graph generation capabilities:
codebase_rag/graph_loader.py: ContainsGraphLoaderclass withload_graph()factory function and indexing logiccodebase_rag/types_defs.py: DefinesGraphData,GraphNode, andGraphRelationshipdata classes used throughout the pipelinecodebase_rag/constants/graph.py: Stores JSON key constants includingKEY_NODES,KEY_RELATIONSHIPS, and property namesexamples/graph_export_example.py: Demonstrates end-to-end graph loading, summary generation, and label-based queryingcodebase_rag/cli.py: Implements thegraph-loadercommand for interactive graph inspection
Summary
- The
GraphLoaderclass incodebase_rag/graph_loader.pybuilds fast in-memory indexes for O(1) node and relationship lookups - Filter nodes using
find_nodes_by_label()for type-based extraction orfind_node_by_property()for path-based filtering - Collect relationships by checking intersection with seed node IDs against
graph.relationships - Export sub-graphs by constructing a new dictionary with
metadata,nodes, andrelationshipskeys - The resulting JSON maintains schema compatibility with the broader code-graph-rag toolchain
Frequently Asked Questions
How does GraphLoader handle large codebases without running out of memory?
GraphLoader keeps the full graph in memory but uses dictionary-based indexes (_nodes_by_id, _nodes_by_label) that store references rather than copies. Property indexes build lazily via _build_property_index, so you only pay memory costs for indexes you actively query. For extremely large graphs, filter early using file path properties rather than loading all nodes into Python lists.
Can I export a sub-graph that includes relationships two or three hops away from my seed nodes?
Yes. After collecting your initial seed_ids, iterate through graph.relationships to find nodes connected to your seeds, add those node IDs to your collection, and repeat for the desired hop depth. The relationship indexes in _outgoing_rels and _incoming_rels make these traversal steps efficient even for multi-hop collections.
What is the difference between load_graph() and the CLI graph-loader command?
load_graph() in codebase_rag/graph_loader.py is the Python API that returns a GraphLoader instance for programmatic manipulation. The graph-loader command in codebase_rag/cli.py wraps this function for interactive terminal use, printing summaries and allowing quick inspection without writing Python scripts.
Does the sub-graph JSON maintain compatibility with Neo4j import tools?
Yes. The output schema follows the same structure as the original export: top-level keys for metadata, nodes, and relationships, with each node containing id, labels, and properties. This format aligns with the Neo4j import specifications used by the broader code-graph-rag pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →