# How to Generate a Graph for a Specific Part of a Codebase with Code-Graph-RAG

> Generate a code graph for a specific part of your codebase using Code-Graph-RAG. Load, query, filter, and export sub-graphs with simple Python collections.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-18

---

**TLDR:** Load a full code-graph JSON with `GraphLoader`, query nodes by labels or file paths, collect connected relationships, and export a filtered sub-graph using standard Python collections.

The vitali87/code-graph-rag repository provides a Python toolkit for extracting focused subgraphs from large codebases. When you need to generate a graph for a specific part of the codebase—such as a single package, module, or class hierarchy—the `GraphLoader` class offers indexed lookup capabilities that make sub-graph extraction both efficient and memory-safe.

## Understanding the GraphLoader Architecture

The `GraphLoader` class in [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) serves as the primary interface for graph manipulation. When you invoke `load_graph(path_to_json)`, the loader constructs several in-memory indexes that enable O(1) lookups without mutating the original graph data.

### In-Memory Indexing Strategy

The loader builds four distinct lookup structures during initialization:

- **`_nodes_by_id`**: Maps `node_id` strings to `GraphNode` objects for direct retrieval
- **`_nodes_by_label`**: Indexes nodes by their `NodeLabel` classification (e.g., FUNCTION, CLASS, MODULE)
- **`_outgoing_rels` / `_incoming_rels`**: Bidirectional relationship indexes mapping `node_id` to lists of `GraphRelationship` objects
- **`_property_indexes`**: Lazy-built indexes that map property names to value-to-node mappings, constructed on first access via `_build_property_index`

These indexes allow you to efficiently query massive code graphs without loading the entire structure into a graph database.

## Querying the Graph for Specific Code Regions

To generate a graph for a specific part of the codebase, you filter nodes using the public API methods exposed by `GraphLoader`.

### By Node Labels

Use `loader.find_nodes_by_label(NodeLabel.FUNCTION)` to retrieve all function nodes, or substitute `NodeLabel.CLASS` and `NodeLabel.MODULE` for other scopes. This leverages the `_nodes_by_label` index for immediate results.

### By File Path Properties

Query nodes by their source file location using `loader.find_node_by_property("path", "src/utils.py")`. The `"path"` property is defined in [`codebase_rag/constants/graph.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/graph.py) as the default key storing relative file paths. This method triggers `_build_property_index` on first call, then caches subsequent lookups.

### By Relationships

Once you identify seed nodes, traverse connections using `loader.get_relationships_for_node(node_id)`. For directional analysis, use `loader.get_outgoing_relationships(node_id)` or `loader.get_incoming_relationships(node_id)`, which reference the `_outgoing_rels` and `_incoming_rels` indexes respectively.

## Step-by-Step: Extracting a Sub-Graph

Follow this five-step workflow to isolate a specific portion of your codebase:

1. **Load the full graph** using `load_graph(path_to_exported_json)`
2. **Select seed nodes** by filtering on properties like file path or labels
3. **Collect transitive relationships** by checking edges that touch your seed node IDs
4. **Build a new `GraphData` structure** containing only the filtered nodes and relationships
5. **Serialize to JSON** for consumption by downstream tools or Neo4j imports

### Complete Implementation Example

The following script extracts all nodes under the `my_pkg/` directory and writes a self-contained sub-graph to [`my_pkg_graph.json`](https://github.com/vitali87/code-graph-rag/blob/main/my_pkg_graph.json):

```python
#!/usr/bin/env python3
import json
from pathlib import Path
from codebase_rag.graph_loader import load_graph, GraphLoader
from codebase_rag.constants import NodeLabel, KEY_PROPERTIES, KEY_LABELS, KEY_RELATIONSHIPS, KEY_NODES

# ----------------------------------------------------------------------

# 1️⃣ Load the full exported graph

# ----------------------------------------------------------------------

GRAPH_FILE = Path("exported_codegraph.json")          # <-- replace with your file

graph: GraphLoader = load_graph(str(GRAPH_FILE))

# ----------------------------------------------------------------------

# 2️⃣ Choose seed nodes (all nodes whose file path starts with "my_pkg/")

# ----------------------------------------------------------------------

seed_nodes = [
    node for node in graph.nodes
    if str(node.properties.get("path", "")).startswith("my_pkg/")
]

# ----------------------------------------------------------------------

# 3️⃣ Collect relationships that touch any of the seed nodes

# ----------------------------------------------------------------------

seed_ids = {node.node_id for node in seed_nodes}
relevant_rels = [
    rel for rel in graph.relationships
    if rel.from_id in seed_ids or rel.to_id in seed_ids
]

# ----------------------------------------------------------------------

# 4️⃣ Build the sub‑graph payload (same shape as the original JSON)

# ----------------------------------------------------------------------

subgraph_payload = {
    "metadata": graph.metadata,
    "nodes": [
        {
            "id": n.node_id,
            "labels": n.labels,
            "properties": n.properties,
        }
        for n in seed_nodes
    ],
    "relationships": [
        {
            "from": r.from_id,
            "to": r.to_id,
            "type": r.type,
            "properties": r.properties,
        }
        for r in relevant_rels
    ],
}

# ----------------------------------------------------------------------

# 5️⃣ Write the sub‑graph to disk

# ----------------------------------------------------------------------

OUT_FILE = Path("my_pkg_graph.json")
OUT_FILE.write_text(json.dumps(subgraph_payload, indent=2))
print(f"✅ Sub‑graph written to {OUT_FILE}")

```

**Key implementation details:**

- Access the complete node and relationship collections via `graph.nodes` and `graph.relationships`
- The `"path"` property comes from the constants defined in [`codebase_rag/constants/graph.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/graph.py)
- The output JSON schema matches the original export format, ensuring compatibility with the CLI and Neo4j import tools

## Core API and File Reference

These source files implement the sub-graph generation capabilities:

- **[`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py)**: Contains `GraphLoader` class with `load_graph()` factory function and indexing logic
- **[`codebase_rag/types_defs.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/types_defs.py)**: Defines `GraphData`, `GraphNode`, and `GraphRelationship` data classes used throughout the pipeline
- **[`codebase_rag/constants/graph.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/graph.py)**: Stores JSON key constants including `KEY_NODES`, `KEY_RELATIONSHIPS`, and property names
- **[`examples/graph_export_example.py`](https://github.com/vitali87/code-graph-rag/blob/main/examples/graph_export_example.py)**: Demonstrates end-to-end graph loading, summary generation, and label-based querying
- **[`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py)**: Implements the `graph-loader` command for interactive graph inspection

## Summary

- The `GraphLoader` class in [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) builds fast in-memory indexes for O(1) node and relationship lookups
- Filter nodes using `find_nodes_by_label()` for type-based extraction or `find_node_by_property()` for path-based filtering
- Collect relationships by checking intersection with seed node IDs against `graph.relationships`
- Export sub-graphs by constructing a new dictionary with `metadata`, `nodes`, and `relationships` keys
- The resulting JSON maintains schema compatibility with the broader code-graph-rag toolchain

## Frequently Asked Questions

### How does GraphLoader handle large codebases without running out of memory?

`GraphLoader` keeps the full graph in memory but uses dictionary-based indexes (`_nodes_by_id`, `_nodes_by_label`) that store references rather than copies. Property indexes build lazily via `_build_property_index`, so you only pay memory costs for indexes you actively query. For extremely large graphs, filter early using file path properties rather than loading all nodes into Python lists.

### Can I export a sub-graph that includes relationships two or three hops away from my seed nodes?

Yes. After collecting your initial `seed_ids`, iterate through `graph.relationships` to find nodes connected to your seeds, add those node IDs to your collection, and repeat for the desired hop depth. The relationship indexes in `_outgoing_rels` and `_incoming_rels` make these traversal steps efficient even for multi-hop collections.

### What is the difference between `load_graph()` and the CLI `graph-loader` command?

`load_graph()` in [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) is the Python API that returns a `GraphLoader` instance for programmatic manipulation. The `graph-loader` command in [`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py) wraps this function for interactive terminal use, printing summaries and allowing quick inspection without writing Python scripts.

### Does the sub-graph JSON maintain compatibility with Neo4j import tools?

Yes. The output schema follows the same structure as the original export: top-level keys for `metadata`, `nodes`, and `relationships`, with each node containing `id`, `labels`, and `properties`. This format aligns with the Neo4j import specifications used by the broader code-graph-rag pipeline.