Code-Graph-RAG Graph Schema: A Complete Guide to Nodes and Edges
The Code-Graph-RAG graph schema defines 17 node types and 18 relationship types as a property graph using Protocol Buffers, with the root schema located in codec/schema.proto.
Code-Graph-RAG models any codebase as a structured property graph, enabling precise retrieval-augmented generation (RAG) over source code. The schema, defined in codec/schema.proto and compiled to codec/schema_pb2.py, provides a type-safe, serializable representation of code entities and their relationships. In this guide, you'll learn every node type, every edge type, and how they connect to form a queryable graph of your software.
Node Types: The 17 Entity Payloads
Nodes in Code-Graph-RAG use a oneof payload pattern—a single Node message holds exactly one strongly-typed entity. This design keeps the graph flat, normalized, and efficient to traverse.
Project and Structural Nodes
These nodes capture the physical and logical organization of code:
Project– The root of every graph, identified byname. Represents the entire codebase.Folder– Physical directory, keyed bypath.File– Individual source file withpath,name, andextensionfields.Package– Language-level package (e.g., Java package), keyed byqualified_name.Module– Language module such as a Python module or Rust crate, keyed byqualified_name.
Definition Nodes
These nodes represent declarations within the code:
Class– Class definitions across any object-oriented language.Function– Top-level functions, keyed byqualified_name.Method– Methods defined inside classes, distinct from standalone functions.Interface– Interface declarations (Java, TypeScript, etc.).Enum– Enumeration definitions.Type– Type aliases or custom type definitions.Union– Union type declarations.
External and System Nodes
These nodes model dependencies and non-code resources:
ExternalPackage– Third-party packages from npm, pip, Maven, etc.ExternalModule– External compiled libraries or modules.ModuleImplementation– Concrete implementation of a module interface.ModuleInterface– The public API surface of a module.Resource– Non-code files such as images, configuration files, or data assets.
Each payload includes auxiliary metadata fields—name, path, decorators, line numbers, and documentation—enabling rich contextual retrieval.
Relationship Types: The 18 Edge Kinds
Relationships are directed edges with a mandatory RelationshipType enum and optional property bag. The Relationship message stores source_id, target_id, source_label, target_label, and a google.protobuf.Struct for arbitrary properties.
Containment Hierarchy
| Relationship | Source | Target | Meaning |
|---|---|---|---|
CONTAINS_PACKAGE |
Project |
Package |
Project owns a package |
CONTAINS_FOLDER |
Project |
Folder |
Project contains a directory |
CONTAINS_FILE |
Folder |
File |
Directory holds a file |
CONTAINS_MODULE |
Folder |
Module |
Directory contains a module |
Definition and Inheritance
| Relationship | Source | Target | Meaning |
|---|---|---|---|
DEFINES |
Module |
Function / Class / Interface / Enum / Type / Union / Resource |
Module declares these entities |
DEFINES_METHOD |
Class |
Method |
Class declares a method |
INHERITS |
Class |
Class |
Class extends superclass |
OVERRIDES |
Method |
Method |
Method overrides parent implementation |
IMPLEMENTS |
Class |
Interface |
Class implements interface |
Dependencies and Calls
| Relationship | Source | Target | Meaning |
|---|---|---|---|
IMPORTS |
Module |
Module / Package |
Module imports another module or package |
CALLS |
Function / Method |
Function / Method |
Call graph edge |
DEPENDS_ON_EXTERNAL |
Project |
ExternalPackage |
Project depends on external library |
Module and Export Relationships
| Relationship | Source | Target | Meaning |
|---|---|---|---|
IMPLEMENTS_MODULE |
ExternalModule |
ModuleImplementation |
External module implements internal module |
EXPORTS |
Module |
(exported symbol) | Module exports a specific symbol |
EXPORTS_MODULE |
Module |
Module |
Module re-exports another module |
Resource Access
| Relationship | Source | Target | Meaning |
|---|---|---|---|
READS_FROM |
Function / Method |
Resource |
Code reads from resource |
WRITES_TO |
Function / Method |
Resource |
Code writes to resource |
Building and Serializing Graphs
The top-level container GraphCodeIndex aggregates all nodes and relationships. Here's how to construct a minimal graph programmatically using the generated schema_pb2 module:
from codec import schema_pb2 as schema
from google.protobuf.struct_pb2 import Struct
# Create a Project node
project = schema.Node(
project=schema.Project(name="code_graph_rag")
)
# Create a File node with metadata
main_file = schema.Node(
file=schema.File(
path="src/main.py",
name="main",
extension=".py"
)
)
# Create a Function node
entry_point = schema.Node(
function=schema.Function(
qualified_name="src.main.entry_point",
name="entry_point",
is_async=False,
decorators=["app.route"]
)
)
# Connect File to Function with DEFINES relationship
defines_rel = schema.Relationship(
type=schema.Relationship.DEFINES,
source_id=main_file.file.path,
target_id=entry_point.function.qualified_name,
source_label="File",
target_label="Function",
properties=Struct() # Add metadata here if needed
)
# Assemble into GraphCodeIndex
index = schema.GraphCodeIndex(
nodes=[project, main_file, entry_point],
relationships=[defines_rel]
)
# Serialize for storage or transmission
serialized = index.SerializeToString()
This pattern—used throughout codebase_rag/graph_updater.py—enables incremental graph construction during code analysis.
Loading and Querying Graphs
Deserialize and traverse graphs using standard protobuf methods:
from codec import schema_pb2 as schema
# Load from binary data
index = schema.GraphCodeIndex()
index.ParseFromString(serialized_bytes)
# Iterate nodes with type checking
for node in index.nodes:
if node.HasField("function"):
print(f"Function: {node.function.qualified_name}")
elif node.HasField("class_"):
print(f"Class: {node.class_.qualified_name}")
# Traverse relationships
for rel in index.relationships:
print(
f"{rel.source_label}({rel.source_id}) "
f"--[{schema.RelationshipType.Name(rel.type)}]--> "
f"{rel.target_label}({rel.target_id})"
)
The codebase_rag/graph_loader.py module implements higher-level loading patterns, including lazy materialization and property indexing.
Key Implementation Files
| File | Purpose | Direct Usage |
|---|---|---|
codec/schema.proto |
Canonical schema definition; single source of truth for all node and edge types | Code generation, documentation reference |
codec/schema_pb2.py |
Generated Python bindings—classes Node, Relationship, GraphCodeIndex, and all payload types |
All graph construction and serialization |
codebase_rag/graph_updater.py |
Emits Node and Relationship messages during AST analysis |
Called by language parsers to populate graphs |
codebase_rag/graph_loader.py |
Deserializes GraphCodeIndex and builds in-memory query structures |
Retrieval pipeline initialization |
codebase_rag/constants/graph.py |
String constants for labels and relationship types (e.g., "CONTAINS_FILE", "Class") |
Consistent labeling across the codebase |
Summary
- 17 node types cover projects, files, modules, classes, functions, methods, interfaces, enums, types, unions, external dependencies, and resources.
- 18 relationship types capture containment, definition, inheritance, calls, imports, exports, and resource access.
- Protocol Buffers schema in
codec/schema.protoensures type safety, efficient serialization, and cross-language compatibility. - Flat ID-based structure via
GraphCodeIndexeliminates duplication and supports incremental updates during code analysis.
Frequently Asked Questions
Where is the Code-Graph-RAG graph schema defined?
The schema is defined in codec/schema.proto at the repository root. This file contains all message definitions for nodes, relationships, and the top-level GraphCodeIndex container. The protoc compiler generates codec/schema_pb2.py for Python usage.
How does a node store its specific entity type?
Each Node message uses a oneof payload with 17 possible options. The HasField() method checks which payload is present, enabling type-safe access to node.function, node.class_, node.file, etc.
What fields does every relationship include?
Every Relationship stores source_id and target_id (primary keys), source_label and target_label (human-readable types), a RelationshipType enum, and an optional properties field of type google.protobuf.Struct for arbitrary key-value metadata.
Can I extend the schema with custom node or edge types?
Yes. Add new messages to the oneof payload in Node, new enum values to RelationshipType, and regenerate schema_pb2.py. The repository's loader and updater treat unknown types gracefully, though query logic in codebase_rag/ may need updates to handle new types.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →