Understanding the Knowledge Graph Schema in Code-Graph-RAG: A Deep Dive into the Protobuf Structure
Code-Graph-RAG stores codebases as property graphs using a protobuf schema defined in codec/schema.proto, with GraphCodeIndex as the root container, polymorphic Node messages for code entities, and typed Relationship edges connecting them by primary keys.
The knowledge graph schema Code-Graph-RAG employs transforms raw source code into a queryable property graph structure. Defined in the vitali87/code-graph-rag repository, this schema captures everything from high-level project structure down to individual function calls and security vulnerabilities using Protocol Buffers. The append-only design ensures backward compatibility while supporting rich semantic relationships between code elements.
Core Architecture of the Knowledge Graph Schema
The schema follows a property graph model where nodes represent code entities and edges represent semantic relationships. Three primary message types define the structure.
The Top-Level Container (GraphCodeIndex)
The GraphCodeIndex message serves as the root container for the entire codebase graph. It maintains two parallel collections: all nodes and all relationships.
message GraphCodeIndex {
repeated Node nodes = 1;
repeated Relationship relationships = 2;
}
This flat structure allows efficient serialization while preserving graph topology through foreign key references.
The Generic Node Wrapper (Node oneof)
Individual nodes use a polymorphic design via the oneof keyword in codec/schema.proto. The Node message wraps specific entity types in a type-safe container, allowing the graph to store heterogeneous entities while maintaining strongly-typed access.
message Node {
oneof payload {
Project project = 1;
Package package = 2;
Folder folder = 3;
Module module = 4;
Class class_node = 5;
Function function = 6;
Method method = 7;
File file = 8;
ExternalPackage external_package = 9;
// ... additional entity types
}
}
Each payload variant corresponds to a specific code entity, from high-level Project containers down to individual Function and Method definitions.
Typed Relationships (Relationship)
Relationships connect nodes using string-based primary keys and explicitly typed semantics via the RelationshipType enum.
message Relationship {
enum RelationshipType {
RELATIONSHIP_TYPE_UNSPECIFIED = 0;
CONTAINS_PACKAGE = 1;
CONTAINS_FOLDER = 2;
CONTAINS_FILE = 3;
CONTAINS_MODULE = 4;
DEFINES = 5;
DEFINES_METHOD = 6;
IMPORTS = 7;
INHERITS = 8;
OVERRIDES = 9;
CALLS = 10;
DEPENDS_ON_EXTERNAL = 11;
IMPLEMENTS_MODULE = 12;
IMPLEMENTS = 13;
EXPORTS = 14;
EXPORTS_MODULE = 15;
READS_FROM = 16;
WRITES_TO = 17;
CONTAINS_SECTION = 18;
EXPOSES = 19;
FLOWS_TO = 20;
HAS_SMELL = 21;
HAS_VULNERABILITY = 22;
IMPLEMENTS_PATTERN = 23;
INSTANTIATES = 24;
LINKS_TO = 25;
REFERENCES = 26;
RESOLVES_TO = 27;
RETURNS = 28;
ACCEPTS = 29;
}
RelationshipType type = 1;
string source_id = 2;
string target_id = 3;
google.protobuf.Struct properties = 4;
string source_label = 5;
string target_label = 6;
}
The source_id and target_id fields reference the primary keys of the connected nodes, while properties stores arbitrary metadata as a google.protobuf.Struct.
Node Types and Primary Keys
Each concrete node type defines a primary key field for unique identification and domain-specific metadata fields. The schema distinguishes between structural containers and behavioral entities.
- Project: Primary key is
name. Represents the root of the graph and the entry point for codebase analysis. - Package: Primary key is
qualified_name. Tracks importable packages withnameandpathattributes. - Folder: Primary key is
path. Represents directory structure within the filesystem. - File: Primary key is
path. Captures individual source files includingnameandextension. - Module: Primary key is
qualified_name. Represents Python modules with decorators, flow coverage flags, and path information. - Function / Method: Primary key is
qualified_name. Stores callable entities withdocstring, line numbers (start_line,end_line), decorators, return types, parameter types, and fingerprint data for deduplication. - Class: Primary key is
qualified_name. Captures class definitions withdocstring, decorators, and export status. - ExternalPackage: Primary key is
name. References third-party dependencies. - Pattern / CodeSmell / SecurityIssue: Primary key is
qualified_name. Represents analysis findings with location, message, and code snippet data. - Resource: Primary key is
qualified_name. Tracks external resources like configuration files or assets. - Section: Primary key is
qualified_name. Represents documentation sections with heading levels and source locations.
Additional node types including Interface, Enum, Union, and ExternalModule follow the same primary key pattern using qualified_name.
Relationship Types and Semantic Meaning
The RelationshipType enum categorizes edges into structural containment, behavioral interaction, and analytical metadata.
Structural containment establishes hierarchy:
CONTAINS_PACKAGE,CONTAINS_FOLDER,CONTAINS_FILE,CONTAINS_MODULElink containers to their contents.DEFINESandDEFINES_METHODconnect modules to the classes and functions they declare.
Behavioral interaction captures runtime semantics:
CALLSrepresents invocations between functions or methods.INHERITSandIMPLEMENTScapture class hierarchy and interface implementation.OVERRIDESlinks overriding methods to their parent implementations.IMPORTStracks module dependencies.INSTANTIATESrecords object creation.
Analytical metadata attaches quality and security data:
HAS_SMELLlinks code to detected code smells.HAS_VULNERABILITYconnects entities to security issues.IMPLEMENTS_PATTERNassociates code with detected design patterns.READS_FROMandWRITES_TOtrack data flow between entities.
Building a Graph Programmatically
The generated Python bindings in codec/schema_pb2.py enable direct graph construction. The following example creates a minimal graph with a project, module, functions, and a call relationship:
from codec import schema_pb2 as schema
from google.protobuf.struct_pb2 import Struct
# 1️⃣ Create nodes
proj = schema.Node(project=schema.Project(name="my_project"))
mod = schema.Node(
module=schema.Module(
qualified_name="my_project.utils",
name="utils",
path="my_project/utils.py",
)
)
func_a = schema.Node(
function=schema.Function(
qualified_name="my_project.utils.func_a",
name="func_a",
start_line=10,
end_line=20,
)
)
func_b = schema.Node(
function=schema.Function(
qualified_name="my_project.utils.func_b",
name="func_b",
start_line=30,
end_line=40,
)
)
# 2️⃣ Create relationships
contains_mod = schema.Relationship(
type=schema.Relationship.CONTAINS_MODULE,
source_id=proj.project.name,
target_id=mod.module.qualified_name,
source_label="Project",
target_label="Module",
properties=Struct(),
)
defines_func_a = schema.Relationship(
type=schema.Relationship.DEFINES,
source_id=mod.module.qualified_name,
target_id=func_a.function.qualified_name,
source_label="Module",
target_label="Function",
properties=Struct(),
)
call_a_to_b = schema.Relationship(
type=schema.Relationship.CALLS,
source_id=func_a.function.qualified_name,
target_id=func_b.function.qualified_name,
source_label="Function",
target_label="Function",
properties=Struct(),
)
# 3️⃣ Assemble the index
graph = schema.GraphCodeIndex(
nodes=[proj, mod, func_a, func_b],
relationships=[contains_mod, defines_func_a, call_a_to_b],
)
# 4️⃣ Serialize to bytes (e.g., write to a protobuf file)
with open("example.graphpb", "wb") as f:
f.write(graph.SerializeToString())
This demonstrates creating concrete payloads, using the RelationshipType enum for typed edges, and serializing the complete GraphCodeIndex for storage.
Key Implementation Files
The knowledge graph schema Code-Graph-RAG relies on specific files for definition, generation, and manipulation:
codec/schema.proto: The canonical protobuf definition containingGraphCodeIndex,Node, andRelationshipmessage declarations.codec/schema_pb2.py: Auto-generated Python bindings used throughout the codebase to construct and parse graph data.codebase_rag/graph_loader.py: Loads serializedGraphCodeIndexprotobuf files into graph databases like Memgraph.codebase_rag/graph_updater.py: Walks abstract syntax trees (AST), builds the protobuf graph representation, and persists updates to the storage layer.codebase_rag/graph_query.py: Provides helper functions to execute queries against the graph using relationship types and node labels.
Summary
- Property graph model: Code-Graph-RAG uses a protobuf-based property graph with separate node and relationship collections.
- Polymorphic nodes: The
Nodemessage usesoneof payloadto type-safely wrap diverse entities from Projects to individual Methods. - Primary key addressing: All entities reference each other via string primary keys (
nameorqualified_name) rather than internal indices. - Rich relationship semantics: Twenty-nine distinct relationship types capture structural containment, behavioral calls, inheritance, and quality analysis.
- Backward compatibility: The schema maintains append-only field numbering to ensure persisted graphs remain readable across versions.
Frequently Asked Questions
What format does Code-Graph-RAG use to store its knowledge graph?
Code-Graph-RAG uses Protocol Buffers (protobuf) as the serialization format. The schema defined in codec/schema.proto compiles to Python bindings in codec/schema_pb2.py, allowing efficient binary storage of the GraphCodeIndex structure while maintaining type safety and cross-language compatibility.
How are different code entities represented in the schema?
The schema employs a polymorphic design where a generic Node message contains a oneof payload field. This field can hold specific typed messages like Project, Module, Class, Function, or Method, each with domain-specific attributes such as qualified_name, start_line, and docstring.
What relationship types are available for connecting code nodes?
The schema defines twenty-nine relationship types in the RelationshipType enum, including structural edges (CONTAINS_FILE, CONTAINS_MODULE), behavioral edges (CALLS, INHERITS, IMPORTS), and analytical edges (HAS_SMELL, HAS_VULNERABILITY, IMPLEMENTS_PATTERN). Each relationship specifies source_id and target_id referencing the primary keys of connected nodes.
How does the schema ensure backward compatibility?
The protobuf schema follows append-only conventions where existing field numbers are never reused or removed. New fields are added with incremental field numbers, ensuring that graphs serialized with older schema versions can be deserialized by newer implementations without data loss.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →