Memgraph Graph Schema for Code-Graph-RAG: Protobuf Definition and Implementation
The Memgraph graph schema for Code-Graph-RAG is defined in codec/schema.proto as a flattened property graph consisting of strongly-typed Node payloads and directed Relationship edges with explicit type enums.
Code-Graph-RAG stores entire parsed codebases in Memgraph as a navigable property graph. The schema architecture, implemented in Protocol Buffers, defines how source code entities map to vertices and edges for graph-based retrieval augmented generation.
Core Schema Architecture
The schema centers on three foundational concepts defined in codec/schema.proto: the GraphCodeIndex container, generic Node wrappers, and typed Relationship edges.
GraphCodeIndex Container
At line 81, the GraphCodeIndex message serves as the top-level collection for all graph data:
message GraphCodeIndex {
repeated Node nodes = 1;
repeated Relationship relationships = 2;
}
When ingested into Memgraph, each entry in the nodes repeated field becomes a vertex, while each Relationship becomes a directed edge connecting source and target vertices by their primary keys.
Node Structure and Payloads
Defined at line 90, the Node message uses a oneof field to enforce type safety across different code entities. This generic wrapper can hold exactly one payload type, determining the vertex label in Memgraph.
Relationship Type System
The Relationship message (lines 114-164) defines directed edges with the following properties:
source_idandtarget_id: Primary keys referencing the connected nodessource_labelandtarget_label: Concrete node types (e.g.,Function,Class)type: ARelationshipTypeenum value determining the edge semanticsproperties: A free-formStructfor additional metadata such as call-site line numbers
Node Payload Types and Primary Keys
Each concrete payload within a Node carries distinct fields and primary key identifiers. The following table maps payload types to their keys and typical attributes:
| Payload | Primary Key | Key Fields |
|---|---|---|
| Project | name |
Container for the entire codebase |
| Package | qualified_name |
name, path |
| Folder | path |
Directory structure representation |
| File | path |
name, extension |
| Module | qualified_name |
name, path, decorators, flow_covered flag |
| Class | qualified_name |
name, docstring, line range, is_exported |
| Method | qualified_name |
name, docstring, line range, decorators |
| Function | qualified_name |
name, docstring, line range, is_exported |
| ExternalPackage | name |
Third-party dependency reference |
| ExternalModule | qualified_name |
External library components |
| Resource | qualified_name |
name, kind |
| Interface / Enum / Type / Union | qualified_name |
Type system constructs |
According to the source code, specific definitions include the Project message starting at line 151, the Function message at line 242, and the Class message at line 264 of codec/schema.proto.
Relationship Type Enumeration
The RelationshipType enum (lines 166-199) captures every logical connection between code entities:
enum RelationshipType {
RELATIONSHIP_TYPE_UNSPECIFIED = 0;
CONTAINS_PACKAGE = 1;
CONTAINS_FOLDER = 2;
CONTAINS_FILE = 3;
CONTAINS_MODULE = 4;
DEFINES = 5;
DEFINES_METHOD = 6;
IMPORTS = 7;
INHERITS = 8;
OVERRIDES = 9;
CALLS = 10;
DEPENDS_ON_EXTERNAL = 11;
IMPLEMENTS_MODULE = 12;
IMPLEMENTS = 13;
EXPORTS = 14;
EXPORTS_MODULE = 15;
READS_FROM = 16;
WRITES_TO = 17;
}
These types enable precise graph traversal patterns, from hierarchical containment (CONTAINS_FILE) to semantic dependencies (CALLS, INHERITS, IMPORTS).
Memgraph Mapping and Ingestion
The Code-Graph-RAG runtime, implemented in [codebase_rag/services/protobuf_service.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/protobuf_service.py), serializes the GraphCodeIndex and streams it to Memgraph via the ingest service. During this process:
- Each Node payload becomes a vertex with a label matching the payload type (e.g.,
Project,Function) - Each Relationship becomes a directed edge labeled with the enum value (e.g.,
CALLS,CONTAINS_MODULE) - Primary keys from the protobuf definitions become vertex identifiers for edge construction
This mapping enables complex Cypher queries across the codebase:
MATCH (p:Project)-[:CONTAINS_PACKAGE]->(pkg:Package)
WHERE pkg.name = "my_lib"
RETURN pkg
Building a Code Graph in Python
The Python-generated classes in [codec/schema_pb2.py](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py) materialize the protobuf definitions for runtime graph construction:
import codec.schema_pb2 as pb
# Create a project vertex
proj = pb.Node()
proj.project.name = "my_project"
# Create a file vertex
file_node = pb.Node()
file_node.file.path = "src/main.py"
file_node.file.name = "main.py"
file_node.file.extension = ".py"
# Establish containment relationship
rel = pb.Relationship(
type=pb.Relationship.CONTAINS_FILE,
source_id=proj.project.name,
target_id=file_node.file.path,
source_label="Project",
target_label="File"
)
# Assemble the complete graph index
index = pb.GraphCodeIndex()
index.nodes.extend([proj, file_node])
index.relationships.append(rel)
# Serialize for Memgraph ingestion
payload = index.SerializeToString()
# Send payload to the Memgraph ingest endpoint
The integration tests in [codebase_rag/tests/test_protobuf_end_to_end.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tests/test_protobuf_end_to_end.py) verify this serialization pipeline end-to-end.
Summary
- The Memgraph graph schema for Code-Graph-RAG is authoritatively defined in
codec/schema.protousing Protocol Buffers. - Nodes are generic wrappers containing strongly-typed payloads (Project, Function, Class, etc.), each identified by primary keys such as
qualified_nameorpath. - Relationships are directed edges with explicit
RelationshipTypeenums includingCONTAINS_FILE,CALLS,INHERITS, andIMPORTS. - The schema maps directly to Memgraph vertices and edges, enabling complex Cypher queries over codebases via the ingest service implemented in
protobuf_service.py.
Frequently Asked Questions
How does Code-Graph-RAG handle different programming languages in the same schema?
The schema uses generic payload types like Function, Class, and Module that are language-agnostic, with language-specific distinctions handled through optional fields (such as C++ specific ModuleImplementation and ModuleInterface messages) and decorator metadata stored in the payload properties.
What primary key should I use to query a specific function in the graph?
Use the qualified_name field, which serves as the primary key for Function, Method, Class, and Module payloads. For Project and ExternalPackage nodes, use the name field, while File and Folder nodes use path as their identifier.
Where does the actual ingestion into Memgraph happen?
The ingestion logic resides in [codebase_rag/services/protobuf_service.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/protobuf_service.py), which builds the GraphCodeIndex from extracted AST data and streams the serialized protobuf to the Memgraph ingest service, creating vertices and edges according to the schema definitions.
Can I add custom properties to relationships?
Yes, the Relationship message includes a properties field of type Struct (defined in the protobuf), allowing arbitrary key-value metadata such as call-site line numbers or dependency versions to be attached to edges without modifying the core schema.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →