# Understanding the Knowledge Graph Schema in Code-Graph-RAG: A Deep Dive into the Protobuf Structure

> Explore the Code-Graph-RAG knowledge graph schema. Understand the Protobuf structure, GraphCodeIndex, Node messages, and Relationship edges for efficient code indexing.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: deep-dive
- Published: 2026-09-04

---

**Code-Graph-RAG stores codebases as property graphs using a protobuf schema defined in `codec/schema.proto`, with `GraphCodeIndex` as the root container, polymorphic `Node` messages for code entities, and typed `Relationship` edges connecting them by primary keys.**

The knowledge graph schema Code-Graph-RAG employs transforms raw source code into a queryable property graph structure. Defined in the `vitali87/code-graph-rag` repository, this schema captures everything from high-level project structure down to individual function calls and security vulnerabilities using Protocol Buffers. The append-only design ensures backward compatibility while supporting rich semantic relationships between code elements.

## Core Architecture of the Knowledge Graph Schema

The schema follows a property graph model where nodes represent code entities and edges represent semantic relationships. Three primary message types define the structure.

### The Top-Level Container (GraphCodeIndex)

The `GraphCodeIndex` message serves as the root container for the entire codebase graph. It maintains two parallel collections: all nodes and all relationships.

```proto
message GraphCodeIndex {
    repeated Node nodes = 1;
    repeated Relationship relationships = 2;
}

```

This flat structure allows efficient serialization while preserving graph topology through foreign key references.

### The Generic Node Wrapper (Node oneof)

Individual nodes use a polymorphic design via the `oneof` keyword in `codec/schema.proto`. The `Node` message wraps specific entity types in a type-safe container, allowing the graph to store heterogeneous entities while maintaining strongly-typed access.

```proto
message Node {
    oneof payload {
        Project project = 1;
        Package package = 2;
        Folder folder = 3;
        Module module = 4;
        Class class_node = 5;
        Function function = 6;
        Method method = 7;
        File file = 8;
        ExternalPackage external_package = 9;
        // ... additional entity types
    }
}

```

Each payload variant corresponds to a specific code entity, from high-level `Project` containers down to individual `Function` and `Method` definitions.

### Typed Relationships (Relationship)

Relationships connect nodes using string-based primary keys and explicitly typed semantics via the `RelationshipType` enum.

```proto
message Relationship {
    enum RelationshipType {
        RELATIONSHIP_TYPE_UNSPECIFIED = 0;
        CONTAINS_PACKAGE = 1;
        CONTAINS_FOLDER = 2;
        CONTAINS_FILE = 3;
        CONTAINS_MODULE = 4;
        DEFINES = 5;
        DEFINES_METHOD = 6;
        IMPORTS = 7;
        INHERITS = 8;
        OVERRIDES = 9;
        CALLS = 10;
        DEPENDS_ON_EXTERNAL = 11;
        IMPLEMENTS_MODULE = 12;
        IMPLEMENTS = 13;
        EXPORTS = 14;
        EXPORTS_MODULE = 15;
        READS_FROM = 16;
        WRITES_TO = 17;
        CONTAINS_SECTION = 18;
        EXPOSES = 19;
        FLOWS_TO = 20;
        HAS_SMELL = 21;
        HAS_VULNERABILITY = 22;
        IMPLEMENTS_PATTERN = 23;
        INSTANTIATES = 24;
        LINKS_TO = 25;
        REFERENCES = 26;
        RESOLVES_TO = 27;
        RETURNS = 28;
        ACCEPTS = 29;
    }
    RelationshipType type = 1;
    string source_id = 2;
    string target_id = 3;
    google.protobuf.Struct properties = 4;
    string source_label = 5;
    string target_label = 6;
}

```

The `source_id` and `target_id` fields reference the primary keys of the connected nodes, while `properties` stores arbitrary metadata as a `google.protobuf.Struct`.

## Node Types and Primary Keys

Each concrete node type defines a primary key field for unique identification and domain-specific metadata fields. The schema distinguishes between structural containers and behavioral entities.

- **Project**: Primary key is `name`. Represents the root of the graph and the entry point for codebase analysis.
- **Package**: Primary key is `qualified_name`. Tracks importable packages with `name` and `path` attributes.
- **Folder**: Primary key is `path`. Represents directory structure within the filesystem.
- **File**: Primary key is `path`. Captures individual source files including `name` and `extension`.
- **Module**: Primary key is `qualified_name`. Represents Python modules with decorators, flow coverage flags, and path information.
- **Function / Method**: Primary key is `qualified_name`. Stores callable entities with `docstring`, line numbers (`start_line`, `end_line`), decorators, return types, parameter types, and fingerprint data for deduplication.
- **Class**: Primary key is `qualified_name`. Captures class definitions with `docstring`, decorators, and export status.
- **ExternalPackage**: Primary key is `name`. References third-party dependencies.
- **Pattern / CodeSmell / SecurityIssue**: Primary key is `qualified_name`. Represents analysis findings with location, message, and code snippet data.
- **Resource**: Primary key is `qualified_name`. Tracks external resources like configuration files or assets.
- **Section**: Primary key is `qualified_name`. Represents documentation sections with heading levels and source locations.

Additional node types including `Interface`, `Enum`, `Union`, and `ExternalModule` follow the same primary key pattern using `qualified_name`.

## Relationship Types and Semantic Meaning

The `RelationshipType` enum categorizes edges into structural containment, behavioral interaction, and analytical metadata.

**Structural containment** establishes hierarchy:
- `CONTAINS_PACKAGE`, `CONTAINS_FOLDER`, `CONTAINS_FILE`, `CONTAINS_MODULE` link containers to their contents.
- `DEFINES` and `DEFINES_METHOD` connect modules to the classes and functions they declare.

**Behavioral interaction** captures runtime semantics:
- `CALLS` represents invocations between functions or methods.
- `INHERITS` and `IMPLEMENTS` capture class hierarchy and interface implementation.
- `OVERRIDES` links overriding methods to their parent implementations.
- `IMPORTS` tracks module dependencies.
- `INSTANTIATES` records object creation.

**Analytical metadata** attaches quality and security data:
- `HAS_SMELL` links code to detected code smells.
- `HAS_VULNERABILITY` connects entities to security issues.
- `IMPLEMENTS_PATTERN` associates code with detected design patterns.
- `READS_FROM` and `WRITES_TO` track data flow between entities.

## Building a Graph Programmatically

The generated Python bindings in [`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py) enable direct graph construction. The following example creates a minimal graph with a project, module, functions, and a call relationship:

```python
from codec import schema_pb2 as schema
from google.protobuf.struct_pb2 import Struct

# 1️⃣ Create nodes

proj = schema.Node(project=schema.Project(name="my_project"))
mod = schema.Node(
    module=schema.Module(
        qualified_name="my_project.utils",
        name="utils",
        path="my_project/utils.py",
    )
)
func_a = schema.Node(
    function=schema.Function(
        qualified_name="my_project.utils.func_a",
        name="func_a",
        start_line=10,
        end_line=20,
    )
)
func_b = schema.Node(
    function=schema.Function(
        qualified_name="my_project.utils.func_b",
        name="func_b",
        start_line=30,
        end_line=40,
    )
)

# 2️⃣ Create relationships

contains_mod = schema.Relationship(
    type=schema.Relationship.CONTAINS_MODULE,
    source_id=proj.project.name,
    target_id=mod.module.qualified_name,
    source_label="Project",
    target_label="Module",
    properties=Struct(),
)

defines_func_a = schema.Relationship(
    type=schema.Relationship.DEFINES,
    source_id=mod.module.qualified_name,
    target_id=func_a.function.qualified_name,
    source_label="Module",
    target_label="Function",
    properties=Struct(),
)

call_a_to_b = schema.Relationship(
    type=schema.Relationship.CALLS,
    source_id=func_a.function.qualified_name,
    target_id=func_b.function.qualified_name,
    source_label="Function",
    target_label="Function",
    properties=Struct(),
)

# 3️⃣ Assemble the index

graph = schema.GraphCodeIndex(
    nodes=[proj, mod, func_a, func_b],
    relationships=[contains_mod, defines_func_a, call_a_to_b],
)

# 4️⃣ Serialize to bytes (e.g., write to a protobuf file)

with open("example.graphpb", "wb") as f:
    f.write(graph.SerializeToString())

```

This demonstrates creating concrete payloads, using the `RelationshipType` enum for typed edges, and serializing the complete `GraphCodeIndex` for storage.

## Key Implementation Files

The knowledge graph schema Code-Graph-RAG relies on specific files for definition, generation, and manipulation:

- **`codec/schema.proto`**: The canonical protobuf definition containing `GraphCodeIndex`, `Node`, and `Relationship` message declarations.
- **[`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py)**: Auto-generated Python bindings used throughout the codebase to construct and parse graph data.
- **[`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py)**: Loads serialized `GraphCodeIndex` protobuf files into graph databases like Memgraph.
- **[`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py)**: Walks abstract syntax trees (AST), builds the protobuf graph representation, and persists updates to the storage layer.
- **[`codebase_rag/graph_query.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_query.py)**: Provides helper functions to execute queries against the graph using relationship types and node labels.

## Summary

- **Property graph model**: Code-Graph-RAG uses a protobuf-based property graph with separate node and relationship collections.
- **Polymorphic nodes**: The `Node` message uses `oneof payload` to type-safely wrap diverse entities from Projects to individual Methods.
- **Primary key addressing**: All entities reference each other via string primary keys (`name` or `qualified_name`) rather than internal indices.
- **Rich relationship semantics**: Twenty-nine distinct relationship types capture structural containment, behavioral calls, inheritance, and quality analysis.
- **Backward compatibility**: The schema maintains append-only field numbering to ensure persisted graphs remain readable across versions.

## Frequently Asked Questions

### What format does Code-Graph-RAG use to store its knowledge graph?

Code-Graph-RAG uses Protocol Buffers (protobuf) as the serialization format. The schema defined in `codec/schema.proto` compiles to Python bindings in [`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py), allowing efficient binary storage of the `GraphCodeIndex` structure while maintaining type safety and cross-language compatibility.

### How are different code entities represented in the schema?

The schema employs a polymorphic design where a generic `Node` message contains a `oneof payload` field. This field can hold specific typed messages like `Project`, `Module`, `Class`, `Function`, or `Method`, each with domain-specific attributes such as `qualified_name`, `start_line`, and `docstring`.

### What relationship types are available for connecting code nodes?

The schema defines twenty-nine relationship types in the `RelationshipType` enum, including structural edges (`CONTAINS_FILE`, `CONTAINS_MODULE`), behavioral edges (`CALLS`, `INHERITS`, `IMPORTS`), and analytical edges (`HAS_SMELL`, `HAS_VULNERABILITY`, `IMPLEMENTS_PATTERN`). Each relationship specifies `source_id` and `target_id` referencing the primary keys of connected nodes.

### How does the schema ensure backward compatibility?

The protobuf schema follows append-only conventions where existing field numbers are never reused or removed. New fields are added with incremental field numbers, ensuring that graphs serialized with older schema versions can be deserialized by newer implementations without data loss.