# Code-Graph-RAG Graph Schema: Complete Protocol Buffers Reference for Property Graph Modeling

> Explore the Code-Graph-RAG graph schema defined in Protocol Buffers. Understand nodes like Project, Class, Function, and relationships like CONTAINS_FILE, CALLS for codebase modeling.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: api-reference
- Published: 2026-09-06

---

**The graph schema for Code-Graph-RAG is defined in a Protocol Buffers file at `codec/schema.proto`, which specifies nodes (Project, Package, Class, Function, etc.) and directed relationships (CONTAINS_FILE, DEFINES, CALLS, INHERITS, etc.) that together model an entire codebase as a serializable property graph.**

Code-Graph-RAG is an open-source tool by vitali87 that transforms source code repositories into queryable knowledge graphs for retrieval-augmented generation (RAG) pipelines. Understanding its graph schema is essential for anyone building custom queries, extending the indexer, or integrating the output with graph databases like Memgraph.

## Schema Location and Structure

The canonical schema definition lives at [`codec/schema.proto`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema.proto). The repository includes auto-generated Python bindings in [[`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py)](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py) that provide programmatic access to all message types.

At the root of every indexed codebase sits a **GraphCodeIndex** message containing two repeated fields:

- `nodes` — a collection of strongly-typed entities wrapped in generic `Node` containers
- `relationships` — directed edges connecting nodes by their primary keys

This flat structure enables efficient serialization while preserving rich semantic relationships.

## Node Types and Their Primary Keys

Each node type uses a **qualified name** or **path** as its stable identifier, ensuring deterministic graph construction across indexing runs.

| Node Type | Primary Key Field | Key Characteristics |
|-----------|-------------------|---------------------|
| **Project** | `name` | Single root node per codebase |
| **Package** | `qualified_name` | Logical grouping (Maven/Gradle/npm equivalent) |
| **Folder** | `path` | Physical directory on disk |
| **File** | `path` | Individual source file with `extension` metadata |
| **Module** | `qualified_name` | Language-specific module with Rust/Python-specific fields like `rust_cfg_test_mods` and `flow_covered` |
| **Class** | `qualified_name` | OO definition with `docstring`, line numbers, `decorators`, `is_exported` |
| **Function** | `qualified_name` | Stand-alone function with `ast_fingerprint`, `return_type`, `param_types` |
| **Method** | `qualified_name` | Same fields as Function with method semantics |
| **ExternalPackage** | `name` | Third-party dependency boundary |
| **ExternalModule** | `qualified_name` | Module from external dependency |
| **Pattern / CodeSmell / SecurityIssue** | `qualified_name` | Quality signals from AST-grep with `snippet` evidence |
| **Resource** | `qualified_name` | I/O resources: files, network, environment variables |
| **Section** | `qualified_name` | Markdown documentation headings |
| **Interface / Enum / Type / Union** | `qualified_name` | Language-specific type constructs |

The `Node` message itself uses a `oneof payload` to enforce type safety while allowing polymorphic storage:

```protobuf
// From codec/schema.proto
message Node {
  oneof payload {
    Project project = 1;
    Package package = 2;
    // ... all other node types
    Function function = 10;
    Method method = 11;
  }
}

```

## Relationship Types and Semantics

Relationships are defined in the `Relationship.RelationshipType` enum. The design follows **append-only versioning**: each enum value is frozen to guarantee wire-format compatibility when new edge types are added.

### Hierarchical Containment

- `CONTAINS_PACKAGE` — Project → Package
- `CONTAINS_FOLDER` — Package/Folder → Folder
- `CONTAINS_FILE` — Folder → File
- `CONTAINS_MODULE` — File → Module

### Definition Edges

- `DEFINES` — Module → Class/Function/Interface/Enum/Type
- `DEFINES_METHOD` — Class → Method

### Code Dependencies

- `IMPORTS` — Module/File → external symbol
- `CALLS` — Function/Method → callable target
- `INHERITS` — Class → parent Class
- `OVERRIDES` — Method → overridden Method

### Cross-Package Relationships

- `DEPENDS_ON_EXTERNAL` — Package → ExternalPackage
- `IMPLEMENTS_MODULE` — Module → external interface/trait
- `EXPORTS` — Module → public API surface

### Quality and Security Signals

- `HAS_SMELL` — any node → CodeSmell
- `HAS_VULNERABILITY` — any node → SecurityIssue
- `HAS_PATTERN` — any node → Pattern (AST-grep match)

### Reference Resolution

- `REFERENCES` — generic usage edge
- `RESOLVES_TO` — symbol → its definition (for import resolution)

Every `Relationship` stores `source_id`, `target_id`, `source_label`, `target_label`, and an extensible `properties` Struct for edge attributes.

## Building and Serializing Graphs

The following example demonstrates constructing a minimal graph programmatically using the Python bindings from [`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py):

```python
from codec import schema_pb2 as schema
from google.protobuf.struct_pb2 import Struct

# Create nodes

project = schema.Node(project=schema.Project(name="my_service"))
source_file = schema.Node(
    file=schema.File(
        path="src/app.py",
        name="app.py",
        extension=".py"
    )
)

# Establish containment relationship

contains_rel = schema.Relationship(
    type=schema.Relationship.CONTAINS_FILE,
    source_id="my_service",      # Project.name

    target_id="src/app.py",       # File.path

    source_label="Project",
    target_label="File",
    properties=Struct()
)

# Assemble and serialize

index = schema.GraphCodeIndex(
    nodes=[project, source_file],
    relationships=[contains_rel]
)

# Write to disk for later ingestion

with open("codebase_index.pb", "wb") as f:
    f.write(index.SerializeToString())

```

## Loading and Querying Existing Indexes

For analysis or database ingestion, load a serialized index as follows:

```python
from codec import schema_pb2 as schema

# Parse from disk

with open("codebase_index.pb", "rb") as f:
    index = schema.GraphCodeIndex()
    index.ParseFromString(f.read())

# Extract all exported functions with their line ranges

exported_functions = [
    (node.function.qualified_name,
     node.function.start_line,
     node.function.end_line)
    for node in index.nodes
    if node.HasField("function") and node.function.is_exported
]

print(f"Found {len(exported_functions)} exported functions")

```

## Key Implementation Files

| File | Role in Schema Pipeline |
|------|------------------------|
| [`codec/schema.proto`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema.proto) | Ground-truth protobuf definition |
| [[`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py)](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py) | Generated Python API |
| [[`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) | Deserializes `GraphCodeIndex` to in-memory objects |
| [[`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) | Incremental updates using schema semantics |
| [[`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py) | Orchestrates indexing and persistence |
| [[`codebase_rag/cypher_queries.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cypher_queries.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cypher_queries.py) | Cypher templates matching schema labels |

## Summary

- The **Code-Graph-RAG graph schema** is a Protocol Buffers property graph defined in `codec/schema.proto`.
- **Nodes** represent 15+ code entities from Project to SecurityIssue, each with a stable primary key.
- **Relationships** are typed directed edges from an append-only enum, covering containment, definition, dependency, and quality signals.
- The **GraphCodeIndex** message provides a flat, serializable container for entire codebases.
- Generated Python bindings enable programmatic graph construction, mutation, and analysis.

## Frequently Asked Questions

### What file format does Code-Graph-RAG use for graph storage?

Code-Graph-RAG uses **Protocol Buffers** as its native serialization format. The `GraphCodeIndex` message can be written to `.pb` files, sent over gRPC, or loaded directly into Python for processing. This binary format is compact and schema-evolution safe.

### How are node types distinguished in the schema?

Node types use a **discriminated union pattern**: the `Node` message contains a `oneof payload` field that holds exactly one strongly-typed message (e.g., `Project`, `Function`, `Class`). The Python API provides `HasField()` to test which type is populated without manual type inspection.

### Can I add custom properties to relationships?

Yes. Every `Relationship` includes a `properties` field of type `google.protobuf.Struct`, which accepts arbitrary JSON-like key-value pairs. This allows the schema to remain stable while supporting extensible edge attributes for specific use cases.

### What guarantees does the relationship enum provide?

The `RelationshipType` enum is **append-only**: existing numeric values never change, ensuring that indexes serialized with older schema versions remain readable. New edge types receive the next available integer, maintaining backward and forward compatibility.