# How to Index a Repository to Protobuf for Offline Use with Code-Graph-RAG

> Index your code repository to Protobuf for offline use with Code-Graph-RAG. Reload your code graphs instantly without re-parsing source files for faster development.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Code-Graph-RAG uses a deterministic Protobuf representation to store repository indexes so graphs can be reloaded without re-parsing source files.**

The `ProtobufFileIngestor` class in [`codebase_rag/services/protobuf_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/protobuf_service.py) handles the entire workflow: collecting nodes and relationships during repository scanning, normalizing paths for canonical IDs, and serializing the result into portable `.pb` files. This article covers the complete indexing process, from direct Python usage to CLI commands and offline loading.

## Architecture Overview

Code-Graph-RAG's Protobuf indexing system consists of four interconnected components:

| Component | Source File | Purpose |
|-----------|-------------|---------|
| **Protobuf schema** | [`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py) | Defines wire-format messages (`GraphCodeIndex`, `Node`, `Relationship`) |
| **Constants** | `codebase_rag/constants/*` | Enumerations (`NodeLabel`), one-of field mappings, output file names |
| **`ProtobufFileIngestor`** | [`codebase_rag/services/protobuf_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/protobuf_service.py) | Accumulates, sorts, and serializes graph data |
| **CLI interface** | [`codebase_rag/workspaces/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/workspaces/cli.py) | User-facing `codebase_rag index` command |

The ingestor guarantees **deterministic output** through sorted keys and `SerializeToString(deterministic=True)`, ensuring identical repositories produce byte-identical files regardless of checkout location.

## Indexing Methods

### Method 1: Direct Python API

For fine-grained control over the indexing process, instantiate `ProtobufFileIngestor` directly and feed it nodes and relationships as you discover them.

```python
from pathlib import Path
from codebase_rag.services.protobuf_service import ProtobufFileIngestor
import codebase_rag.constants as cs

# Initialize ingestor

ingestor = ProtobufFileIngestor(
    output_path="out/protobuf",
    split_index=False,        # True creates separate nodes.pb + relationships.pb

    repo_path="/path/to/repo"  # Enables path canonicalization for stable IDs

)

# Add a MODULE node

ingestor.ensure_node_batch(
    label=cs.NodeLabel.MODULE.name,
    properties={
        cs.KEY_QUALIFIED_NAME: "my_package.utils",
        cs.KEY_PATH: "my_package/utils.py",
        "docstring": "Utility functions for data processing."
    }
)

# Add a FUNCTION node

ingestor.ensure_node_batch(
    label=cs.NodeLabel.FUNCTION.name,
    properties={
        cs.KEY_QUALIFIED_NAME: "my_package.utils.process_data",
        cs.KEY_PATH: "my_package/utils.py",
        "signature": "def process_data(raw: dict) -> list:"
    }
)

# Create relationship: MODULE DEFINES FUNCTION

ingestor.ensure_relationship_batch(
    from_spec=(cs.NodeLabel.MODULE.name, cs.KEY_QUALIFIED_NAME, "my_package.utils"),
    rel_type="DEFINES",
    to_spec=(cs.NodeLabel.FUNCTION.name, cs.KEY_QUALIFIED_NAME, "my_package.utils.process_data"),
    properties=None
)

# Write to disk

ingestor.flush_all()

```

**Key methods in `ProtobufFileIngestor`:**

- `ensure_node_batch(label, properties)` — Adds or updates a node; batching defers serialization until flush
- `ensure_relationship_batch(from_spec, rel_type, to_spec, properties)` — Creates edges using `(label, key_field, key_value)` tuples to reference endpoints
- `flush_all()` — Finalizes and writes `.pb` files; automatically selects joint or split mode based on `split_index`

### Method 2: Built-in CLI

For standard repository indexing, use the `codebase_rag index` command without writing code.

```bash

# Install with Protobuf support

pip install "code-graph-rag[protobuf]"

# Index to single joint file

codebase_rag index \
    --repo /path/to/source \
    --output ./index_output

# Index to split files (nodes.pb + relationships.pb)

codebase_rag index \
    --repo /path/to/source \
    --output ./index_output \
    --split

```

The CLI in [`codebase_rag/workspaces/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/workspaces/cli.py) orchestrates:

1. Repository scanning via the graph builder
2. `ProtobufFileIngestor` instantiation with parsed arguments
3. Feeding discovered nodes/relationships to `ensure_node_batch` and `ensure_relationship_batch`
4. Calling `flush_all()` to persist results

## Loading Indexed Protobuf Data Offline

Once indexed, the `.pb` files enable fast graph restoration without source code access.

### Joint Mode (Single File)

```python
import codec.schema_pb2 as pb
from pathlib import Path

data = Path("out/protobuf/graph_code_index.pb").read_bytes()
index = pb.GraphCodeIndex()
index.ParseFromString(data)

print(f"Nodes: {len(index.nodes)}")
print(f"Relationships: {len(index.relationships)}")

# Access first node

first_node = index.nodes[0]
print(first_node.qualified_name)  # Access via one-of field

```

### Split Mode (Separate Files)

```python
nodes_data = Path("out/protobuf/nodes.pb").read_bytes()
rels_data = Path("out/protobuf/relationships.pb").read_bytes()

nodes_msg = pb.NodeList()
rels_msg = pb.RelationshipList()
nodes_msg.ParseFromString(nodes_data)
rels_msg.ParseFromString(rels_data)

# Combine manually if needed

combined = pb.GraphCodeIndex()
combined.nodes.extend(nodes_msg.nodes)
combined.relationships.extend(rels_msg.relationships)

```

## Determinism and Caching Benefits

The `ProtobufFileIngestor` implements several mechanisms to ensure reproducible output:

- **Path canonicalization** — Strips `repo_path` prefix from file paths, making IDs portable across machines
- **Sorted serialization** — `_sorted_nodes` and `_sorted_relationships` order elements before writing
- **Deterministic protobuf** — Google's `deterministic=True` flag for `SerializeToString()`

These properties enable content-addressed caching: the same repository commit always produces the same file hash, allowing build systems to skip redundant re-indexing.

## File Reference Guide

| File | Location | Role in Indexing |
|------|----------|------------------|
| [`protobuf_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/protobuf_service.py) | `codebase_rag/services/` | Core `ProtobufFileIngestor` implementation |
| [`schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/schema_pb2.py) | `codec/` | Auto-generated protobuf message classes |
| `constants/` | `codebase_rag/constants/` | `NodeLabel` enum, `LABEL_TO_ONEOF_FIELD` mapping, file name constants |
| [`cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/cli.py) | `codebase_rag/workspaces/` | Entry point for `codebase_rag index` command |
| [`test_protobuf_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/test_protobuf_service.py) | `codebase_rag/tests/` | Verification suite for ingestion logic |

## Summary

- **`ProtobufFileIngestor`** in [`codebase_rag/services/protobuf_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/protobuf_service.py) is the central class for repository-to-protobuf conversion
- **`ensure_node_batch`** and **`ensure_relationship_batch`** accumulate graph data during scanning
- **`flush_all()`** writes deterministic `.pb` files in either joint or split configuration
- The **CLI** at `codebase_rag index` provides zero-code indexing for standard use cases
- **Deterministic serialization** enables reliable caching and offline graph reloading

## Frequently Asked Questions

### What is the difference between joint and split index mode?

**Joint mode** writes a single `graph_code_index.pb` file containing both nodes and relationships. **Split mode** produces separate `nodes.pb` and `relationships.pb` files, useful when you need to load or process nodes and edges independently. Set `split_index=True` in `ProtobufFileIngestor` or pass `--split` to the CLI.

### Why are my node IDs different when indexing the same repository on different machines?

The `ProtobufFileIngestor` uses `repo_path` to canonicalize file paths into stable identifiers. Ensure you instantiate the ingestor with the actual repository root: `ProtobufFileIngestor(..., repo_path="/absolute/path/to/repo")`. This strips the prefix from all paths, making IDs location-independent.

### How do I add custom properties to nodes or relationships?

Pass arbitrary string-keyed properties to `ensure_node_batch` or `ensure_relationship_batch`:

```python
ingestor.ensure_node_batch(
    label=cs.NodeLabel.FUNCTION.name,
    properties={
        cs.KEY_QUALIFIED_NAME: "module.func",
        "custom_metric": "complexity:12",
        "last_modified": "2024-01-15"
    }
)

```

The protobuf schema stores all properties as string key-value pairs.

### Can I extend the protobuf schema with new node or relationship types?

The schema in [`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py) is generated from `.proto` definitions. New `NodeLabel` values require updating the source `.proto` file, regenerating [`schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/schema_pb2.py), and adding corresponding entries to `LABEL_TO_ONEOF_FIELD` in the constants module. The one-of field pattern in `Node` messages ensures type-safe access while maintaining backward compatibility.