How to Index a Repository to Protobuf for Offline Use with Code-Graph-RAG

Code-Graph-RAG uses a deterministic Protobuf representation to store repository indexes so graphs can be reloaded without re-parsing source files.

The ProtobufFileIngestor class in codebase_rag/services/protobuf_service.py handles the entire workflow: collecting nodes and relationships during repository scanning, normalizing paths for canonical IDs, and serializing the result into portable .pb files. This article covers the complete indexing process, from direct Python usage to CLI commands and offline loading.

Architecture Overview

Code-Graph-RAG's Protobuf indexing system consists of four interconnected components:

Component Source File Purpose
Protobuf schema codec/schema_pb2.py Defines wire-format messages (GraphCodeIndex, Node, Relationship)
Constants codebase_rag/constants/* Enumerations (NodeLabel), one-of field mappings, output file names
ProtobufFileIngestor codebase_rag/services/protobuf_service.py Accumulates, sorts, and serializes graph data
CLI interface codebase_rag/workspaces/cli.py User-facing codebase_rag index command

The ingestor guarantees deterministic output through sorted keys and SerializeToString(deterministic=True), ensuring identical repositories produce byte-identical files regardless of checkout location.

Indexing Methods

Method 1: Direct Python API

For fine-grained control over the indexing process, instantiate ProtobufFileIngestor directly and feed it nodes and relationships as you discover them.

from pathlib import Path
from codebase_rag.services.protobuf_service import ProtobufFileIngestor
import codebase_rag.constants as cs

# Initialize ingestor

ingestor = ProtobufFileIngestor(
    output_path="out/protobuf",
    split_index=False,        # True creates separate nodes.pb + relationships.pb

    repo_path="/path/to/repo"  # Enables path canonicalization for stable IDs

)

# Add a MODULE node

ingestor.ensure_node_batch(
    label=cs.NodeLabel.MODULE.name,
    properties={
        cs.KEY_QUALIFIED_NAME: "my_package.utils",
        cs.KEY_PATH: "my_package/utils.py",
        "docstring": "Utility functions for data processing."
    }
)

# Add a FUNCTION node

ingestor.ensure_node_batch(
    label=cs.NodeLabel.FUNCTION.name,
    properties={
        cs.KEY_QUALIFIED_NAME: "my_package.utils.process_data",
        cs.KEY_PATH: "my_package/utils.py",
        "signature": "def process_data(raw: dict) -> list:"
    }
)

# Create relationship: MODULE DEFINES FUNCTION

ingestor.ensure_relationship_batch(
    from_spec=(cs.NodeLabel.MODULE.name, cs.KEY_QUALIFIED_NAME, "my_package.utils"),
    rel_type="DEFINES",
    to_spec=(cs.NodeLabel.FUNCTION.name, cs.KEY_QUALIFIED_NAME, "my_package.utils.process_data"),
    properties=None
)

# Write to disk

ingestor.flush_all()

Key methods in ProtobufFileIngestor:

  • ensure_node_batch(label, properties) — Adds or updates a node; batching defers serialization until flush
  • ensure_relationship_batch(from_spec, rel_type, to_spec, properties) — Creates edges using (label, key_field, key_value) tuples to reference endpoints
  • flush_all() — Finalizes and writes .pb files; automatically selects joint or split mode based on split_index

Method 2: Built-in CLI

For standard repository indexing, use the codebase_rag index command without writing code.


# Install with Protobuf support

pip install "code-graph-rag[protobuf]"

# Index to single joint file

codebase_rag index \
    --repo /path/to/source \
    --output ./index_output

# Index to split files (nodes.pb + relationships.pb)

codebase_rag index \
    --repo /path/to/source \
    --output ./index_output \
    --split

The CLI in codebase_rag/workspaces/cli.py orchestrates:

  1. Repository scanning via the graph builder
  2. ProtobufFileIngestor instantiation with parsed arguments
  3. Feeding discovered nodes/relationships to ensure_node_batch and ensure_relationship_batch
  4. Calling flush_all() to persist results

Loading Indexed Protobuf Data Offline

Once indexed, the .pb files enable fast graph restoration without source code access.

Joint Mode (Single File)

import codec.schema_pb2 as pb
from pathlib import Path

data = Path("out/protobuf/graph_code_index.pb").read_bytes()
index = pb.GraphCodeIndex()
index.ParseFromString(data)

print(f"Nodes: {len(index.nodes)}")
print(f"Relationships: {len(index.relationships)}")

# Access first node

first_node = index.nodes[0]
print(first_node.qualified_name)  # Access via one-of field

Split Mode (Separate Files)

nodes_data = Path("out/protobuf/nodes.pb").read_bytes()
rels_data = Path("out/protobuf/relationships.pb").read_bytes()

nodes_msg = pb.NodeList()
rels_msg = pb.RelationshipList()
nodes_msg.ParseFromString(nodes_data)
rels_msg.ParseFromString(rels_data)

# Combine manually if needed

combined = pb.GraphCodeIndex()
combined.nodes.extend(nodes_msg.nodes)
combined.relationships.extend(rels_msg.relationships)

Determinism and Caching Benefits

The ProtobufFileIngestor implements several mechanisms to ensure reproducible output:

  • Path canonicalization — Strips repo_path prefix from file paths, making IDs portable across machines
  • Sorted serialization — _sorted_nodes and _sorted_relationships order elements before writing
  • Deterministic protobuf — Google's deterministic=True flag for SerializeToString()

These properties enable content-addressed caching: the same repository commit always produces the same file hash, allowing build systems to skip redundant re-indexing.

File Reference Guide

File Location Role in Indexing
protobuf_service.py codebase_rag/services/ Core ProtobufFileIngestor implementation
schema_pb2.py codec/ Auto-generated protobuf message classes
constants/ codebase_rag/constants/ NodeLabel enum, LABEL_TO_ONEOF_FIELD mapping, file name constants
cli.py codebase_rag/workspaces/ Entry point for codebase_rag index command
test_protobuf_service.py codebase_rag/tests/ Verification suite for ingestion logic

Summary

  • ProtobufFileIngestor in codebase_rag/services/protobuf_service.py is the central class for repository-to-protobuf conversion
  • ensure_node_batch and ensure_relationship_batch accumulate graph data during scanning
  • flush_all() writes deterministic .pb files in either joint or split configuration
  • The CLI at codebase_rag index provides zero-code indexing for standard use cases
  • Deterministic serialization enables reliable caching and offline graph reloading

Frequently Asked Questions

What is the difference between joint and split index mode?

Joint mode writes a single graph_code_index.pb file containing both nodes and relationships. Split mode produces separate nodes.pb and relationships.pb files, useful when you need to load or process nodes and edges independently. Set split_index=True in ProtobufFileIngestor or pass --split to the CLI.

Why are my node IDs different when indexing the same repository on different machines?

The ProtobufFileIngestor uses repo_path to canonicalize file paths into stable identifiers. Ensure you instantiate the ingestor with the actual repository root: ProtobufFileIngestor(..., repo_path="/absolute/path/to/repo"). This strips the prefix from all paths, making IDs location-independent.

How do I add custom properties to nodes or relationships?

Pass arbitrary string-keyed properties to ensure_node_batch or ensure_relationship_batch:

ingestor.ensure_node_batch(
    label=cs.NodeLabel.FUNCTION.name,
    properties={
        cs.KEY_QUALIFIED_NAME: "module.func",
        "custom_metric": "complexity:12",
        "last_modified": "2024-01-15"
    }
)

The protobuf schema stores all properties as string key-value pairs.

Can I extend the protobuf schema with new node or relationship types?

The schema in codec/schema_pb2.py is generated from .proto definitions. New NodeLabel values require updating the source .proto file, regenerating schema_pb2.py, and adding corresponding entries to LABEL_TO_ONEOF_FIELD in the constants module. The one-of field pattern in Node messages ensures type-safe access while maintaining backward compatibility.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →