How to Index a Repository with Code-Graph-RAG: A Complete Guide

To index a repository with code-graph-rag, run the cgr index command against your repository root, which triggers a multi-stage pipeline that parses source files, constructs a semantic graph, and persists vectors for natural-language querying.

Code-Graph-RAG (cgr) is an open-source tool from vitali87/code-graph-rag that transforms source code into a semantic code graph stored in a vector database. Indexing is the foundational step that enables downstream retrieval-augmented generation (RAG) capabilities.

The Three-Layer Indexing Architecture

The indexing pipeline operates through three distinct layers, each handled by specific modules in the codebase.

Ingestion Layer

The ingestion layer walks your file system and extracts code structure. In codebase_rag/utils/source_extraction.py, language-specific parsers generate abstract syntax trees (ASTs) to discover function definitions, class declarations, and call relationships. These are recorded as nodes and edges in the graph.

The orchestration happens in codebase_rag/graph_updater.py, where the GraphUpdater class manages "re-ingest" operations that remove stale sub-trees and write fresh graph data.

Graph Construction Layer

Raw ingestion data is transformed into a Neo4j-compatible graph in codebase_rag/graph_loader.py. This layer resolves fuzzy identifiers and builds property indexes on fields like qualified_name and file_path to accelerate queries. The decorators.py module provides helpers for lazy property evaluation during this phase.

Storage and Retrieval Layer

The vector store abstraction in codebase_rag/vector_store.py supports multiple backends including FAISS and Milvus. During indexing, codebase_rag/embedder.py encodes code snippets into dense vectors using your configured LLM, while VectorStore.batch_write persists the data.

Using the CLI to Index a Repository

The primary interface for indexing is the Typer-based CLI defined in codebase_rag/graph_cli.py.

Basic Indexing Command

The signature for indexing follows this pattern:

cgr index [OPTIONS] PATH

Where PATH is the root directory of your repository.

Index a fresh repository:

cgr index ./my-project --store ./graph.db

This creates a new graph database at ./graph.db containing the parsed structure of your codebase.

Force a full re-index:

cgr index ./my-project --store ./graph.db --force

The --force flag skips incremental checks and rebuilds the entire graph from scratch.

Index a subdirectory (mono-repo support):

cgr index ./my-project/services/auth --store ./graph.db

Watch mode for continuous indexing:

cgr watch ./my-project --store ./graph.db

Watch mode monitors the repository for changes and automatically triggers re-indexing when files are modified.

Programmatic Indexing with Python

You can trigger indexing directly from Python by instantiating the GraphUpdater class from codebase_rag/graph_updater.py:

from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.vector_store import VectorStore

store = VectorStore(path="./graph.db")
updater = GraphUpdater(
    store=store,
    repo_path="./my-project",
    parsers=["python", "java"],
    queries=None
)

# Trigger full re-ingest

updater.reingest(())

The reingest method handles file deletion detection, content hashing, and batch writes to the vector store.

The Five-Step Internal Indexing Process

When you execute cgr index, the system performs these operations in sequence:

  1. Lock acquisition – GraphUpdater acquires a file lock to prevent concurrent indexing operations that could corrupt the graph.

  2. File-system walk – The reingest method enumerates files, computes content hashes, and determines which sub-trees require rebuilding based on changes since the last index.

  3. Parsing – Language-specific parsers in source_extraction.py convert source files into ingestor records containing nodes (functions, classes) and relationships (calls, imports).

  4. Batch write – The ingestor streams batches to VectorStore.batch_write, which persists nodes and edges to your chosen backend.

  5. Index refresh – GraphLoader._build_property_index rebuilds property indexes (such as qualified_name) to ensure subsequent semantic searches remain fast.

Partial vs. Full Re-indexing

The decision logic for incremental versus full re-indexing resides in realtime_updater.py. According to the comment on line 114, the system evaluates whether a partial update is sufficient or a full re-index is required based on the scope of detected changes and the integrity of existing graph components.

Key Files in the Indexing Pipeline

File Role
codebase_rag/graph_cli.py Typer entry point defining the cgr command interface
codebase_rag/graph_updater.py Core logic for deletion handling and batch ingestion
codebase_rag/graph_loader.py Graph persistence and property index construction
codebase_rag/utils/source_extraction.py Language-agnostic AST parsing and symbol extraction
codebase_rag/vector_store.py Backend abstraction for FAISS, Milvus, and other stores
codebase_rag/embedder.py LLM-based vectorization of code fragments
realtime_updater.py Determines re-indexing strategy (partial vs. full)

Summary

  • Use cgr index ./path --store ./graph.db to create a searchable semantic graph from your repository.
  • The pipeline comprises ingestion (AST parsing), graph construction (Neo4j-compatible structure), and vector storage (FAISS/Milvus).
  • GraphUpdater in codebase_rag/graph_updater.py orchestrates the process with file locking and batch writes.
  • Incremental indexing is supported, but --force triggers a complete rebuild.
  • The protobuf capture file graph_code_index.pb serves as an intermediate format between parsing and graph materialization.

Frequently Asked Questions

What is the difference between incremental and full re-indexing?

Incremental re-indexing updates only modified sub-trees by comparing content hashes, while full re-indexing deletes the existing graph and rebuilds it from scratch. The realtime_updater.py module contains logic on line 114 that determines which approach is necessary based on the extent of repository changes and graph integrity checks.

Which programming languages are supported for parsing?

The parser configuration is controlled via the --parsers option or the parsers parameter in GraphUpdater. The source_extraction.py module handles language-specific AST extraction, though the specific supported languages depend on the tree-sitter grammars available in your installation (commonly including Python, Java, JavaScript, and Go).

How does code-graph-rag prevent concurrent indexing operations?

The system implements a file-locking mechanism at the start of the indexing process. When GraphUpdater initializes a reingest operation, it acquires a lock that prevents other cgr index processes from running simultaneously, ensuring graph consistency and preventing race conditions during batch writes.

Can I index only a subdirectory of a repository?

Yes. Pass a subdirectory path instead of the repository root to the cgr index command. This is particularly useful for mono-repos where you want to index only specific services or packages. The resulting graph will contain only the symbols and relationships present in that subdirectory and its children.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →