# How to Index a Repository with Code-Graph-RAG: A Complete Guide

> Learn how to index a repository with code-graph-rag. This guide details the cgr index command, semantic graph construction, and vector persistence for efficient querying.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-05

---

**To index a repository with code-graph-rag, run the `cgr index` command against your repository root, which triggers a multi-stage pipeline that parses source files, constructs a semantic graph, and persists vectors for natural-language querying.**

Code-Graph-RAG (cgr) is an open-source tool from `vitali87/code-graph-rag` that transforms source code into a **semantic code graph** stored in a vector database. Indexing is the foundational step that enables downstream retrieval-augmented generation (RAG) capabilities.

## The Three-Layer Indexing Architecture

The indexing pipeline operates through three distinct layers, each handled by specific modules in the codebase.

### Ingestion Layer

The ingestion layer walks your file system and extracts code structure. In [`codebase_rag/utils/source_extraction.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/source_extraction.py), language-specific parsers generate abstract syntax trees (ASTs) to discover function definitions, class declarations, and call relationships. These are recorded as **nodes** and **edges** in the graph.

The orchestration happens in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py), where the `GraphUpdater` class manages "re-ingest" operations that remove stale sub-trees and write fresh graph data.

### Graph Construction Layer

Raw ingestion data is transformed into a Neo4j-compatible graph in [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py). This layer resolves fuzzy identifiers and builds property indexes on fields like `qualified_name` and `file_path` to accelerate queries. The [`decorators.py`](https://github.com/vitali87/code-graph-rag/blob/main/decorators.py) module provides helpers for lazy property evaluation during this phase.

### Storage and Retrieval Layer

The **vector store** abstraction in [`codebase_rag/vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/vector_store.py) supports multiple backends including FAISS and Milvus. During indexing, [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) encodes code snippets into dense vectors using your configured LLM, while `VectorStore.batch_write` persists the data.

## Using the CLI to Index a Repository

The primary interface for indexing is the Typer-based CLI defined in [`codebase_rag/graph_cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_cli.py).

### Basic Indexing Command

The signature for indexing follows this pattern:

```bash
cgr index [OPTIONS] PATH

```

Where `PATH` is the root directory of your repository.

**Index a fresh repository:**

```bash
cgr index ./my-project --store ./graph.db

```

This creates a new graph database at `./graph.db` containing the parsed structure of your codebase.

**Force a full re-index:**

```bash
cgr index ./my-project --store ./graph.db --force

```

The `--force` flag skips incremental checks and rebuilds the entire graph from scratch.

**Index a subdirectory (mono-repo support):**

```bash
cgr index ./my-project/services/auth --store ./graph.db

```

**Watch mode for continuous indexing:**

```bash
cgr watch ./my-project --store ./graph.db

```

Watch mode monitors the repository for changes and automatically triggers re-indexing when files are modified.

## Programmatic Indexing with Python

You can trigger indexing directly from Python by instantiating the `GraphUpdater` class from [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py):

```python
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.vector_store import VectorStore

store = VectorStore(path="./graph.db")
updater = GraphUpdater(
    store=store,
    repo_path="./my-project",
    parsers=["python", "java"],
    queries=None
)

# Trigger full re-ingest

updater.reingest(())

```

The `reingest` method handles file deletion detection, content hashing, and batch writes to the vector store.

## The Five-Step Internal Indexing Process

When you execute `cgr index`, the system performs these operations in sequence:

1.  **Lock acquisition** – `GraphUpdater` acquires a file lock to prevent concurrent indexing operations that could corrupt the graph.

2.  **File-system walk** – The `reingest` method enumerates files, computes content hashes, and determines which sub-trees require rebuilding based on changes since the last index.

3.  **Parsing** – Language-specific parsers in [`source_extraction.py`](https://github.com/vitali87/code-graph-rag/blob/main/source_extraction.py) convert source files into ingestor records containing nodes (functions, classes) and relationships (calls, imports).

4.  **Batch write** – The ingestor streams batches to `VectorStore.batch_write`, which persists nodes and edges to your chosen backend.

5.  **Index refresh** – `GraphLoader._build_property_index` rebuilds property indexes (such as `qualified_name`) to ensure subsequent semantic searches remain fast.

### Partial vs. Full Re-indexing

The decision logic for incremental versus full re-indexing resides in [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py). According to the comment on line 114, the system evaluates whether a **partial** update is sufficient or a **full** re-index is required based on the scope of detected changes and the integrity of existing graph components.

## Key Files in the Indexing Pipeline

| File | Role |
|------|------|
| [`codebase_rag/graph_cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_cli.py) | Typer entry point defining the `cgr` command interface |
| [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) | Core logic for deletion handling and batch ingestion |
| [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) | Graph persistence and property index construction |
| [`codebase_rag/utils/source_extraction.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/source_extraction.py) | Language-agnostic AST parsing and symbol extraction |
| [`codebase_rag/vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/vector_store.py) | Backend abstraction for FAISS, Milvus, and other stores |
| [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) | LLM-based vectorization of code fragments |
| [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py) | Determines re-indexing strategy (partial vs. full) |

## Summary

-   Use `cgr index ./path --store ./graph.db` to create a searchable semantic graph from your repository.
-   The pipeline comprises ingestion (AST parsing), graph construction (Neo4j-compatible structure), and vector storage (FAISS/Milvus).
-   `GraphUpdater` in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) orchestrates the process with file locking and batch writes.
-   Incremental indexing is supported, but `--force` triggers a complete rebuild.
-   The `protobuf` capture file `graph_code_index.pb` serves as an intermediate format between parsing and graph materialization.

## Frequently Asked Questions

### What is the difference between incremental and full re-indexing?

Incremental re-indexing updates only modified sub-trees by comparing content hashes, while full re-indexing deletes the existing graph and rebuilds it from scratch. The [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py) module contains logic on line 114 that determines which approach is necessary based on the extent of repository changes and graph integrity checks.

### Which programming languages are supported for parsing?

The parser configuration is controlled via the `--parsers` option or the `parsers` parameter in `GraphUpdater`. The [`source_extraction.py`](https://github.com/vitali87/code-graph-rag/blob/main/source_extraction.py) module handles language-specific AST extraction, though the specific supported languages depend on the tree-sitter grammars available in your installation (commonly including Python, Java, JavaScript, and Go).

### How does code-graph-rag prevent concurrent indexing operations?

The system implements a file-locking mechanism at the start of the indexing process. When `GraphUpdater` initializes a `reingest` operation, it acquires a lock that prevents other `cgr index` processes from running simultaneously, ensuring graph consistency and preventing race conditions during batch writes.

### Can I index only a subdirectory of a repository?

Yes. Pass a subdirectory path instead of the repository root to the `cgr index` command. This is particularly useful for mono-repos where you want to index only specific services or packages. The resulting graph will contain only the symbols and relationships present in that subdirectory and its children.