How to Index a Repository with Code-Graph-RAG: A Complete Guide
To index a repository with code-graph-rag, run the cgr index command against your repository root, which triggers a multi-stage pipeline that parses source files, constructs a semantic graph, and persists vectors for natural-language querying.
Code-Graph-RAG (cgr) is an open-source tool from vitali87/code-graph-rag that transforms source code into a semantic code graph stored in a vector database. Indexing is the foundational step that enables downstream retrieval-augmented generation (RAG) capabilities.
The Three-Layer Indexing Architecture
The indexing pipeline operates through three distinct layers, each handled by specific modules in the codebase.
Ingestion Layer
The ingestion layer walks your file system and extracts code structure. In codebase_rag/utils/source_extraction.py, language-specific parsers generate abstract syntax trees (ASTs) to discover function definitions, class declarations, and call relationships. These are recorded as nodes and edges in the graph.
The orchestration happens in codebase_rag/graph_updater.py, where the GraphUpdater class manages "re-ingest" operations that remove stale sub-trees and write fresh graph data.
Graph Construction Layer
Raw ingestion data is transformed into a Neo4j-compatible graph in codebase_rag/graph_loader.py. This layer resolves fuzzy identifiers and builds property indexes on fields like qualified_name and file_path to accelerate queries. The decorators.py module provides helpers for lazy property evaluation during this phase.
Storage and Retrieval Layer
The vector store abstraction in codebase_rag/vector_store.py supports multiple backends including FAISS and Milvus. During indexing, codebase_rag/embedder.py encodes code snippets into dense vectors using your configured LLM, while VectorStore.batch_write persists the data.
Using the CLI to Index a Repository
The primary interface for indexing is the Typer-based CLI defined in codebase_rag/graph_cli.py.
Basic Indexing Command
The signature for indexing follows this pattern:
cgr index [OPTIONS] PATH
Where PATH is the root directory of your repository.
Index a fresh repository:
cgr index ./my-project --store ./graph.db
This creates a new graph database at ./graph.db containing the parsed structure of your codebase.
Force a full re-index:
cgr index ./my-project --store ./graph.db --force
The --force flag skips incremental checks and rebuilds the entire graph from scratch.
Index a subdirectory (mono-repo support):
cgr index ./my-project/services/auth --store ./graph.db
Watch mode for continuous indexing:
cgr watch ./my-project --store ./graph.db
Watch mode monitors the repository for changes and automatically triggers re-indexing when files are modified.
Programmatic Indexing with Python
You can trigger indexing directly from Python by instantiating the GraphUpdater class from codebase_rag/graph_updater.py:
from codebase_rag.graph_updater import GraphUpdater
from codebase_rag.vector_store import VectorStore
store = VectorStore(path="./graph.db")
updater = GraphUpdater(
store=store,
repo_path="./my-project",
parsers=["python", "java"],
queries=None
)
# Trigger full re-ingest
updater.reingest(())
The reingest method handles file deletion detection, content hashing, and batch writes to the vector store.
The Five-Step Internal Indexing Process
When you execute cgr index, the system performs these operations in sequence:
-
Lock acquisition –
GraphUpdateracquires a file lock to prevent concurrent indexing operations that could corrupt the graph. -
File-system walk – The
reingestmethod enumerates files, computes content hashes, and determines which sub-trees require rebuilding based on changes since the last index. -
Parsing – Language-specific parsers in
source_extraction.pyconvert source files into ingestor records containing nodes (functions, classes) and relationships (calls, imports). -
Batch write – The ingestor streams batches to
VectorStore.batch_write, which persists nodes and edges to your chosen backend. -
Index refresh –
GraphLoader._build_property_indexrebuilds property indexes (such asqualified_name) to ensure subsequent semantic searches remain fast.
Partial vs. Full Re-indexing
The decision logic for incremental versus full re-indexing resides in realtime_updater.py. According to the comment on line 114, the system evaluates whether a partial update is sufficient or a full re-index is required based on the scope of detected changes and the integrity of existing graph components.
Key Files in the Indexing Pipeline
| File | Role |
|---|---|
codebase_rag/graph_cli.py |
Typer entry point defining the cgr command interface |
codebase_rag/graph_updater.py |
Core logic for deletion handling and batch ingestion |
codebase_rag/graph_loader.py |
Graph persistence and property index construction |
codebase_rag/utils/source_extraction.py |
Language-agnostic AST parsing and symbol extraction |
codebase_rag/vector_store.py |
Backend abstraction for FAISS, Milvus, and other stores |
codebase_rag/embedder.py |
LLM-based vectorization of code fragments |
realtime_updater.py |
Determines re-indexing strategy (partial vs. full) |
Summary
- Use
cgr index ./path --store ./graph.dbto create a searchable semantic graph from your repository. - The pipeline comprises ingestion (AST parsing), graph construction (Neo4j-compatible structure), and vector storage (FAISS/Milvus).
GraphUpdaterincodebase_rag/graph_updater.pyorchestrates the process with file locking and batch writes.- Incremental indexing is supported, but
--forcetriggers a complete rebuild. - The
protobufcapture filegraph_code_index.pbserves as an intermediate format between parsing and graph materialization.
Frequently Asked Questions
What is the difference between incremental and full re-indexing?
Incremental re-indexing updates only modified sub-trees by comparing content hashes, while full re-indexing deletes the existing graph and rebuilds it from scratch. The realtime_updater.py module contains logic on line 114 that determines which approach is necessary based on the extent of repository changes and graph integrity checks.
Which programming languages are supported for parsing?
The parser configuration is controlled via the --parsers option or the parsers parameter in GraphUpdater. The source_extraction.py module handles language-specific AST extraction, though the specific supported languages depend on the tree-sitter grammars available in your installation (commonly including Python, Java, JavaScript, and Go).
How does code-graph-rag prevent concurrent indexing operations?
The system implements a file-locking mechanism at the start of the indexing process. When GraphUpdater initializes a reingest operation, it acquires a lock that prevents other cgr index processes from running simultaneously, ensuring graph consistency and preventing race conditions during batch writes.
Can I index only a subdirectory of a repository?
Yes. Pass a subdirectory path instead of the repository root to the cgr index command. This is particularly useful for mono-repos where you want to index only specific services or packages. The resulting graph will contain only the symbols and relationships present in that subdirectory and its children.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →