Best Practices for Indexing Large Monorepos with Code-Graph-RAG: A Complete Guide

Code-Graph-RAG's indexing pipeline uses Tree-sitter parsing, parallel workers, and incremental updates to scale from small repositories to multi-million line monorepos without exhausting memory.

The vitali87/code-graph-rag project constructs language-agnostic knowledge graphs by parsing source files with Tree-sitter and persisting nodes in Memgraph. When targeting large monorepos, unconstrained ingestion triggers memory pressure and excessive parse times. The following practices leverage specific modules in the codebase to keep indexing fast, memory-efficient, and incremental.

Selective Ingestion with Ignore Patterns

Large monorepos contain generated artifacts, vendor directories, and test fixtures that pollute the graph and waste compute cycles. Code-Graph-RAG implements selective ingestion through patterns defined in a .cgrignore file (Git-ignore syntax) or via the --exclude CLI flag.

The ignore logic resides in codebase_rag/utils/path_utils.py, which traverses the repository tree while filtering paths against glob patterns. Exclude node_modules, build/, and dist/ directories before indexing to reduce graph size by 40-60% in typical JavaScript or Python monorepos.

cgr start --repo-path /path/to/huge/monorepo \
          --exclude "*/node_modules/**" "*/build/**" "*/.git/**"

Parallel Parsing to Maximize CPU Utilization

Tree-sitter parsing is CPU-intensive but embarrassingly parallel. The parser runner spawns a process pool controlled by the CG_RAG_MAX_WORKERS setting, defaulting to the CPU count. For dedicated indexing servers, override this via environment variable or CLI flag to saturate available cores.

Configuration resolution happens in codebase_rag/config.py, which centralizes worker counts, batch sizes, and cache locations. Setting --max-workers 16 on a 32-core machine leaves headroom for the Memgraph ingestion thread while maximizing parse throughput.

export CG_RAG_MAX_WORKERS=16
cgr start --repo-path /path/to/huge/monorepo

Incremental Indexing for Fast Updates

Full re-indexing is unsustainable for active monorepos. The cgr trace sub-command records runtime call traces and merges them into the existing graph, while the incremental evaluator in evals/incremental.py demonstrates how to re-index only changed files.

This approach skips unchanged ASTs entirely, reducing update latency from hours to minutes for large codebases. The incremental validator compares file hashes against the serialized cache to determine invalidation.


# After modifying source files

cgr trace --repo-path /path/to/huge/monorepo --incremental

Chunked Project Handling

Monorepos often comprise logical sub-projects (e.g., frontend/, backend/, shared/). Treat these as separate projects within the shared graph by invoking cgr start --repo-path <subdir> for each directory.

The graph schema documented in docs/architecture/graph-schema.md stores a project property on every node, isolating data while preserving cross-project relationships. This prevents namespace collisions and enables project-scoped queries.

cgr start --repo-path /monorepo/frontend --project-name "web-client"
cgr start --repo-path /monorepo/backend --project-name "api-server"

Memory Management and Batch Processing

Streaming ingestion prevents out-of-memory errors when loading millions of nodes. The MemgraphIngestor class in codebase_rag/services/graph_service.py honors the CG_RAG_BATCH_SIZE configuration, committing nodes in configurable chunks rather than holding the entire graph in RAM.

Tune this value based on available memory; smaller batches reduce peak heap usage at the cost of slightly higher transaction overhead. The default configuration provides a safe baseline for 8GB systems.

AST Caching for Reproducible Speed

Tree-sitter tree construction dominates indexing time for large files. The codebase_rag/parsers/ast_cache.py module (invoked indirectly by GraphLoader) serializes parsed ASTs to disk and invalidates entries only when source files change.

This cache survives process restarts, making subsequent indexing runs—whether full or incremental—significantly faster. Ensure the cache directory resides on fast storage (NVMe SSD) to avoid I/O bottlenecks.

Query Optimization with Lookup Indexes

Fast traversal requires efficient node lookup. The GraphLoader._build_property_index method constructs a hash map of qualified names (e.g., module.Class.method) to internal node IDs, implemented in optimize/memory_profile.py.

Building this index once during graph loading accelerates downstream queries by eliminating repeated label scans. The index is maintained in-memory alongside the graph and is regenerated whenever the graph structure changes.

Complete CLI Workflow for Massive Monorepos

Combine these practices to index a multi-million line repository efficiently:


# 1. Start the Memgraph and Qdrant services

cgr daemon up

# 2. Initial index with exclusions and maximum parallelism

cgr start --repo-path /path/to/huge/monorepo \
          --exclude "*/node_modules/**" "*/build/**" "*/vendor/**" \
          --max-workers 16 \
          --project-name "production-monorepo"

# 3. Verify incremental capability (simulates a change)

python -c "
from codebase_rag.graph_loader import GraphLoader
loader = GraphLoader('/path/to/huge/monorepo/graph.json')
print(f'Indexed {len(loader.nodes)} nodes')
"

# 4. Daily incremental update after git pull

cgr trace --repo-path /path/to/huge/monorepo --incremental

Summary

Frequently Asked Questions

How do I exclude generated files from indexing?

Create a .cgrignore file in your repository root using standard Git-ignore syntax. The parser consults codebase_rag/utils/path_utils.py to resolve these patterns, filtering out paths before Tree-sitter ever touches them. You can also pass --exclude flags directly to the CLI for one-off runs without modifying the ignore file.

What environment variables control memory usage during indexing?

Set CG_RAG_BATCH_SIZE to limit how many nodes MemgraphIngestor buffers before committing to the database (found in codebase_rag/services/graph_service.py). Additionally, CG_RAG_MAX_WORKERS (defined in codebase_rag/config.py) controls parser parallelism, which indirectly affects peak memory by limiting concurrent AST construction.

Can I index only specific subdirectories of a monorepo?

Yes. Invoke cgr start --repo-path <subdir> multiple times for each logical component, assigning distinct --project-name values. The graph schema stores a project property on every node, allowing cross-project queries while keeping data isolated. This strategy is documented in docs/architecture/graph-schema.md.

How does incremental indexing determine which files to re-parse?

The incremental evaluator in evals/incremental.py compares file modification times and content hashes against the serialized AST cache in codebase_rag/parsers/ast_cache.py. Only files with mismatched hashes trigger new Tree-sitter parsing; unchanged files reuse cached graph nodes, typically reducing update time by 90% in large repositories.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →