Best Practices for Indexing Large Monorepos with Code-Graph-RAG: A Complete Guide
Code-Graph-RAG's indexing pipeline uses Tree-sitter parsing, parallel workers, and incremental updates to scale from small repositories to multi-million line monorepos without exhausting memory.
The vitali87/code-graph-rag project constructs language-agnostic knowledge graphs by parsing source files with Tree-sitter and persisting nodes in Memgraph. When targeting large monorepos, unconstrained ingestion triggers memory pressure and excessive parse times. The following practices leverage specific modules in the codebase to keep indexing fast, memory-efficient, and incremental.
Selective Ingestion with Ignore Patterns
Large monorepos contain generated artifacts, vendor directories, and test fixtures that pollute the graph and waste compute cycles. Code-Graph-RAG implements selective ingestion through patterns defined in a .cgrignore file (Git-ignore syntax) or via the --exclude CLI flag.
The ignore logic resides in codebase_rag/utils/path_utils.py, which traverses the repository tree while filtering paths against glob patterns. Exclude node_modules, build/, and dist/ directories before indexing to reduce graph size by 40-60% in typical JavaScript or Python monorepos.
cgr start --repo-path /path/to/huge/monorepo \
--exclude "*/node_modules/**" "*/build/**" "*/.git/**"
Parallel Parsing to Maximize CPU Utilization
Tree-sitter parsing is CPU-intensive but embarrassingly parallel. The parser runner spawns a process pool controlled by the CG_RAG_MAX_WORKERS setting, defaulting to the CPU count. For dedicated indexing servers, override this via environment variable or CLI flag to saturate available cores.
Configuration resolution happens in codebase_rag/config.py, which centralizes worker counts, batch sizes, and cache locations. Setting --max-workers 16 on a 32-core machine leaves headroom for the Memgraph ingestion thread while maximizing parse throughput.
export CG_RAG_MAX_WORKERS=16
cgr start --repo-path /path/to/huge/monorepo
Incremental Indexing for Fast Updates
Full re-indexing is unsustainable for active monorepos. The cgr trace sub-command records runtime call traces and merges them into the existing graph, while the incremental evaluator in evals/incremental.py demonstrates how to re-index only changed files.
This approach skips unchanged ASTs entirely, reducing update latency from hours to minutes for large codebases. The incremental validator compares file hashes against the serialized cache to determine invalidation.
# After modifying source files
cgr trace --repo-path /path/to/huge/monorepo --incremental
Chunked Project Handling
Monorepos often comprise logical sub-projects (e.g., frontend/, backend/, shared/). Treat these as separate projects within the shared graph by invoking cgr start --repo-path <subdir> for each directory.
The graph schema documented in docs/architecture/graph-schema.md stores a project property on every node, isolating data while preserving cross-project relationships. This prevents namespace collisions and enables project-scoped queries.
cgr start --repo-path /monorepo/frontend --project-name "web-client"
cgr start --repo-path /monorepo/backend --project-name "api-server"
Memory Management and Batch Processing
Streaming ingestion prevents out-of-memory errors when loading millions of nodes. The MemgraphIngestor class in codebase_rag/services/graph_service.py honors the CG_RAG_BATCH_SIZE configuration, committing nodes in configurable chunks rather than holding the entire graph in RAM.
Tune this value based on available memory; smaller batches reduce peak heap usage at the cost of slightly higher transaction overhead. The default configuration provides a safe baseline for 8GB systems.
AST Caching for Reproducible Speed
Tree-sitter tree construction dominates indexing time for large files. The codebase_rag/parsers/ast_cache.py module (invoked indirectly by GraphLoader) serializes parsed ASTs to disk and invalidates entries only when source files change.
This cache survives process restarts, making subsequent indexing runs—whether full or incremental—significantly faster. Ensure the cache directory resides on fast storage (NVMe SSD) to avoid I/O bottlenecks.
Query Optimization with Lookup Indexes
Fast traversal requires efficient node lookup. The GraphLoader._build_property_index method constructs a hash map of qualified names (e.g., module.Class.method) to internal node IDs, implemented in optimize/memory_profile.py.
Building this index once during graph loading accelerates downstream queries by eliminating repeated label scans. The index is maintained in-memory alongside the graph and is regenerated whenever the graph structure changes.
Complete CLI Workflow for Massive Monorepos
Combine these practices to index a multi-million line repository efficiently:
# 1. Start the Memgraph and Qdrant services
cgr daemon up
# 2. Initial index with exclusions and maximum parallelism
cgr start --repo-path /path/to/huge/monorepo \
--exclude "*/node_modules/**" "*/build/**" "*/vendor/**" \
--max-workers 16 \
--project-name "production-monorepo"
# 3. Verify incremental capability (simulates a change)
python -c "
from codebase_rag.graph_loader import GraphLoader
loader = GraphLoader('/path/to/huge/monorepo/graph.json')
print(f'Indexed {len(loader.nodes)} nodes')
"
# 4. Daily incremental update after git pull
cgr trace --repo-path /path/to/huge/monorepo --incremental
Summary
- Filter early: Use
.cgrignoreor--excludepatterns viacodebase_rag/utils/path_utils.pyto skip generated artifacts. - Scale horizontally: Adjust
CG_RAG_MAX_WORKERSincodebase_rag/config.pyto utilize all CPU cores for parsing. - Update incrementally: Run
cgr trace --incrementaland referenceevals/incremental.pyto avoid full re-parsing. - Partition logically: Pass
--repo-path <subdir>to create distinct project nodes as defined indocs/architecture/graph-schema.md. - Manage memory: Configure
CG_RAG_BATCH_SIZEto controlMemgraphIngestorbatching incodebase_rag/services/graph_service.py. - Cache aggressively: Leverage
codebase_rag/parsers/ast_cache.pyto reuse Tree-sitter trees across runs. - Index for speed: Build qualified-name lookups using
GraphLoader._build_property_indexfromoptimize/memory_profile.py.
Frequently Asked Questions
How do I exclude generated files from indexing?
Create a .cgrignore file in your repository root using standard Git-ignore syntax. The parser consults codebase_rag/utils/path_utils.py to resolve these patterns, filtering out paths before Tree-sitter ever touches them. You can also pass --exclude flags directly to the CLI for one-off runs without modifying the ignore file.
What environment variables control memory usage during indexing?
Set CG_RAG_BATCH_SIZE to limit how many nodes MemgraphIngestor buffers before committing to the database (found in codebase_rag/services/graph_service.py). Additionally, CG_RAG_MAX_WORKERS (defined in codebase_rag/config.py) controls parser parallelism, which indirectly affects peak memory by limiting concurrent AST construction.
Can I index only specific subdirectories of a monorepo?
Yes. Invoke cgr start --repo-path <subdir> multiple times for each logical component, assigning distinct --project-name values. The graph schema stores a project property on every node, allowing cross-project queries while keeping data isolated. This strategy is documented in docs/architecture/graph-schema.md.
How does incremental indexing determine which files to re-parse?
The incremental evaluator in evals/incremental.py compares file modification times and content hashes against the serialized AST cache in codebase_rag/parsers/ast_cache.py. Only files with mismatched hashes trigger new Tree-sitter parsing; unchanged files reuse cached graph nodes, typically reducing update time by 90% in large repositories.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →