# Best Practices for Indexing Large Monorepos with Code-Graph-RAG: A Complete Guide

> Master indexing large monorepos with Code-Graph-RAG. Learn best practices for efficient scaling, memory management, and parallel processing in this complete guide.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: best-practices
- Published: 2026-08-19

---

**Code-Graph-RAG's indexing pipeline uses Tree-sitter parsing, parallel workers, and incremental updates to scale from small repositories to multi-million line monorepos without exhausting memory.**

The `vitali87/code-graph-rag` project constructs language-agnostic knowledge graphs by parsing source files with Tree-sitter and persisting nodes in Memgraph. When targeting large monorepos, unconstrained ingestion triggers memory pressure and excessive parse times. The following practices leverage specific modules in the codebase to keep indexing fast, memory-efficient, and incremental.

## Selective Ingestion with Ignore Patterns

Large monorepos contain generated artifacts, vendor directories, and test fixtures that pollute the graph and waste compute cycles. Code-Graph-RAG implements selective ingestion through patterns defined in a `.cgrignore` file (Git-ignore syntax) or via the `--exclude` CLI flag.

The ignore logic resides in **[`codebase_rag/utils/path_utils.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/path_utils.py)**, which traverses the repository tree while filtering paths against glob patterns. Exclude `node_modules`, `build/`, and `dist/` directories before indexing to reduce graph size by 40-60% in typical JavaScript or Python monorepos.

```bash
cgr start --repo-path /path/to/huge/monorepo \
          --exclude "*/node_modules/**" "*/build/**" "*/.git/**"

```

## Parallel Parsing to Maximize CPU Utilization

Tree-sitter parsing is CPU-intensive but embarrassingly parallel. The parser runner spawns a process pool controlled by the **`CG_RAG_MAX_WORKERS`** setting, defaulting to the CPU count. For dedicated indexing servers, override this via environment variable or CLI flag to saturate available cores.

Configuration resolution happens in **[`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py)**, which centralizes worker counts, batch sizes, and cache locations. Setting `--max-workers 16` on a 32-core machine leaves headroom for the Memgraph ingestion thread while maximizing parse throughput.

```bash
export CG_RAG_MAX_WORKERS=16
cgr start --repo-path /path/to/huge/monorepo

```

## Incremental Indexing for Fast Updates

Full re-indexing is unsustainable for active monorepos. The **`cgr trace`** sub-command records runtime call traces and merges them into the existing graph, while the incremental evaluator in **[`evals/incremental.py`](https://github.com/vitali87/code-graph-rag/blob/main/evals/incremental.py)** demonstrates how to re-index only changed files.

This approach skips unchanged ASTs entirely, reducing update latency from hours to minutes for large codebases. The incremental validator compares file hashes against the serialized cache to determine invalidation.

```bash

# After modifying source files

cgr trace --repo-path /path/to/huge/monorepo --incremental

```

## Chunked Project Handling

Monorepos often comprise logical sub-projects (e.g., `frontend/`, `backend/`, `shared/`). Treat these as separate projects within the shared graph by invoking **`cgr start --repo-path <subdir>`** for each directory.

The graph schema documented in **[`docs/architecture/graph-schema.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/architecture/graph-schema.md)** stores a `project` property on every node, isolating data while preserving cross-project relationships. This prevents namespace collisions and enables project-scoped queries.

```bash
cgr start --repo-path /monorepo/frontend --project-name "web-client"
cgr start --repo-path /monorepo/backend --project-name "api-server"

```

## Memory Management and Batch Processing

Streaming ingestion prevents out-of-memory errors when loading millions of nodes. The **`MemgraphIngestor`** class in **[`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py)** honors the **`CG_RAG_BATCH_SIZE`** configuration, committing nodes in configurable chunks rather than holding the entire graph in RAM.

Tune this value based on available memory; smaller batches reduce peak heap usage at the cost of slightly higher transaction overhead. The default configuration provides a safe baseline for 8GB systems.

## AST Caching for Reproducible Speed

Tree-sitter tree construction dominates indexing time for large files. The **[`codebase_rag/parsers/ast_cache.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/ast_cache.py)** module (invoked indirectly by **`GraphLoader`**) serializes parsed ASTs to disk and invalidates entries only when source files change.

This cache survives process restarts, making subsequent indexing runs—whether full or incremental—significantly faster. Ensure the cache directory resides on fast storage (NVMe SSD) to avoid I/O bottlenecks.

## Query Optimization with Lookup Indexes

Fast traversal requires efficient node lookup. The **`GraphLoader._build_property_index`** method constructs a hash map of qualified names (e.g., `module.Class.method`) to internal node IDs, implemented in **[`optimize/memory_profile.py`](https://github.com/vitali87/code-graph-rag/blob/main/optimize/memory_profile.py)**.

Building this index once during graph loading accelerates downstream queries by eliminating repeated label scans. The index is maintained in-memory alongside the graph and is regenerated whenever the graph structure changes.

## Complete CLI Workflow for Massive Monorepos

Combine these practices to index a multi-million line repository efficiently:

```bash

# 1. Start the Memgraph and Qdrant services

cgr daemon up

# 2. Initial index with exclusions and maximum parallelism

cgr start --repo-path /path/to/huge/monorepo \
          --exclude "*/node_modules/**" "*/build/**" "*/vendor/**" \
          --max-workers 16 \
          --project-name "production-monorepo"

# 3. Verify incremental capability (simulates a change)

python -c "
from codebase_rag.graph_loader import GraphLoader
loader = GraphLoader('/path/to/huge/monorepo/graph.json')
print(f'Indexed {len(loader.nodes)} nodes')
"

# 4. Daily incremental update after git pull

cgr trace --repo-path /path/to/huge/monorepo --incremental

```

## Summary

- **Filter early**: Use `.cgrignore` or `--exclude` patterns via [`codebase_rag/utils/path_utils.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/path_utils.py) to skip generated artifacts.
- **Scale horizontally**: Adjust `CG_RAG_MAX_WORKERS` in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py) to utilize all CPU cores for parsing.
- **Update incrementally**: Run `cgr trace --incremental` and reference [`evals/incremental.py`](https://github.com/vitali87/code-graph-rag/blob/main/evals/incremental.py) to avoid full re-parsing.
- **Partition logically**: Pass `--repo-path <subdir>` to create distinct project nodes as defined in [`docs/architecture/graph-schema.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/architecture/graph-schema.md).
- **Manage memory**: Configure `CG_RAG_BATCH_SIZE` to control `MemgraphIngestor` batching in [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py).
- **Cache aggressively**: Leverage [`codebase_rag/parsers/ast_cache.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/ast_cache.py) to reuse Tree-sitter trees across runs.
- **Index for speed**: Build qualified-name lookups using `GraphLoader._build_property_index` from [`optimize/memory_profile.py`](https://github.com/vitali87/code-graph-rag/blob/main/optimize/memory_profile.py).

## Frequently Asked Questions

### How do I exclude generated files from indexing?

Create a `.cgrignore` file in your repository root using standard Git-ignore syntax. The parser consults [`codebase_rag/utils/path_utils.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/path_utils.py) to resolve these patterns, filtering out paths before Tree-sitter ever touches them. You can also pass `--exclude` flags directly to the CLI for one-off runs without modifying the ignore file.

### What environment variables control memory usage during indexing?

Set `CG_RAG_BATCH_SIZE` to limit how many nodes `MemgraphIngestor` buffers before committing to the database (found in [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py)). Additionally, `CG_RAG_MAX_WORKERS` (defined in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py)) controls parser parallelism, which indirectly affects peak memory by limiting concurrent AST construction.

### Can I index only specific subdirectories of a monorepo?

Yes. Invoke `cgr start --repo-path <subdir>` multiple times for each logical component, assigning distinct `--project-name` values. The graph schema stores a `project` property on every node, allowing cross-project queries while keeping data isolated. This strategy is documented in [`docs/architecture/graph-schema.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/architecture/graph-schema.md).

### How does incremental indexing determine which files to re-parse?

The incremental evaluator in [`evals/incremental.py`](https://github.com/vitali87/code-graph-rag/blob/main/evals/incremental.py) compares file modification times and content hashes against the serialized AST cache in [`codebase_rag/parsers/ast_cache.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/ast_cache.py). Only files with mismatched hashes trigger new Tree-sitter parsing; unchanged files reuse cached graph nodes, typically reducing update time by 90% in large repositories.