Code-Graph-RAG CLI Commands for Parsing and Querying Codebases: A Complete Guide

Code-Graph-RAG provides two primary CLI commands—cgr index for parsing repositories into portable protobuf graphs and cgr start for querying codebases through an interactive LLM-driven chat interface.

Code-Graph-RAG is an open-source tool that transforms software repositories into queryable knowledge graphs. Built with Typer, the cgr executable exposes specialized CLI commands for parsing and querying codebases, enabling developers to analyze complex projects using natural language. This guide examines the implementation details in vitali87/code-graph-rag to explain how these commands function.

Parsing Codebases with cgr index

The cgr index command converts a local repository into a serialized graph representation. Defined at line 645 in cli.py, this command orchestrates the ingestion pipeline without requiring a live database connection.

How Repository Parsing Works

The parsing workflow executes five distinct phases:

  1. Repository validation – _resolve_and_validate_repo() verifies the supplied path exists, is a directory, and contains a .git folder (warning if absent).
  2. Ignore pattern handling – load_ignore_patterns() reads .cgrignore files; the CLI can interactively prompt for un-ignoring specific paths.
  3. Ingestor initialization – ProtobufFileIngestor(output_path=..., split_index=...) prepares the output destination for serialized graph data.
  4. Graph construction – GraphUpdater.run() traverses the repository tree, executes language-specific parsers, and stores symbols and relationships.
  5. Protobuf export – The system writes a .proto database file that can be version-controlled or loaded later via cgr graph-loader.

The GraphUpdater class in graph_updater.py serves as the core engine, receiving the ingestor, repository path, and parser configurations to populate the graph.


# Parse a repository into a protobuf graph

cgr index \
  --repo-path /path/to/my/project \
  --output-proto-dir ./graph-output \
  --split-index

Querying Codebases with cgr start

The cgr start command enables LLM-driven interrogation of codebases, implemented at line 495 in cli.py. This command synchronizes graph data to Memgraph and launches either an interactive chat session or single-shot queries.

Interactive Chat and Single-Shot Queries

The querying workflow follows this sequence:

  1. Repository resolution – Derives the project name from the repo path or uses the --project-name override.
  2. Database preparation – --clean empties the Memgraph database; --update-graph forces a fresh sync before accepting queries.
  3. Graph synchronization – _run_graph_sync() instantiates GraphUpdater with a live MemgraphIngestor from connect_memgraph() in config.py, then calls updater.run() to populate Memgraph with nodes representing functions, calls, and imports.
  4. Query execution – main_async() launches an asynchronous loop that streams LLM responses while retrieving relevant graph fragments via Cypher queries.
  5. Agent invocation – With --ask-agent, the system executes main_single_query() for one-time questions without entering the interactive chat loop.

The vector_store.py module manages embedding storage for semantic search during the querying phase.


# Start interactive chat with fresh graph sync

cgr start \
  --repo-path /path/to/my/project \
  --update-graph \
  --orchestrator gpt-4o-mini \
  --cypher gpt-4o-mini

# Single-shot query without interactive mode

cgr start \
  --repo-path /path/to/my/project \
  --ask-agent "Which functions modify global state?" \
  --output-format json

Shared Infrastructure and Configuration

Both CLI commands rely on common utilities defined across the codebase:

  • load_parsers() in parser_loader.py – Dynamically discovers language-specific parsers (Python, Java, Go, etc.) and registers Cypher queries that enable LLM reasoning.
  • connect_memgraph() in config.py – Provides a context manager yielding a MemgraphIngestor with configurable batch sizes for efficient database writes.
  • graph_loader.py – Contains utilities for loading protobuf graph files and printing summaries via the cgr graph-loader helper command.

These shared components ensure consistent behavior whether you are exporting static graphs or querying live databases.

Practical CLI Workflow Examples

Complete workflows combining parsing and querying operations:


# 1. Parse and export a repository

cgr index \
  --repo-path ./my-project \
  --output-proto-dir ./graphs \
  --split-index

# 2. Inspect the exported graph (optional)

cgr graph-loader ./graphs/graph.proto

# 3. Query with automatic graph updates

cgr start \
  --repo-path ./my-project \
  --update-graph \
  --orchestrator gpt-4o-mini

Common flags available to both commands include --exclude, --capture, --interactive-setup, --batch-size, and --orchestrator, providing fine-grained control over ingestion scope and LLM model selection.

Summary

  • cgr index parses repositories into portable protobuf graph files using ProtobufFileIngestor and GraphUpdater as defined in cli.py.
  • cgr start synchronizes graphs to Memgraph and launches interactive LLM chat or single-shot queries via main_async() or main_single_query().
  • Both commands utilize load_parsers() for multi-language support and connect_memgraph() for database connectivity.
  • The CLI supports granular control through shared flags for exclusion patterns, batch sizes, and orchestrator selection.

Frequently Asked Questions

What is the difference between cgr index and cgr start?

cgr index performs offline parsing, writing repository structure to protobuf files without requiring a database connection. cgr start requires Memgraph and provides interactive querying capabilities, optionally syncing the graph before launching the LLM interface.

How does Code-Graph-RAG handle large repositories?

The CLI provides --split-index for cgr index to partition large graphs into multiple protobuf files. Additionally, --batch-size controls the number of nodes processed per transaction when syncing to Memgraph, preventing memory exhaustion during ingestion.

Can I query a codebase without parsing it first?

No, the graph must be populated before querying. However, cgr start with --update-graph performs parsing and syncing automatically before launching the chat interface, combining both steps into a single command execution.

Which programming languages are supported by the parser?

The system dynamically discovers language-specific parsers through load_parsers() in parser_loader.py, with built-in support for Python, Java, Go, and other languages. Each parser registers specific Cypher queries that enable the LLM to reason about language-specific constructs like imports, class hierarchies, and function calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →