Code-Graph-RAG CLI Commands for Parsing and Querying Codebases: A Complete Guide
Code-Graph-RAG provides two primary CLI commands—cgr index for parsing repositories into portable protobuf graphs and cgr start for querying codebases through an interactive LLM-driven chat interface.
Code-Graph-RAG is an open-source tool that transforms software repositories into queryable knowledge graphs. Built with Typer, the cgr executable exposes specialized CLI commands for parsing and querying codebases, enabling developers to analyze complex projects using natural language. This guide examines the implementation details in vitali87/code-graph-rag to explain how these commands function.
Parsing Codebases with cgr index
The cgr index command converts a local repository into a serialized graph representation. Defined at line 645 in cli.py, this command orchestrates the ingestion pipeline without requiring a live database connection.
How Repository Parsing Works
The parsing workflow executes five distinct phases:
- Repository validation –
_resolve_and_validate_repo()verifies the supplied path exists, is a directory, and contains a.gitfolder (warning if absent). - Ignore pattern handling –
load_ignore_patterns()reads.cgrignorefiles; the CLI can interactively prompt for un-ignoring specific paths. - Ingestor initialization –
ProtobufFileIngestor(output_path=..., split_index=...)prepares the output destination for serialized graph data. - Graph construction –
GraphUpdater.run()traverses the repository tree, executes language-specific parsers, and stores symbols and relationships. - Protobuf export – The system writes a
.protodatabase file that can be version-controlled or loaded later viacgr graph-loader.
The GraphUpdater class in graph_updater.py serves as the core engine, receiving the ingestor, repository path, and parser configurations to populate the graph.
# Parse a repository into a protobuf graph
cgr index \
--repo-path /path/to/my/project \
--output-proto-dir ./graph-output \
--split-index
Querying Codebases with cgr start
The cgr start command enables LLM-driven interrogation of codebases, implemented at line 495 in cli.py. This command synchronizes graph data to Memgraph and launches either an interactive chat session or single-shot queries.
Interactive Chat and Single-Shot Queries
The querying workflow follows this sequence:
- Repository resolution – Derives the project name from the repo path or uses the
--project-nameoverride. - Database preparation –
--cleanempties the Memgraph database;--update-graphforces a fresh sync before accepting queries. - Graph synchronization –
_run_graph_sync()instantiatesGraphUpdaterwith a liveMemgraphIngestorfromconnect_memgraph()inconfig.py, then callsupdater.run()to populate Memgraph with nodes representing functions, calls, and imports. - Query execution –
main_async()launches an asynchronous loop that streams LLM responses while retrieving relevant graph fragments via Cypher queries. - Agent invocation – With
--ask-agent, the system executesmain_single_query()for one-time questions without entering the interactive chat loop.
The vector_store.py module manages embedding storage for semantic search during the querying phase.
# Start interactive chat with fresh graph sync
cgr start \
--repo-path /path/to/my/project \
--update-graph \
--orchestrator gpt-4o-mini \
--cypher gpt-4o-mini
# Single-shot query without interactive mode
cgr start \
--repo-path /path/to/my/project \
--ask-agent "Which functions modify global state?" \
--output-format json
Shared Infrastructure and Configuration
Both CLI commands rely on common utilities defined across the codebase:
load_parsers()inparser_loader.py– Dynamically discovers language-specific parsers (Python, Java, Go, etc.) and registers Cypher queries that enable LLM reasoning.connect_memgraph()inconfig.py– Provides a context manager yielding aMemgraphIngestorwith configurable batch sizes for efficient database writes.graph_loader.py– Contains utilities for loading protobuf graph files and printing summaries via thecgr graph-loaderhelper command.
These shared components ensure consistent behavior whether you are exporting static graphs or querying live databases.
Practical CLI Workflow Examples
Complete workflows combining parsing and querying operations:
# 1. Parse and export a repository
cgr index \
--repo-path ./my-project \
--output-proto-dir ./graphs \
--split-index
# 2. Inspect the exported graph (optional)
cgr graph-loader ./graphs/graph.proto
# 3. Query with automatic graph updates
cgr start \
--repo-path ./my-project \
--update-graph \
--orchestrator gpt-4o-mini
Common flags available to both commands include --exclude, --capture, --interactive-setup, --batch-size, and --orchestrator, providing fine-grained control over ingestion scope and LLM model selection.
Summary
cgr indexparses repositories into portable protobuf graph files usingProtobufFileIngestorandGraphUpdateras defined incli.py.cgr startsynchronizes graphs to Memgraph and launches interactive LLM chat or single-shot queries viamain_async()ormain_single_query().- Both commands utilize
load_parsers()for multi-language support andconnect_memgraph()for database connectivity. - The CLI supports granular control through shared flags for exclusion patterns, batch sizes, and orchestrator selection.
Frequently Asked Questions
What is the difference between cgr index and cgr start?
cgr index performs offline parsing, writing repository structure to protobuf files without requiring a database connection. cgr start requires Memgraph and provides interactive querying capabilities, optionally syncing the graph before launching the LLM interface.
How does Code-Graph-RAG handle large repositories?
The CLI provides --split-index for cgr index to partition large graphs into multiple protobuf files. Additionally, --batch-size controls the number of nodes processed per transaction when syncing to Memgraph, preventing memory exhaustion during ingestion.
Can I query a codebase without parsing it first?
No, the graph must be populated before querying. However, cgr start with --update-graph performs parsing and syncing automatically before launching the chat interface, combining both steps into a single command execution.
Which programming languages are supported by the parser?
The system dynamically discovers language-specific parsers through load_parsers() in parser_loader.py, with built-in support for Python, Java, Go, and other languages. Each parser registers specific Cypher queries that enable the LLM to reason about language-specific constructs like imports, class hierarchies, and function calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →