# Code-Graph-RAG CLI Commands for Parsing and Querying Codebases: A Complete Guide

> Discover Code-Graph-RAG CLI commands for parsing codebases into graphs with cgr index and querying with cgr start. Explore your code interactively with this complete guide.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-18

---

**Code-Graph-RAG provides two primary CLI commands—`cgr index` for parsing repositories into portable protobuf graphs and `cgr start` for querying codebases through an interactive LLM-driven chat interface.**

Code-Graph-RAG is an open-source tool that transforms software repositories into queryable knowledge graphs. Built with **Typer**, the `cgr` executable exposes specialized CLI commands for parsing and querying codebases, enabling developers to analyze complex projects using natural language. This guide examines the implementation details in `vitali87/code-graph-rag` to explain how these commands function.

## Parsing Codebases with `cgr index`

The `cgr index` command converts a local repository into a serialized graph representation. Defined at line 645 in **[`cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/cli.py)**, this command orchestrates the ingestion pipeline without requiring a live database connection.

### How Repository Parsing Works

The parsing workflow executes five distinct phases:

1. **Repository validation** – `_resolve_and_validate_repo()` verifies the supplied path exists, is a directory, and contains a `.git` folder (warning if absent).
2. **Ignore pattern handling** – `load_ignore_patterns()` reads `.cgrignore` files; the CLI can interactively prompt for un-ignoring specific paths.
3. **Ingestor initialization** – `ProtobufFileIngestor(output_path=..., split_index=...)` prepares the output destination for serialized graph data.
4. **Graph construction** – `GraphUpdater.run()` traverses the repository tree, executes language-specific parsers, and stores symbols and relationships.
5. **Protobuf export** – The system writes a **`.proto`** database file that can be version-controlled or loaded later via `cgr graph-loader`.

The **`GraphUpdater`** class in **[`graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_updater.py)** serves as the core engine, receiving the ingestor, repository path, and parser configurations to populate the graph.

```bash

# Parse a repository into a protobuf graph

cgr index \
  --repo-path /path/to/my/project \
  --output-proto-dir ./graph-output \
  --split-index

```

## Querying Codebases with `cgr start`

The `cgr start` command enables LLM-driven interrogation of codebases, implemented at line 495 in **[`cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/cli.py)**. This command synchronizes graph data to Memgraph and launches either an interactive chat session or single-shot queries.

### Interactive Chat and Single-Shot Queries

The querying workflow follows this sequence:

1. **Repository resolution** – Derives the project name from the repo path or uses the `--project-name` override.
2. **Database preparation** – `--clean` empties the Memgraph database; `--update-graph` forces a fresh sync before accepting queries.
3. **Graph synchronization** – `_run_graph_sync()` instantiates `GraphUpdater` with a live `MemgraphIngestor` from `connect_memgraph()` in **[`config.py`](https://github.com/vitali87/code-graph-rag/blob/main/config.py)**, then calls `updater.run()` to populate Memgraph with nodes representing functions, calls, and imports.
4. **Query execution** – `main_async()` launches an asynchronous loop that streams LLM responses while retrieving relevant graph fragments via Cypher queries.
5. **Agent invocation** – With `--ask-agent`, the system executes `main_single_query()` for one-time questions without entering the interactive chat loop.

The **[`vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/vector_store.py)** module manages embedding storage for semantic search during the querying phase.

```bash

# Start interactive chat with fresh graph sync

cgr start \
  --repo-path /path/to/my/project \
  --update-graph \
  --orchestrator gpt-4o-mini \
  --cypher gpt-4o-mini

# Single-shot query without interactive mode

cgr start \
  --repo-path /path/to/my/project \
  --ask-agent "Which functions modify global state?" \
  --output-format json

```

## Shared Infrastructure and Configuration

Both CLI commands rely on common utilities defined across the codebase:

- **`load_parsers()`** in **[`parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/parser_loader.py)** – Dynamically discovers language-specific parsers (Python, Java, Go, etc.) and registers Cypher queries that enable LLM reasoning.
- **`connect_memgraph()`** in **[`config.py`](https://github.com/vitali87/code-graph-rag/blob/main/config.py)** – Provides a context manager yielding a `MemgraphIngestor` with configurable batch sizes for efficient database writes.
- **[`graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_loader.py)** – Contains utilities for loading protobuf graph files and printing summaries via the `cgr graph-loader` helper command.

These shared components ensure consistent behavior whether you are exporting static graphs or querying live databases.

## Practical CLI Workflow Examples

Complete workflows combining parsing and querying operations:

```bash

# 1. Parse and export a repository

cgr index \
  --repo-path ./my-project \
  --output-proto-dir ./graphs \
  --split-index

# 2. Inspect the exported graph (optional)

cgr graph-loader ./graphs/graph.proto

# 3. Query with automatic graph updates

cgr start \
  --repo-path ./my-project \
  --update-graph \
  --orchestrator gpt-4o-mini

```

Common flags available to both commands include `--exclude`, `--capture`, `--interactive-setup`, `--batch-size`, and `--orchestrator`, providing fine-grained control over ingestion scope and LLM model selection.

## Summary

- **`cgr index`** parses repositories into portable protobuf graph files using `ProtobufFileIngestor` and `GraphUpdater` as defined in **[`cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/cli.py)**.
- **`cgr start`** synchronizes graphs to Memgraph and launches interactive LLM chat or single-shot queries via `main_async()` or `main_single_query()`.
- Both commands utilize **`load_parsers()`** for multi-language support and **`connect_memgraph()`** for database connectivity.
- The CLI supports granular control through shared flags for exclusion patterns, batch sizes, and orchestrator selection.

## Frequently Asked Questions

### What is the difference between `cgr index` and `cgr start`?

`cgr index` performs offline parsing, writing repository structure to protobuf files without requiring a database connection. `cgr start` requires Memgraph and provides interactive querying capabilities, optionally syncing the graph before launching the LLM interface.

### How does Code-Graph-RAG handle large repositories?

The CLI provides `--split-index` for `cgr index` to partition large graphs into multiple protobuf files. Additionally, `--batch-size` controls the number of nodes processed per transaction when syncing to Memgraph, preventing memory exhaustion during ingestion.

### Can I query a codebase without parsing it first?

No, the graph must be populated before querying. However, `cgr start` with `--update-graph` performs parsing and syncing automatically before launching the chat interface, combining both steps into a single command execution.

### Which programming languages are supported by the parser?

The system dynamically discovers language-specific parsers through `load_parsers()` in **[`parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/parser_loader.py)**, with built-in support for Python, Java, Go, and other languages. Each parser registers specific Cypher queries that enable the LLM to reason about language-specific constructs like imports, class hierarchies, and function calls.