How to Set Up Code-Graph-RAG for a New Project: Complete Installation Guide
Set up Code-Graph-RAG by installing the Python CLI with Tree-sitter support, launching the Memgraph stack via cgr daemon up, and indexing your repository with cgr start --repo-path /path/to/project --update-graph.
Code-Graph-RAG transforms multi-language codebases into interactive knowledge graphs that you can query and edit using plain English. This guide walks you through installing the vitali87/code-graph-rag toolkit, configuring the Memgraph backend, and indexing your first project for AI-assisted code navigation.
Understanding the Code-Graph-RAG Architecture
Before installation, it helps to understand how the system processes your code. Code-Graph-RAG consists of two integrated components working together to provide semantic code search and editing capabilities.
Multi-Language Parser
The Multi-Language Parser uses Tree-sitter to analyze source files across Python, Rust, Go, and other languages. It extracts functions, classes, methods, and their relationships, then writes them into Memgraph using a language-agnostic schema. The core logic resides in codebase_rag/parser_loader.py, with language-specific AST definitions located in codebase_rag/constants/ast_*.py files.
RAG Query System
The RAG System provides the interactive CLI (cgr command) that converts natural-language queries into Cypher queries. It fetches relevant nodes from Memgraph and optionally invokes LLMs to edit or optimize code. Key implementations include codebase_rag/cli.py for the command interface, codebase_rag/graph_query.py for database interactions, and codebase_rag/prompts.py for LLM prompt templates.
Prerequisites and Installation
Setting up Code-Graph-RAG requires specific system dependencies and Python tooling.
System Requirements
Ensure your environment includes:
- Python 3.12 or higher for runtime execution
- Docker and Docker Compose to run the Memgraph database and Qdrant vector store
- cmake for building the native
pymgclientdriver - ripgrep (
rg) for fast file search during indexing
These tools support the graph database backend, native client compilation, and efficient repository scanning.
Install Code-Graph-RAG from PyPI
Install the latest stable release using uv or pipx with full language support:
# Using uv (recommended)
uv tool install "code-graph-rag[treesitter-full,semantic]"
# Using pipx
pipx install "code-graph-rag[treesitter-full,semantic]"
This command installs the cgr CLI along with all Tree-sitter grammars and UniXcoder embeddings required for semantic code analysis.
Alternative: Install from Source
To access features ahead of the PyPI release, install directly from Git:
uv tool install "code-graph-rag[treesitter-full,semantic] @ git+https://github.com/vitali87/code-graph-rag@main"
This approach pulls the most recent commits from the main branch, including experimental features not yet packaged for distribution.
Configuring Your Environment
After installation, configure the database stack and authentication credentials.
Start the Memgraph Database Stack
Launch the bundled Memgraph and Qdrant containers:
cgr daemon up
This command spins up the required infrastructure on port 7687, enabling the CLI to communicate with the graph database.
Set Up Environment Variables
Create and configure your environment file:
cp .env.example .env
Edit .env to include your LLM provider credentials (OpenAI, Gemini, etc.) and optional Memgraph authentication details. The application reads these variables at runtime to authenticate with external AI services.
Verify Installation with cgr doctor
Run the diagnostic tool to validate your setup:
cgr doctor
This command performs health checks for Docker connectivity, Memgraph availability, required binaries (cmake, rg), and LLM API accessibility. Resolve any reported issues before proceeding.
Indexing Your Codebase
With the infrastructure running, parse your target repository into the knowledge graph.
Initial Repository Parsing
Index your project using the start command:
cgr start --repo-path /path/to/your/project --update-graph
This invokes codebase_rag/parser_loader.py to traverse your source tree, build abstract syntax trees, and ingest nodes into Memgraph via codebase_rag/graph_loader.py. The process maps functions, classes, modules, and their dependencies into a queryable graph structure.
Real-Time Graph Updates
For active development, enable filesystem watching to synchronize changes:
cgr realtime-updater --repo-path /path/to/project
Implemented in codebase_rag/realtime_updater.py, this service monitors your repository for modifications and incrementally updates the graph without full re-indexing.
Querying and Editing Code
Once indexed, interact with your codebase through natural language commands.
Natural Language Queries
Ask questions about your code structure:
cgr ask "How many functions does module X contain?"
cgr ask "Show me all classes that inherit from BaseController"
The CLI translates these queries into Cypher statements using templates from codebase_rag/prompts.py, executes them against Memgraph, and returns structured results.
AI-Assisted Code Editing
Modify code through natural language instructions:
cgr edit --function my_module.my_func --instruction "Rename parameter x to count"
This command locates the function in the graph, retrieves its context and dependencies from Memgraph, and prompts your configured LLM to generate the edit. The system handles multi-file refactoring by understanding cross-references stored in the knowledge graph.
Summary
- Install Code-Graph-RAG via PyPI with
uv tool install "code-graph-rag[treesitter-full,semantic]"to get thecgrCLI and Tree-sitter parsers. - Start infrastructure using
cgr daemon upto launch Memgraph and Qdrant containers on port 7687. - Configure API keys in
.envand verify setup withcgr doctorto ensure all dependencies are functional. - Index repositories with
cgr start --repo-path /path/to/project --update-graphto parse code into the knowledge graph viacodebase_rag/parser_loader.py. - Query naturally using
cgr askand edit code withcgr edit, leveraging the RAG system incodebase_rag/cli.pyandcodebase_rag/graph_query.py.
Frequently Asked Questions
What programming languages does Code-Graph-RAG support?
Code-Graph-RAG supports multiple languages through Tree-sitter grammars included in the treesitter-full extra. The parser definitions in codebase_rag/constants/ast_*.py handle Python, Rust, Go, and other popular languages, extracting functions, classes, and their relationships into a language-agnostic graph schema stored in Memgraph.
How do I update the knowledge graph when my code changes?
Run cgr start --repo-path /path/to/project --update-graph to perform a full re-index, or use cgr realtime-updater --repo-path /path/to/project for continuous synchronization. The realtime updater, implemented in codebase_rag/realtime_updater.py, monitors filesystem events and incrementally updates the graph without re-parsing the entire repository.
Can I export the knowledge graph for external analysis?
Yes. Use the example script at examples/graph_export_example.py to export your Memgraph data to JSON or CSV formats. This allows you to analyze code relationships in external tools or create backups of your graph structure outside the database.
Why does the installation require cmake and ripgrep?
The cmake dependency builds the native pymgclient driver required for high-performance communication with Memgraph. The ripgrep (rg) binary enables fast file search during repository indexing, significantly speeding up the initial parsing phase in codebase_rag/parser_loader.py when scanning large codebases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →