Prerequisites for Installing Code-Graph-RAG: Complete System Requirements Guide
Before installing Code-Graph-RAG, you must have Python 3.12 or newer, Docker with Docker Compose, cmake, ripgrep, access to a Google Gemini or OpenAI API key (or a local Ollama instance), and a Python package manager such as uv.
The vitali87/code-graph-rag repository provides a graph-based RAG system for codebases that requires specific system-level dependencies to compile parsers, run containerized databases, and connect to AI backends. Meeting these prerequisites for installing Code-Graph-RAG ensures the Tree-sitter parsers, Memgraph graph database, and vector stores function correctly before you execute the installation commands. The following guide details every requirement based on the project's pyproject.toml and installation documentation.
Essential Runtime and System Dependencies
Python 3.12 or Newer
The package ships as a pure-Python wheel that strictly requires Python 3.12 or newer. According to README.md (lines 95-99), attempting to install on earlier versions such as Debian Bookworm's Python 3.11 will fail. Verify your version:
python --version
Docker and Docker Compose
You need Docker and Docker Compose to run the bundled Memgraph graph database and Qdrant vector store. The cgr daemon up command launches these services via Docker Compose as specified in the quick-start guide (README.md lines 49-52). Verify installation:
docker version
docker compose version
Build Tools and Search Utilities
Two command-line tools are mandatory:
cmake: Required to build the nativepymgclientdependency that connects to Memgraph, as documented indocs/getting-started/installation.md(lines 11-12) anddocs/advanced/building-binaries.md.ripgrep(rg): Used by the CLI for fast source-code searching and shell-command integrations (referenced indocs/getting-started/installation.mdlines 12-13).
Check both:
cmake --version
rg --version
AI Model Provider Access
The RAG engine requires a language model to generate Cypher queries and code patches. You must provide either:
- Cloud models: A Google Gemini or OpenAI API key set via environment variables (see
.env.example). - Local models: An Ollama instance running locally for self-hosted models.
This requirement is documented in the installation prerequisites (docs/getting-started/installation.md lines 13-14).
Python Package Management
While pip or pipx work, the documentation recommends uv for handling isolated environments and supporting extra-dependency syntax like [treesitter-full,semantic] (README.md lines 99-105). Install uv if you haven't already, then proceed with the tool installation:
uv tool install "code-graph-rag[treesitter-full,semantic]"
Or use pipx as an alternative:
pipx install "code-graph-rag[treesitter-full,semantic]"
Optional Dependencies for Full Language Support
Tree-sitter Language Grammars
For parsing multiple programming languages beyond basic subsets, install the treesitter-full extra. This provides parsers for Python, TypeScript, Rust, Go, Java, and C/C++ that codebase_rag/graph_loader.py uses to transform ASTs into Memgraph nodes. Without these grammars, the indexer supports only a limited language set (docs/getting-started/installation.md lines 64-80).
Semantic Search Embeddings
To enable vector-based code search powered by the UniXcoder model, include the semantic extra when installing. This functionality requires the embeddings package referenced in pyproject.toml (docs/getting-started/installation.md lines 70-78).
Installation Verification and Startup
Once prerequisites are met and the package is installed, verify your environment with the built-in health check:
cgr doctor
This command validates your Python version, Docker connectivity, cmake availability, and ripgrep installation.
Start the required services:
cgr daemon up
Then parse a repository and query it:
cgr start --repo-path /path/to/your/project --update-graph
cgr query "What functions call `foo`?"
Summary
- Python 3.12+ is strictly required; earlier versions will not install the package according to
README.md. - Docker and Docker Compose are mandatory for running the Memgraph and Qdrant containers via
cgr daemon up. - cmake and ripgrep must be present at the system level to build native dependencies and enable fast code searching.
- AI access requires either cloud API keys (Gemini/OpenAI) or a local Ollama installation defined in
.env.example. - Optional extras (
treesitter-full,semantic) unlock full multi-language parsing and semantic vector search capabilities defined inpyproject.toml.
Frequently Asked Questions
Can I install Code-Graph-RAG with Python 3.11?
No. The package requires Python 3.12 or newer because it ships a pure-Python wheel incompatible with earlier versions. According to the README.md, systems like Debian Bookworm that default to Python 3.11 will fail during installation.
Is Docker mandatory for Code-Graph-RAG?
Yes. Docker and Docker Compose are required to run Memgraph (the graph database) and Qdrant (the vector store). The cgr daemon up command orchestrates these services via Docker Compose, and the system cannot function without them.
What is the difference between the treesitter-full and semantic extras?
The treesitter-full extra installs all Tree-sitter grammars needed by codebase_rag/graph_loader.py to parse multiple programming languages into ASTs. The semantic extra provides UniXcoder embeddings for vector-based code search. You can install both together using code-graph-rag[treesitter-full,semantic].
How do I verify all prerequisites are met before installing?
Run the cgr doctor command after installing the package. This health check validates your Python version, Docker daemon, cmake installation, ripgrep availability, and AI model connectivity before you attempt to parse repositories.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →