How to Set Up Code-Graph-RAG with Memgraph: Complete Configuration Guide
Code-Graph-RAG persists code graphs in Memgraph via the MemgraphIngestor class, configured through environment variables and initialized with the connect_memgraph helper function.
The Code-Graph-RAG (cgr) CLI tool extracts code structure into a graph format and stores it in Memgraph, a high-performance graph database. This setup enables fast Cypher queries for code analysis, duplicate detection, and dependency exploration. According to the vitali87/code-graph-rag source code, Memgraph integration is handled through a dedicated ingestor class with batched Bolt protocol writes.
Prerequisites and Memgraph Installation
Before configuring Code-Graph-RAG, you need a running Memgraph instance. The default configuration expects Memgraph on localhost:7687 using the Bolt protocol.
Start Memgraph with Docker
The quickest method is the official Memgraph Docker image, which exposes ports 7687 (Bolt) and 7444 (web interface):
docker run -d --name memgraph \
-p 7687:7687 -p 7444:7444 \
memgraph/memgraph:latest
For persistent storage, add a volume mount:
docker run -d --name memgraph \
-p 7687:7687 -p 7444:7444 \
-v mg_data:/var/lib/memgraph \
memgraph/memgraph:latest
Verify the instance is healthy:
docker exec -it memgraph mgconsole
Configure Memgraph Connection Parameters
Code-Graph-RAG loads connection settings from codebase_rag/config.py via the AppConfig class. The relevant defaults are defined at lines 162-168:
| Environment Variable | Default | Purpose |
|---|---|---|
MEMGRAPH_HOST |
localhost |
Bolt server hostname |
MEMGRAPH_PORT |
7687 |
Bolt server port |
MEMGRAPH_USERNAME |
None |
Authentication username |
MEMGRAPH_PASSWORD |
None |
Authentication password |
MEMGRAPH_BATCH_SIZE |
1000 |
Graph mutation batch size |
Create a .env Configuration File
Place this in your repository root for automatic loading:
cat > .env <<'EOF'
MEMGRAPH_HOST=localhost
MEMGRAPH_PORT=7687
# Uncomment if Memgraph authentication is enabled
# MEMGRAPH_USERNAME=admin
# MEMGRAPH_PASSWORD=secret
# Optional: tune batch size for large codebases
MEMGRAPH_BATCH_SIZE=2000
EOF
Code-Graph-RAG uses python-dotenv to load these values into AppConfig at runtime.
CLI Usage: Index and Query Code Graphs
Once configured, use the --graph-backend memgraph flag with cgr commands. The CLI internally calls connect_memgraph from codebase_rag/main.py (lines 1541-1548) to obtain a context-managed MemgraphIngestor.
Index a Python Project
cgr index --path /path/to/your/repo \
--graph-backend memgraph \
--batch-size 2000
The --batch-size parameter overrides MEMGRAPH_BATCH_SIZE. The ingestor accumulates nodes and edges, flushing to Memgraph when the batch threshold is reached or at context exit.
Query the Stored Graph
cgr query "MATCH (f:Function) RETURN f.name LIMIT 10"
Under the hood, this opens a MemgraphIngestor context and executes Cypher through the Bolt connection.
Find Duplicate Code Patterns
cgr duplicates --graph-backend memgraph --threshold 0.85
Programmatic Memgraph Integration
For custom workflows, import the connection helper directly:
from pathlib import Path
from codebase_rag.main import connect_memgraph
from codebase_rag.graph_loader import load_graph_from_repo
# Batch size can be passed directly or falls back to config
with connect_memgraph(batch_size=500) as ingestor:
load_graph_from_repo(Path("/path/to/repo"), ingestor)
# Automatic flush on context exit
The connect_memgraph function yields a MemgraphIngestor that implements:
- Batched mutations – accumulates
CREATEandMERGEstatements - Automatic flush – commits remaining operations on context exit
- Connection pooling – reuses Bolt sessions efficiently
Health Checking and Troubleshooting
The codebase_rag/tools/health_checker.py module provides connectivity verification. Commands that require Memgraph run this check before executing.
Common Issues
| Symptom | Cause | Solution |
|---|---|---|
Connection refused |
Memgraph not running | Start Docker container, verify port 7687 |
Authentication failed |
Credentials mismatch | Check MEMGRAPH_USERNAME/MEMGRAPH_PASSWORD |
Slow ingestion |
Small batch size | Increase MEMGRAPH_BATCH_SIZE or use --batch-size |
Host not found |
Docker network issues | Use host.docker.internal for cross-container access |
For multi-project setups, codebase_rag/stack/manager.py (lines 73-78) resolves Memgraph host and port through the stack manager, enabling centralized configuration across repositories.
Key Source Files Reference
Understanding these locations aids debugging and extension:
codebase_rag/config.py– Default connection parameters andAppConfigclasscodebase_rag/main.py–connect_memgraphhelper (lines 1541-1548)codebase_rag/cli.py– CLI entry point forwarding to ingestorcodebase_rag/tools/health_checker.py– Connectivity verification utilitiescodebase_rag/stack/manager.py– Multi-project stack resolution (lines 73-78)
Summary
- Memgraph setup requires a running instance on port 7687, typically via Docker
- Configuration uses environment variables loaded by
AppConfiginconfig.py - Connection is established through
connect_memgraph, yielding a batchedMemgraphIngestor - CLI commands (
index,query,duplicates) use--graph-backend memgraphto enable persistence - Batch size tunes performance: default 1000, adjustable via
MEMGRAPH_BATCH_SIZEor CLI flag - Health checks verify connectivity before operations to fail fast on misconfiguration
Frequently Asked Questions
What port does Code-Graph-RAG use to connect to Memgraph?
Code-Graph-RAG connects to Memgraph via the Bolt protocol on port 7687 by default. This is defined in codebase_rag/config.py as the MEMGRAPH_PORT configuration value. If running Memgraph on a non-standard port, set the environment variable or pass a custom configuration.
Can I use authenticated Memgraph with Code-Graph-RAG?
Yes. Set MEMGRAPH_USERNAME and MEMGRAPH_PASSWORD in your .env file or environment. These values are read by AppConfig in config.py and passed to the Bolt driver when connect_memgraph establishes the connection. Leave both unset for unauthenticated local development.
How do I tune ingestion performance for large codebases?
Increase the batch size above the default 1000. Use --batch-size 5000 with CLI commands or set MEMGRAPH_BATCH_SIZE=5000 in your environment. Larger batches reduce network round-trips but consume more memory. The MemgraphIngestor flushes automatically at batch threshold and on context exit.
Is Docker required to run Memgraph with Code-Graph-RAG?
No. Docker is the recommended method, but any Memgraph instance accessible via Bolt works. Install Memgraph natively, configure MEMGRAPH_HOST and MEMGRAPH_PORT to match your deployment, and ensure the Code-Graph-RAG host can reach the Memgraph server on that address.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →