How to Set Up Code-Graph-RAG with Memgraph: Complete Configuration Guide

Code-Graph-RAG persists code graphs in Memgraph via the MemgraphIngestor class, configured through environment variables and initialized with the connect_memgraph helper function.

The Code-Graph-RAG (cgr) CLI tool extracts code structure into a graph format and stores it in Memgraph, a high-performance graph database. This setup enables fast Cypher queries for code analysis, duplicate detection, and dependency exploration. According to the vitali87/code-graph-rag source code, Memgraph integration is handled through a dedicated ingestor class with batched Bolt protocol writes.

Prerequisites and Memgraph Installation

Before configuring Code-Graph-RAG, you need a running Memgraph instance. The default configuration expects Memgraph on localhost:7687 using the Bolt protocol.

Start Memgraph with Docker

The quickest method is the official Memgraph Docker image, which exposes ports 7687 (Bolt) and 7444 (web interface):

docker run -d --name memgraph \
  -p 7687:7687 -p 7444:7444 \
  memgraph/memgraph:latest

For persistent storage, add a volume mount:

docker run -d --name memgraph \
  -p 7687:7687 -p 7444:7444 \
  -v mg_data:/var/lib/memgraph \
  memgraph/memgraph:latest

Verify the instance is healthy:

docker exec -it memgraph mgconsole

Configure Memgraph Connection Parameters

Code-Graph-RAG loads connection settings from codebase_rag/config.py via the AppConfig class. The relevant defaults are defined at lines 162-168:

Environment Variable Default Purpose
MEMGRAPH_HOST localhost Bolt server hostname
MEMGRAPH_PORT 7687 Bolt server port
MEMGRAPH_USERNAME None Authentication username
MEMGRAPH_PASSWORD None Authentication password
MEMGRAPH_BATCH_SIZE 1000 Graph mutation batch size

Create a .env Configuration File

Place this in your repository root for automatic loading:

cat > .env <<'EOF'
MEMGRAPH_HOST=localhost
MEMGRAPH_PORT=7687

# Uncomment if Memgraph authentication is enabled

# MEMGRAPH_USERNAME=admin

# MEMGRAPH_PASSWORD=secret

# Optional: tune batch size for large codebases

MEMGRAPH_BATCH_SIZE=2000
EOF

Code-Graph-RAG uses python-dotenv to load these values into AppConfig at runtime.

CLI Usage: Index and Query Code Graphs

Once configured, use the --graph-backend memgraph flag with cgr commands. The CLI internally calls connect_memgraph from codebase_rag/main.py (lines 1541-1548) to obtain a context-managed MemgraphIngestor.

Index a Python Project

cgr index --path /path/to/your/repo \
  --graph-backend memgraph \
  --batch-size 2000

The --batch-size parameter overrides MEMGRAPH_BATCH_SIZE. The ingestor accumulates nodes and edges, flushing to Memgraph when the batch threshold is reached or at context exit.

Query the Stored Graph

cgr query "MATCH (f:Function) RETURN f.name LIMIT 10"

Under the hood, this opens a MemgraphIngestor context and executes Cypher through the Bolt connection.

Find Duplicate Code Patterns

cgr duplicates --graph-backend memgraph --threshold 0.85

Programmatic Memgraph Integration

For custom workflows, import the connection helper directly:

from pathlib import Path
from codebase_rag.main import connect_memgraph
from codebase_rag.graph_loader import load_graph_from_repo

# Batch size can be passed directly or falls back to config

with connect_memgraph(batch_size=500) as ingestor:
    load_graph_from_repo(Path("/path/to/repo"), ingestor)
    # Automatic flush on context exit

The connect_memgraph function yields a MemgraphIngestor that implements:

  • Batched mutations – accumulates CREATE and MERGE statements
  • Automatic flush – commits remaining operations on context exit
  • Connection pooling – reuses Bolt sessions efficiently

Health Checking and Troubleshooting

The codebase_rag/tools/health_checker.py module provides connectivity verification. Commands that require Memgraph run this check before executing.

Common Issues

Symptom Cause Solution
Connection refused Memgraph not running Start Docker container, verify port 7687
Authentication failed Credentials mismatch Check MEMGRAPH_USERNAME/MEMGRAPH_PASSWORD
Slow ingestion Small batch size Increase MEMGRAPH_BATCH_SIZE or use --batch-size
Host not found Docker network issues Use host.docker.internal for cross-container access

For multi-project setups, codebase_rag/stack/manager.py (lines 73-78) resolves Memgraph host and port through the stack manager, enabling centralized configuration across repositories.

Key Source Files Reference

Understanding these locations aids debugging and extension:

Summary

  • Memgraph setup requires a running instance on port 7687, typically via Docker
  • Configuration uses environment variables loaded by AppConfig in config.py
  • Connection is established through connect_memgraph, yielding a batched MemgraphIngestor
  • CLI commands (index, query, duplicates) use --graph-backend memgraph to enable persistence
  • Batch size tunes performance: default 1000, adjustable via MEMGRAPH_BATCH_SIZE or CLI flag
  • Health checks verify connectivity before operations to fail fast on misconfiguration

Frequently Asked Questions

What port does Code-Graph-RAG use to connect to Memgraph?

Code-Graph-RAG connects to Memgraph via the Bolt protocol on port 7687 by default. This is defined in codebase_rag/config.py as the MEMGRAPH_PORT configuration value. If running Memgraph on a non-standard port, set the environment variable or pass a custom configuration.

Can I use authenticated Memgraph with Code-Graph-RAG?

Yes. Set MEMGRAPH_USERNAME and MEMGRAPH_PASSWORD in your .env file or environment. These values are read by AppConfig in config.py and passed to the Bolt driver when connect_memgraph establishes the connection. Leave both unset for unauthenticated local development.

How do I tune ingestion performance for large codebases?

Increase the batch size above the default 1000. Use --batch-size 5000 with CLI commands or set MEMGRAPH_BATCH_SIZE=5000 in your environment. Larger batches reduce network round-trips but consume more memory. The MemgraphIngestor flushes automatically at batch threshold and on context exit.

Is Docker required to run Memgraph with Code-Graph-RAG?

No. Docker is the recommended method, but any Memgraph instance accessible via Bolt works. Install Memgraph natively, configure MEMGRAPH_HOST and MEMGRAPH_PORT to match your deployment, and ensure the Code-Graph-RAG host can reach the Memgraph server on that address.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →