# How to Set Up Code-Graph-RAG with Memgraph: Complete Configuration Guide

> Master Code-Graph-RAG setup with Memgraph. Follow our guide to configure persistence using MemgraphIngestor and connect_memgraph for efficient code graph management. Get started today!

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Code-Graph-RAG persists code graphs in Memgraph via the `MemgraphIngestor` class, configured through environment variables and initialized with the `connect_memgraph` helper function.**

The Code-Graph-RAG (`cgr`) CLI tool extracts code structure into a graph format and stores it in **Memgraph**, a high-performance graph database. This setup enables fast Cypher queries for code analysis, duplicate detection, and dependency exploration. According to the `vitali87/code-graph-rag` source code, Memgraph integration is handled through a dedicated ingestor class with batched Bolt protocol writes.

## Prerequisites and Memgraph Installation

Before configuring Code-Graph-RAG, you need a running Memgraph instance. The default configuration expects Memgraph on **localhost:7687** using the Bolt protocol.

### Start Memgraph with Docker

The quickest method is the official Memgraph Docker image, which exposes ports **7687** (Bolt) and **7444** (web interface):

```bash
docker run -d --name memgraph \
  -p 7687:7687 -p 7444:7444 \
  memgraph/memgraph:latest

```

For persistent storage, add a volume mount:

```bash
docker run -d --name memgraph \
  -p 7687:7687 -p 7444:7444 \
  -v mg_data:/var/lib/memgraph \
  memgraph/memgraph:latest

```

Verify the instance is healthy:

```bash
docker exec -it memgraph mgconsole

```

## Configure Memgraph Connection Parameters

Code-Graph-RAG loads connection settings from **[`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py)** via the `AppConfig` class. The relevant defaults are defined at **lines 162-168**:

| Environment Variable | Default | Purpose |
|---------------------|---------|---------|
| `MEMGRAPH_HOST` | `localhost` | Bolt server hostname |
| `MEMGRAPH_PORT` | `7687` | Bolt server port |
| `MEMGRAPH_USERNAME` | `None` | Authentication username |
| `MEMGRAPH_PASSWORD` | `None` | Authentication password |
| `MEMGRAPH_BATCH_SIZE` | `1000` | Graph mutation batch size |

### Create a .env Configuration File

Place this in your repository root for automatic loading:

```bash
cat > .env <<'EOF'
MEMGRAPH_HOST=localhost
MEMGRAPH_PORT=7687

# Uncomment if Memgraph authentication is enabled

# MEMGRAPH_USERNAME=admin

# MEMGRAPH_PASSWORD=secret

# Optional: tune batch size for large codebases

MEMGRAPH_BATCH_SIZE=2000
EOF

```

Code-Graph-RAG uses `python-dotenv` to load these values into `AppConfig` at runtime.

## CLI Usage: Index and Query Code Graphs

Once configured, use the `--graph-backend memgraph` flag with `cgr` commands. The CLI internally calls `connect_memgraph` from **[`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py)** (lines 1541-1548) to obtain a context-managed `MemgraphIngestor`.

### Index a Python Project

```bash
cgr index --path /path/to/your/repo \
  --graph-backend memgraph \
  --batch-size 2000

```

The `--batch-size` parameter overrides `MEMGRAPH_BATCH_SIZE`. The ingestor accumulates nodes and edges, flushing to Memgraph when the batch threshold is reached or at context exit.

### Query the Stored Graph

```bash
cgr query "MATCH (f:Function) RETURN f.name LIMIT 10"

```

Under the hood, this opens a `MemgraphIngestor` context and executes Cypher through the Bolt connection.

### Find Duplicate Code Patterns

```bash
cgr duplicates --graph-backend memgraph --threshold 0.85

```

## Programmatic Memgraph Integration

For custom workflows, import the connection helper directly:

```python
from pathlib import Path
from codebase_rag.main import connect_memgraph
from codebase_rag.graph_loader import load_graph_from_repo

# Batch size can be passed directly or falls back to config

with connect_memgraph(batch_size=500) as ingestor:
    load_graph_from_repo(Path("/path/to/repo"), ingestor)
    # Automatic flush on context exit

```

The `connect_memgraph` function yields a `MemgraphIngestor` that implements:

- **Batched mutations** – accumulates `CREATE` and `MERGE` statements
- **Automatic flush** – commits remaining operations on context exit
- **Connection pooling** – reuses Bolt sessions efficiently

## Health Checking and Troubleshooting

The **[`codebase_rag/tools/health_checker.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/health_checker.py)** module provides connectivity verification. Commands that require Memgraph run this check before executing.

### Common Issues

| Symptom | Cause | Solution |
|---------|-------|----------|
| `Connection refused` | Memgraph not running | Start Docker container, verify port 7687 |
| `Authentication failed` | Credentials mismatch | Check `MEMGRAPH_USERNAME`/`MEMGRAPH_PASSWORD` |
| `Slow ingestion` | Small batch size | Increase `MEMGRAPH_BATCH_SIZE` or use `--batch-size` |
| `Host not found` | Docker network issues | Use `host.docker.internal` for cross-container access |

For multi-project setups, **[`codebase_rag/stack/manager.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/stack/manager.py)** (lines 73-78) resolves Memgraph host and port through the stack manager, enabling centralized configuration across repositories.

## Key Source Files Reference

Understanding these locations aids debugging and extension:

- **[`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py)** – Default connection parameters and `AppConfig` class
- **[`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py)** – `connect_memgraph` helper (lines 1541-1548)
- **[`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py)** – CLI entry point forwarding to ingestor
- **[`codebase_rag/tools/health_checker.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/health_checker.py)** – Connectivity verification utilities
- **[`codebase_rag/stack/manager.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/stack/manager.py)** – Multi-project stack resolution (lines 73-78)

## Summary

- **Memgraph setup** requires a running instance on port 7687, typically via Docker
- **Configuration** uses environment variables loaded by `AppConfig` in [`config.py`](https://github.com/vitali87/code-graph-rag/blob/main/config.py)
- **Connection** is established through `connect_memgraph`, yielding a batched `MemgraphIngestor`
- **CLI commands** (`index`, `query`, `duplicates`) use `--graph-backend memgraph` to enable persistence
- **Batch size** tunes performance: default 1000, adjustable via `MEMGRAPH_BATCH_SIZE` or CLI flag
- **Health checks** verify connectivity before operations to fail fast on misconfiguration

## Frequently Asked Questions

### What port does Code-Graph-RAG use to connect to Memgraph?

Code-Graph-RAG connects to Memgraph via the **Bolt protocol on port 7687** by default. This is defined in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py) as the `MEMGRAPH_PORT` configuration value. If running Memgraph on a non-standard port, set the environment variable or pass a custom configuration.

### Can I use authenticated Memgraph with Code-Graph-RAG?

Yes. Set `MEMGRAPH_USERNAME` and `MEMGRAPH_PASSWORD` in your `.env` file or environment. These values are read by `AppConfig` in [`config.py`](https://github.com/vitali87/code-graph-rag/blob/main/config.py) and passed to the Bolt driver when `connect_memgraph` establishes the connection. Leave both unset for unauthenticated local development.

### How do I tune ingestion performance for large codebases?

Increase the **batch size** above the default 1000. Use `--batch-size 5000` with CLI commands or set `MEMGRAPH_BATCH_SIZE=5000` in your environment. Larger batches reduce network round-trips but consume more memory. The `MemgraphIngestor` flushes automatically at batch threshold and on context exit.

### Is Docker required to run Memgraph with Code-Graph-RAG?

No. Docker is the recommended method, but any Memgraph instance accessible via Bolt works. Install Memgraph natively, configure `MEMGRAPH_HOST` and `MEMGRAPH_PORT` to match your deployment, and ensure the Code-Graph-RAG host can reach the Memgraph server on that address.