# Default Memgraph Settings for code-graph-rag: Complete Configuration Reference

> Discover the default Memgraph settings for code-graph-rag. Explore connection details, batch sizes, and query memory limits in this complete configuration reference.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: api-reference
- Published: 2026-09-05

---

**The `code-graph-rag` project configures Memgraph defaults in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py), using localhost:7687 for connections, batch sizes of 1000 nodes/relationships, and a 4096 MB query memory limit.**

The open-source `vitali87/code-graph-rag` repository leverages Memgraph as its primary graph database for storing code relationships. All default Memgraph settings for code-graph-rag are centralized in the `AppConfig` Pydantic settings class located in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py), which automatically loads environment variables from `.env` files to override hardcoded defaults.

## Connection Parameters

### Host, Port, and Protocol

The foundation of the Memgraph connection relies on three key settings defined in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py):

- **MEMGRAPH_HOST**: Defaults to `"localhost"` (line 62). This specifies the hostname where the Memgraph instance is reachable.
- **MEMGRAPH_PORT**: Defaults to `7687` (line 63). This is the Bolt-protocol port used by the Python `mgclient` driver to establish binary connections.
- **MEMGRAPH_HTTP_PORT**: Defaults to `7444` (line 64). This port serves Memgraph Lab and optional REST endpoints, though the primary ingestion uses the Bolt protocol.

### Authentication Defaults

By default, `code-graph-rag` assumes an unsecured local Memgraph instance:

- **MEMGRAPH_USERNAME**: Defaults to `None` (line 65).
- **MEMGRAPH_PASSWORD**: Defaults to `None` (line 66).

When both values remain `None`, the `MemgraphIngestor` class in [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py) connects without authentication. For production deployments, these can be overridden via environment variables to enable Memgraph's role-based access control.

## Performance and Batching Configuration

### Ingestion Batch Size

**MEMGRAPH_BATCH_SIZE** defaults to `1000` (lines 68-69 in [`config.py`](https://github.com/vitali87/code-graph-rag/blob/main/config.py)). This parameter controls how many nodes and relationships the system buffers before executing a bulk write operation to Memgraph. The `AppConfig.resolve_batch_size()` method returns this value when no explicit size is provided to the ingester.

### Parallel Flushing

**FLUSH_THREAD_POOL_SIZE** defaults to `4` (line 30), defining the number of worker threads available for parallel node and relationship flushing operations. This setting optimizes throughput when processing large codebases by allowing concurrent write operations to the database.

### Query Memory Limits

**QUERY_MEMORY_LIMIT_MB** defaults to `4096` MB (line 364). This safety mechanism appends a memory limit clause to every Cypher query via the `_apply_memory_limit` helper in [`graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_service.py) (lines 76-84). If a query does not already contain `CYPHER_MEMORY_LIMIT_TOKEN` (defined in [`codebase_rag/constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants.py)), the system automatically appends `CYPHER_MEMORY_LIMIT_SUFFIX` to prevent runaway queries from consuming excessive resources.

## Configuration Loading Mechanism

`AppConfig` uses Pydantic's `BaseSettings` to merge hardcoded defaults with runtime overrides. The resolution order follows:

1. Hardcoded defaults in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py)
2. Environment variables (e.g., `export MEMGRAPH_HOST=memgraph.db`)
3. `.env` file variables in the project root

This architecture allows developers to modify default Memgraph settings for code-graph-rag without altering source code, ensuring containerized deployments can inject connection details dynamically.

## Implementation in MemgraphIngestor

The `MemgraphIngestor` class in [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py) consumes these defaults during initialization. When entering the context manager (`__enter__`), it establishes the connection using:

```python
self.conn = mgclient.connect(
    host=self._host,
    port=self._port,
    username=self._username,
    password=self._password,
)

```

See lines 16-24 in [`graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_service.py). The batch size controls the internal `node_buffer` and `_rel_groups` dictionaries, flushing automatically when the buffer reaches `MEMGRAPH_BATCH_SIZE` or when the context manager exits.

## Practical Configuration Example

```python
from codebase_rag.services.graph_service import MemgraphIngestor
from codebase_rag.config import settings

# Connect using default localhost:7687 with 1000-item batches

with MemgraphIngestor(
        host=settings.MEMGRAPH_HOST,
        port=settings.MEMGRAPH_PORT,
        batch_size=settings.MEMGRAPH_BATCH_SIZE,
) as ingestor:
    # Buffer entities for bulk insertion

    ingestor.ensure_node_batch("Function", {"name": "parse_code", "file": "parser.py"})
    ingestor.ensure_relationship_batch(
        ("Function", "name", "parse_code"),
        "CALLS",
        ("Function", "name", "tokenize"),
    )
    # Automatic flush occurs at batch size threshold or context exit

```

## Summary

- **Connection defaults**: Host is `localhost`, Bolt port is `7687`, HTTP port is `7444`, with optional authentication disabled by default.
- **Batch processing**: 1000 items per batch with 4 parallel flush threads optimize ingestion throughput.
- **Memory safety**: 4096 MB query memory limits prevent resource exhaustion via automatic Cypher suffix injection.
- **Configuration source**: All defaults reside in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py) within the `AppConfig` class, overridable via environment variables or `.env` files.

## Frequently Asked Questions

### How do I change the default Memgraph host in code-graph-rag?

Set the `MEMGRAPH_HOST` environment variable before starting your application. The `AppConfig` class in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py) automatically picks up this value, overriding the default `"localhost"`. For Docker deployments, this typically becomes `MEMGRAPH_HOST=memgraph` to match the service name.

### Is authentication required for the default Memgraph configuration?

No. The default `MEMGRAPH_USERNAME` and `MEMGRAPH_PASSWORD` values are `None`, allowing unauthenticated connections to local development instances. For production environments hosting sensitive code graphs, configure these variables to match your Memgraph user credentials.

### What happens if I increase the MEMGRAPH_BATCH_SIZE value?

Increasing `MEMGRAPH_BATCH_SIZE` beyond the default 1000 reduces the frequency of network round-trips to Memgraph, potentially improving ingestion speed for massive codebases. However, larger batches consume more client-side memory and may delay visibility of newly created nodes. The optimal value depends on available RAM and graph complexity.

### How does the query memory limit affect Cypher execution?

The `QUERY_MEMORY_LIMIT_MB` setting (default 4096 MB) applies a hard ceiling to query memory consumption via the `_apply_memory_limit` helper in [`graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_service.py). If a query approaches this limit, Memgraph terminates it to protect database stability. Increase this value only when processing extremely large dependency graphs that legitimately require more memory for path-finding operations.