Default Memgraph Settings for code-graph-rag: Complete Configuration Reference

The code-graph-rag project configures Memgraph defaults in codebase_rag/config.py, using localhost:7687 for connections, batch sizes of 1000 nodes/relationships, and a 4096 MB query memory limit.

The open-source vitali87/code-graph-rag repository leverages Memgraph as its primary graph database for storing code relationships. All default Memgraph settings for code-graph-rag are centralized in the AppConfig Pydantic settings class located in codebase_rag/config.py, which automatically loads environment variables from .env files to override hardcoded defaults.

Connection Parameters

Host, Port, and Protocol

The foundation of the Memgraph connection relies on three key settings defined in codebase_rag/config.py:

  • MEMGRAPH_HOST: Defaults to "localhost" (line 62). This specifies the hostname where the Memgraph instance is reachable.
  • MEMGRAPH_PORT: Defaults to 7687 (line 63). This is the Bolt-protocol port used by the Python mgclient driver to establish binary connections.
  • MEMGRAPH_HTTP_PORT: Defaults to 7444 (line 64). This port serves Memgraph Lab and optional REST endpoints, though the primary ingestion uses the Bolt protocol.

Authentication Defaults

By default, code-graph-rag assumes an unsecured local Memgraph instance:

  • MEMGRAPH_USERNAME: Defaults to None (line 65).
  • MEMGRAPH_PASSWORD: Defaults to None (line 66).

When both values remain None, the MemgraphIngestor class in codebase_rag/services/graph_service.py connects without authentication. For production deployments, these can be overridden via environment variables to enable Memgraph's role-based access control.

Performance and Batching Configuration

Ingestion Batch Size

MEMGRAPH_BATCH_SIZE defaults to 1000 (lines 68-69 in config.py). This parameter controls how many nodes and relationships the system buffers before executing a bulk write operation to Memgraph. The AppConfig.resolve_batch_size() method returns this value when no explicit size is provided to the ingester.

Parallel Flushing

FLUSH_THREAD_POOL_SIZE defaults to 4 (line 30), defining the number of worker threads available for parallel node and relationship flushing operations. This setting optimizes throughput when processing large codebases by allowing concurrent write operations to the database.

Query Memory Limits

QUERY_MEMORY_LIMIT_MB defaults to 4096 MB (line 364). This safety mechanism appends a memory limit clause to every Cypher query via the _apply_memory_limit helper in graph_service.py (lines 76-84). If a query does not already contain CYPHER_MEMORY_LIMIT_TOKEN (defined in codebase_rag/constants.py), the system automatically appends CYPHER_MEMORY_LIMIT_SUFFIX to prevent runaway queries from consuming excessive resources.

Configuration Loading Mechanism

AppConfig uses Pydantic's BaseSettings to merge hardcoded defaults with runtime overrides. The resolution order follows:

  1. Hardcoded defaults in codebase_rag/config.py
  2. Environment variables (e.g., export MEMGRAPH_HOST=memgraph.db)
  3. .env file variables in the project root

This architecture allows developers to modify default Memgraph settings for code-graph-rag without altering source code, ensuring containerized deployments can inject connection details dynamically.

Implementation in MemgraphIngestor

The MemgraphIngestor class in codebase_rag/services/graph_service.py consumes these defaults during initialization. When entering the context manager (__enter__), it establishes the connection using:

self.conn = mgclient.connect(
    host=self._host,
    port=self._port,
    username=self._username,
    password=self._password,
)

See lines 16-24 in graph_service.py. The batch size controls the internal node_buffer and _rel_groups dictionaries, flushing automatically when the buffer reaches MEMGRAPH_BATCH_SIZE or when the context manager exits.

Practical Configuration Example

from codebase_rag.services.graph_service import MemgraphIngestor
from codebase_rag.config import settings

# Connect using default localhost:7687 with 1000-item batches

with MemgraphIngestor(
        host=settings.MEMGRAPH_HOST,
        port=settings.MEMGRAPH_PORT,
        batch_size=settings.MEMGRAPH_BATCH_SIZE,
) as ingestor:
    # Buffer entities for bulk insertion

    ingestor.ensure_node_batch("Function", {"name": "parse_code", "file": "parser.py"})
    ingestor.ensure_relationship_batch(
        ("Function", "name", "parse_code"),
        "CALLS",
        ("Function", "name", "tokenize"),
    )
    # Automatic flush occurs at batch size threshold or context exit

Summary

  • Connection defaults: Host is localhost, Bolt port is 7687, HTTP port is 7444, with optional authentication disabled by default.
  • Batch processing: 1000 items per batch with 4 parallel flush threads optimize ingestion throughput.
  • Memory safety: 4096 MB query memory limits prevent resource exhaustion via automatic Cypher suffix injection.
  • Configuration source: All defaults reside in codebase_rag/config.py within the AppConfig class, overridable via environment variables or .env files.

Frequently Asked Questions

How do I change the default Memgraph host in code-graph-rag?

Set the MEMGRAPH_HOST environment variable before starting your application. The AppConfig class in codebase_rag/config.py automatically picks up this value, overriding the default "localhost". For Docker deployments, this typically becomes MEMGRAPH_HOST=memgraph to match the service name.

Is authentication required for the default Memgraph configuration?

No. The default MEMGRAPH_USERNAME and MEMGRAPH_PASSWORD values are None, allowing unauthenticated connections to local development instances. For production environments hosting sensitive code graphs, configure these variables to match your Memgraph user credentials.

What happens if I increase the MEMGRAPH_BATCH_SIZE value?

Increasing MEMGRAPH_BATCH_SIZE beyond the default 1000 reduces the frequency of network round-trips to Memgraph, potentially improving ingestion speed for massive codebases. However, larger batches consume more client-side memory and may delay visibility of newly created nodes. The optimal value depends on available RAM and graph complexity.

How does the query memory limit affect Cypher execution?

The QUERY_MEMORY_LIMIT_MB setting (default 4096 MB) applies a hard ceiling to query memory consumption via the _apply_memory_limit helper in graph_service.py. If a query approaches this limit, Memgraph terminates it to protect database stability. Increase this value only when processing extremely large dependency graphs that legitimately require more memory for path-finding operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →