Default Memgraph Settings for code-graph-rag: Complete Configuration Reference
The code-graph-rag project configures Memgraph defaults in codebase_rag/config.py, using localhost:7687 for connections, batch sizes of 1000 nodes/relationships, and a 4096 MB query memory limit.
The open-source vitali87/code-graph-rag repository leverages Memgraph as its primary graph database for storing code relationships. All default Memgraph settings for code-graph-rag are centralized in the AppConfig Pydantic settings class located in codebase_rag/config.py, which automatically loads environment variables from .env files to override hardcoded defaults.
Connection Parameters
Host, Port, and Protocol
The foundation of the Memgraph connection relies on three key settings defined in codebase_rag/config.py:
- MEMGRAPH_HOST: Defaults to
"localhost"(line 62). This specifies the hostname where the Memgraph instance is reachable. - MEMGRAPH_PORT: Defaults to
7687(line 63). This is the Bolt-protocol port used by the Pythonmgclientdriver to establish binary connections. - MEMGRAPH_HTTP_PORT: Defaults to
7444(line 64). This port serves Memgraph Lab and optional REST endpoints, though the primary ingestion uses the Bolt protocol.
Authentication Defaults
By default, code-graph-rag assumes an unsecured local Memgraph instance:
- MEMGRAPH_USERNAME: Defaults to
None(line 65). - MEMGRAPH_PASSWORD: Defaults to
None(line 66).
When both values remain None, the MemgraphIngestor class in codebase_rag/services/graph_service.py connects without authentication. For production deployments, these can be overridden via environment variables to enable Memgraph's role-based access control.
Performance and Batching Configuration
Ingestion Batch Size
MEMGRAPH_BATCH_SIZE defaults to 1000 (lines 68-69 in config.py). This parameter controls how many nodes and relationships the system buffers before executing a bulk write operation to Memgraph. The AppConfig.resolve_batch_size() method returns this value when no explicit size is provided to the ingester.
Parallel Flushing
FLUSH_THREAD_POOL_SIZE defaults to 4 (line 30), defining the number of worker threads available for parallel node and relationship flushing operations. This setting optimizes throughput when processing large codebases by allowing concurrent write operations to the database.
Query Memory Limits
QUERY_MEMORY_LIMIT_MB defaults to 4096 MB (line 364). This safety mechanism appends a memory limit clause to every Cypher query via the _apply_memory_limit helper in graph_service.py (lines 76-84). If a query does not already contain CYPHER_MEMORY_LIMIT_TOKEN (defined in codebase_rag/constants.py), the system automatically appends CYPHER_MEMORY_LIMIT_SUFFIX to prevent runaway queries from consuming excessive resources.
Configuration Loading Mechanism
AppConfig uses Pydantic's BaseSettings to merge hardcoded defaults with runtime overrides. The resolution order follows:
- Hardcoded defaults in
codebase_rag/config.py - Environment variables (e.g.,
export MEMGRAPH_HOST=memgraph.db) .envfile variables in the project root
This architecture allows developers to modify default Memgraph settings for code-graph-rag without altering source code, ensuring containerized deployments can inject connection details dynamically.
Implementation in MemgraphIngestor
The MemgraphIngestor class in codebase_rag/services/graph_service.py consumes these defaults during initialization. When entering the context manager (__enter__), it establishes the connection using:
self.conn = mgclient.connect(
host=self._host,
port=self._port,
username=self._username,
password=self._password,
)
See lines 16-24 in graph_service.py. The batch size controls the internal node_buffer and _rel_groups dictionaries, flushing automatically when the buffer reaches MEMGRAPH_BATCH_SIZE or when the context manager exits.
Practical Configuration Example
from codebase_rag.services.graph_service import MemgraphIngestor
from codebase_rag.config import settings
# Connect using default localhost:7687 with 1000-item batches
with MemgraphIngestor(
host=settings.MEMGRAPH_HOST,
port=settings.MEMGRAPH_PORT,
batch_size=settings.MEMGRAPH_BATCH_SIZE,
) as ingestor:
# Buffer entities for bulk insertion
ingestor.ensure_node_batch("Function", {"name": "parse_code", "file": "parser.py"})
ingestor.ensure_relationship_batch(
("Function", "name", "parse_code"),
"CALLS",
("Function", "name", "tokenize"),
)
# Automatic flush occurs at batch size threshold or context exit
Summary
- Connection defaults: Host is
localhost, Bolt port is7687, HTTP port is7444, with optional authentication disabled by default. - Batch processing: 1000 items per batch with 4 parallel flush threads optimize ingestion throughput.
- Memory safety: 4096 MB query memory limits prevent resource exhaustion via automatic Cypher suffix injection.
- Configuration source: All defaults reside in
codebase_rag/config.pywithin theAppConfigclass, overridable via environment variables or.envfiles.
Frequently Asked Questions
How do I change the default Memgraph host in code-graph-rag?
Set the MEMGRAPH_HOST environment variable before starting your application. The AppConfig class in codebase_rag/config.py automatically picks up this value, overriding the default "localhost". For Docker deployments, this typically becomes MEMGRAPH_HOST=memgraph to match the service name.
Is authentication required for the default Memgraph configuration?
No. The default MEMGRAPH_USERNAME and MEMGRAPH_PASSWORD values are None, allowing unauthenticated connections to local development instances. For production environments hosting sensitive code graphs, configure these variables to match your Memgraph user credentials.
What happens if I increase the MEMGRAPH_BATCH_SIZE value?
Increasing MEMGRAPH_BATCH_SIZE beyond the default 1000 reduces the frequency of network round-trips to Memgraph, potentially improving ingestion speed for massive codebases. However, larger batches consume more client-side memory and may delay visibility of newly created nodes. The optimal value depends on available RAM and graph complexity.
How does the query memory limit affect Cypher execution?
The QUERY_MEMORY_LIMIT_MB setting (default 4096 MB) applies a hard ceiling to query memory consumption via the _apply_memory_limit helper in graph_service.py. If a query approaches this limit, Memgraph terminates it to protect database stability. Increase this value only when processing extremely large dependency graphs that legitimately require more memory for path-finding operations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →