How to Ingest Parsed Code into Memgraph with Code-Graph-RAG
Code-Graph-RAG stores parsed source code as a knowledge graph inside Memgraph using the MemgraphIngestor class, which batch-loads nodes and relationships via efficient UNWIND-based Cypher queries.
Code-Graph-RAG transforms entire codebases into rich, queryable knowledge graphs by parsing source files into abstract syntax tree (AST) representations and loading them into Memgraph. The ingestion pipeline centers on the MemgraphIngestor class defined in codebase_rag/services/graph_service.py, which manages Bolt protocol connections, memory-efficient batching, and transactional integrity when you ingest parsed code into Memgraph with Code-Graph-RAG.
Creating the Memgraph Connection with Context Managers
The MemgraphIngestor class wraps the mgclient Bolt driver to provide thread-safe connection pooling and automatic resource cleanup. Initialize the ingestor with your Memgraph host, port, and desired batch size, then use it as a context manager to ensure connections open and close automatically.
from codebase_rag.services.graph_service import MemgraphIngestor
with MemgraphIngestor(host="localhost", port=7687, batch_size=500) as ingestor:
# Batching operations execute here
pass # Connection closes automatically on exit
The batch_size parameter configures the internal buffer threshold. When the buffer reaches this limit, the ingestor automatically flushes accumulated entities to Memgraph using optimized Cypher statements.
Batching Nodes and Relationships
The ingestion pipeline queues parsed entities into in-memory buffers through two primary methods: ensure_node_batch() and ensure_relationship_batch(). Both methods group entities by pattern to minimize query complexity and maximize throughput when you ingest parsed code into Memgraph with Code-Graph-RAG.
Queuing Node Entities
Call ensure_node_batch(label, properties) to stage a node definition. The method stores the entity in an internal dictionary keyed by label, batching until the configured batch_size triggers an automatic flush.
# Ingesting a parsed class node
ingestor.ensure_node_batch(
label="Class",
properties={
"name": "DataProcessor",
"file_path": "/src/processor.py",
"line_number": 42
}
)
Queuing Relationship Entities
For edges, use ensure_relationship_batch(from_spec, rel_type, to_spec, properties), where from_spec and to_spec are tuples defining the endpoint identifiers: (label, key, value).
# Ingesting an inheritance relationship
ingestor.ensure_relationship_batch(
from_spec=("Class", "name", "DataProcessor"),
rel_type="INHERITS_FROM",
to_spec=("Class", "name", "BaseProcessor"),
properties={}
)
This method groups relationships by their source label, target label, and relationship type, enabling the ingestor to generate efficient MERGE or CREATE statements that handle multiple edges in a single query.
Flushing Buffers to Memgraph
After queuing all entities, execute flush_all() to write remaining buffers to the database. When using the context manager, this method invokes automatically during __exit__, ensuring no data loss even if exceptions occur.
The flush_all() method (defined at line 150 of codebase_rag/services/graph_service.py) generates UNWIND-based Cypher queries that process entire batches in single transactions. This approach minimizes round-trips and leverages Memgraph's bulk import capabilities.
# Explicit flush (optional when using context manager)
ingestor.flush_all()
Behind the scenes, the ingestor applies configurable memory limits via _apply_memory_limit() to prevent large queries from exhausting server resources. It also respects the use_merge flag, choosing between MERGE (idempotent) and CREATE (performant for confirmed new data) based on your deduplication requirements.
Complete End-to-End Ingestion Workflow
Combine the parser and ingestor to transform a local repository into a Memgraph knowledge graph. The GraphUpdater class in codebase_rag/graph_updater.py handles directory traversal and AST extraction, returning lightweight node and relationship objects.
from codebase_rag.services.graph_service import MemgraphIngestor
from codebase_rag.graph_updater import GraphUpdater
# Initialize ingestor with 500-item batch threshold
with MemgraphIngestor(host="localhost", port=7687, batch_size=500) as ingestor:
# Parse repository into graph structure
updater = GraphUpdater()
graph = updater.parse_path("/path/to/your/project")
# Batch load all nodes
for node in graph.nodes:
ingestor.ensure_node_batch(
label=node.label,
properties=node.props
)
# Batch load all relationships (calls, imports, inheritance)
for rel in graph.relationships:
ingestor.ensure_relationship_batch(
from_spec=(rel.from_label, rel.from_key, rel.from_val),
rel_type=rel.type,
to_spec=(rel.to_label, rel.to_key, rel.to_val),
properties=rel.props
)
# Final flush happens automatically on context exit
This workflow efficiently handles large codebases by streaming parsed entities through memory-bounded buffers rather than loading entire graphs into RAM.
Implementation Details and Advanced Features
The MemgraphIngestor provides several mechanisms to optimize throughput and reliability:
- Parallel Execution: When a
ThreadPoolExecutoris available, the ingestor executes independent batch writes concurrently, saturating network bandwidth without overwhelming the Memgraph instance. - Memory-Constrained Queries: The
_apply_memory_limit()helper appends optional query suffixes to prevent out-of-memory errors on massive batches. - Diagnostic Utilities: Methods like
export_graph_to_dict(),fetch_all(), andexecute_write()support debugging and validation by allowing direct Cypher execution and result inspection. - Constraint Enforcement: The ingestor respects schema constraints defined in
codebase_rag/constants/graph.py, ensuring that node labels and relationship types match the expected ontology.
CLI users can trigger this entire pipeline via cgr ingest defined in codebase_rag/graph_cli.py, which wires the GraphUpdater and MemgraphIngestor together for command-line repository ingestion.
Summary
- MemgraphIngestor in
codebase_rag/services/graph_service.pyprovides the primary interface for loading parsed code into Memgraph. - Use context managers to handle connection lifecycle automatically, ensuring buffers flush and connections close via
__exit__. - Stage entities with batch methods:
ensure_node_batch()for vertices andensure_relationship_batch()for edges. - Automatic flushing occurs at
batch_sizethresholds; manualflush_all()ensures persistence before shutdown. - The parser in GraphUpdater (
codebase_rag/graph_updater.py) produces the node/relationship streams consumed by the ingestor. - UNWIND-based Cypher and optional parallel execution provide high-throughput ingestion suitable for enterprise codebases.
Frequently Asked Questions
What is the optimal batch_size for ingesting large repositories?
Set batch_size between 500 and 2000 depending on available RAM and relationship complexity. Smaller batches reduce memory pressure but increase network round-trips; larger batches maximize throughput but may trigger Memgraph's memory limits without the _apply_memory_limit safeguard.
How does Code-Graph-RAG handle duplicate nodes during ingestion?
The ingestor respects the use_merge boolean flag. When True, it generates MERGE statements that match existing nodes before creation, preventing duplicates. When False, it uses faster CREATE statements suitable for fresh databases or immutable snapshots.
Can I ingest multiple repositories into the same Memgraph instance?
Yes. Instantiate separate GraphUpdater sessions for each repository path and stream them through a single MemgraphIngestor instance. Ensure your node properties include repository identifiers to prevent cross-project collisions, or use MERGE semantics with unique constraints.
Where are the node labels and relationship types defined?
Schema constants reside in codebase_rag/constants/graph.py, which centralizes label definitions (e.g., Class, Function, Module) and relationship types (e.g., CALLS, IMPORTS, INHERITS_FROM). The ingestor validates queued entities against these constants to maintain graph consistency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →