# How to Ingest Parsed Code into Memgraph with Code-Graph-RAG

> Learn to ingest parsed code into Memgraph using Code-Graph-RAG. Discover how the MemgraphIngestor class efficiently loads nodes and relationships via Cypher for powerful code knowledge graphs.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-04

---

**Code-Graph-RAG stores parsed source code as a knowledge graph inside Memgraph using the `MemgraphIngestor` class, which batch-loads nodes and relationships via efficient UNWIND-based Cypher queries.**

Code-Graph-RAG transforms entire codebases into rich, queryable knowledge graphs by parsing source files into abstract syntax tree (AST) representations and loading them into Memgraph. The ingestion pipeline centers on the `MemgraphIngestor` class defined in [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py), which manages Bolt protocol connections, memory-efficient batching, and transactional integrity when you ingest parsed code into Memgraph with Code-Graph-RAG.

## Creating the Memgraph Connection with Context Managers

The `MemgraphIngestor` class wraps the `mgclient` Bolt driver to provide thread-safe connection pooling and automatic resource cleanup. Initialize the ingestor with your Memgraph host, port, and desired batch size, then use it as a context manager to ensure connections open and close automatically.

```python
from codebase_rag.services.graph_service import MemgraphIngestor

with MemgraphIngestor(host="localhost", port=7687, batch_size=500) as ingestor:
    # Batching operations execute here

    pass  # Connection closes automatically on exit

```

The `batch_size` parameter configures the internal buffer threshold. When the buffer reaches this limit, the ingestor automatically flushes accumulated entities to Memgraph using optimized Cypher statements.

## Batching Nodes and Relationships

The ingestion pipeline queues parsed entities into in-memory buffers through two primary methods: `ensure_node_batch()` and `ensure_relationship_batch()`. Both methods group entities by pattern to minimize query complexity and maximize throughput when you ingest parsed code into Memgraph with Code-Graph-RAG.

### Queuing Node Entities

Call `ensure_node_batch(label, properties)` to stage a node definition. The method stores the entity in an internal dictionary keyed by label, batching until the configured `batch_size` triggers an automatic flush.

```python

# Ingesting a parsed class node

ingestor.ensure_node_batch(
    label="Class",
    properties={
        "name": "DataProcessor",
        "file_path": "/src/processor.py",
        "line_number": 42
    }
)

```

### Queuing Relationship Entities

For edges, use `ensure_relationship_batch(from_spec, rel_type, to_spec, properties)`, where `from_spec` and `to_spec` are tuples defining the endpoint identifiers: `(label, key, value)`.

```python

# Ingesting an inheritance relationship

ingestor.ensure_relationship_batch(
    from_spec=("Class", "name", "DataProcessor"),
    rel_type="INHERITS_FROM",
    to_spec=("Class", "name", "BaseProcessor"),
    properties={}
)

```

This method groups relationships by their source label, target label, and relationship type, enabling the ingestor to generate efficient `MERGE` or `CREATE` statements that handle multiple edges in a single query.

## Flushing Buffers to Memgraph

After queuing all entities, execute `flush_all()` to write remaining buffers to the database. When using the context manager, this method invokes automatically during `__exit__`, ensuring no data loss even if exceptions occur.

The `flush_all()` method (defined at line 150 of [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py)) generates UNWIND-based Cypher queries that process entire batches in single transactions. This approach minimizes round-trips and leverages Memgraph's bulk import capabilities.

```python

# Explicit flush (optional when using context manager)

ingestor.flush_all()

```

Behind the scenes, the ingestor applies configurable memory limits via `_apply_memory_limit()` to prevent large queries from exhausting server resources. It also respects the `use_merge` flag, choosing between `MERGE` (idempotent) and `CREATE` (performant for confirmed new data) based on your deduplication requirements.

## Complete End-to-End Ingestion Workflow

Combine the parser and ingestor to transform a local repository into a Memgraph knowledge graph. The `GraphUpdater` class in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) handles directory traversal and AST extraction, returning lightweight node and relationship objects.

```python
from codebase_rag.services.graph_service import MemgraphIngestor
from codebase_rag.graph_updater import GraphUpdater

# Initialize ingestor with 500-item batch threshold

with MemgraphIngestor(host="localhost", port=7687, batch_size=500) as ingestor:
    
    # Parse repository into graph structure

    updater = GraphUpdater()
    graph = updater.parse_path("/path/to/your/project")
    
    # Batch load all nodes

    for node in graph.nodes:
        ingestor.ensure_node_batch(
            label=node.label, 
            properties=node.props
        )
    
    # Batch load all relationships (calls, imports, inheritance)

    for rel in graph.relationships:
        ingestor.ensure_relationship_batch(
            from_spec=(rel.from_label, rel.from_key, rel.from_val),
            rel_type=rel.type,
            to_spec=(rel.to_label, rel.to_key, rel.to_val),
            properties=rel.props
        )
    
    # Final flush happens automatically on context exit

```

This workflow efficiently handles large codebases by streaming parsed entities through memory-bounded buffers rather than loading entire graphs into RAM.

## Implementation Details and Advanced Features

The `MemgraphIngestor` provides several mechanisms to optimize throughput and reliability:

- **Parallel Execution**: When a `ThreadPoolExecutor` is available, the ingestor executes independent batch writes concurrently, saturating network bandwidth without overwhelming the Memgraph instance.
- **Memory-Constrained Queries**: The `_apply_memory_limit()` helper appends optional query suffixes to prevent out-of-memory errors on massive batches.
- **Diagnostic Utilities**: Methods like `export_graph_to_dict()`, `fetch_all()`, and `execute_write()` support debugging and validation by allowing direct Cypher execution and result inspection.
- **Constraint Enforcement**: The ingestor respects schema constraints defined in [`codebase_rag/constants/graph.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/graph.py), ensuring that node labels and relationship types match the expected ontology.

CLI users can trigger this entire pipeline via `cgr ingest` defined in [`codebase_rag/graph_cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_cli.py), which wires the `GraphUpdater` and `MemgraphIngestor` together for command-line repository ingestion.

## Summary

- **MemgraphIngestor** in [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py) provides the primary interface for loading parsed code into Memgraph.
- Use **context managers** to handle connection lifecycle automatically, ensuring buffers flush and connections close via `__exit__`.
- Stage entities with **batch methods**: `ensure_node_batch()` for vertices and `ensure_relationship_batch()` for edges.
- **Automatic flushing** occurs at `batch_size` thresholds; manual `flush_all()` ensures persistence before shutdown.
- The parser in **GraphUpdater** ([`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py)) produces the node/relationship streams consumed by the ingestor.
- **UNWIND-based Cypher** and optional parallel execution provide high-throughput ingestion suitable for enterprise codebases.

## Frequently Asked Questions

### What is the optimal batch_size for ingesting large repositories?

Set `batch_size` between 500 and 2000 depending on available RAM and relationship complexity. Smaller batches reduce memory pressure but increase network round-trips; larger batches maximize throughput but may trigger Memgraph's memory limits without the `_apply_memory_limit` safeguard.

### How does Code-Graph-RAG handle duplicate nodes during ingestion?

The ingestor respects the `use_merge` boolean flag. When `True`, it generates `MERGE` statements that match existing nodes before creation, preventing duplicates. When `False`, it uses faster `CREATE` statements suitable for fresh databases or immutable snapshots.

### Can I ingest multiple repositories into the same Memgraph instance?

Yes. Instantiate separate `GraphUpdater` sessions for each repository path and stream them through a single `MemgraphIngestor` instance. Ensure your node properties include repository identifiers to prevent cross-project collisions, or use `MERGE` semantics with unique constraints.

### Where are the node labels and relationship types defined?

Schema constants reside in [`codebase_rag/constants/graph.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/graph.py), which centralizes label definitions (e.g., `Class`, `Function`, `Module`) and relationship types (e.g., `CALLS`, `IMPORTS`, `INHERITS_FROM`). The ingestor validates queued entities against these constants to maintain graph consistency.