# How to Export the Knowledge Graph from Code-Graph-RAG

> Learn how to export the knowledge graph from Code-Graph-RAG. Follow simple steps to serialize your code entities into a portable JSON file for analysis and sharing.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-19

---

**Export the knowledge graph from Code-Graph-RAG by initializing a `MemgraphIngestor`, running the ingestion pipeline to populate it with code entities, and calling `export_graph_to_file()` to serialize the graph to a portable JSON file.** Code-Graph-RAG indexes functions, classes, and modules into Memgraph, and the library provides a built-in export pipeline in [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py) that converts these database records into a schema-compliant JSON document suitable for backup, migration, or downstream analysis.

## Initialize the MemgraphIngestor

Before exporting, you must establish a connection to Memgraph through the **`MemgraphIngestor`** class. This component, defined in [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py) at line 86, manages the buffer of nodes and relationships while the ingestion runs.

Use the `connect_memgraph` factory function from [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py) to instantiate the ingestor with your desired batch size:

```python
from codebase_rag.main import connect_memgraph

ingestor = connect_memgraph(batch_size=1000)

```

## Populate the Knowledge Graph

Once the ingestor is initialized, execute the code analysis pipeline to fill the graph with entities. You can trigger this programmatically or via the CLI. The `MemgraphIngestor` accumulates nodes (functions, classes, modules) and their relationships (imports, inheritance, calls) in memory until you explicitly request an export.

```python
from codebase_rag.main import connect_memgraph, export_graph_to_file

# Initialize

ingestor = connect_memgraph(batch_size=2000)

# Run your ingestion (example: processing a codebase)

# This populates the ingestor with graph data

ingestor.process_codebase("./src")

```

## Serialize to JSON with export_graph_to_file

The **`export_graph_to_file()`** function in [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py) (around line 1403) provides the primary interface for exporting. It calls the ingestor's internal `export_graph_to_dict()` method, converts the result to JSON, and writes it to disk.

```python
from codebase_rag.main import export_graph_to_file

success = export_graph_to_file(ingestor, "graph_export.json")

if success:
    print("Export complete")
else:
    print("Export failed - check logs for details")

```

The function returns a boolean indicating success. On completion, it prints a console summary showing total node and relationship counts. Internally, `export_graph_to_dict` (implemented in [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py)) serializes the buffered graph data according to the protobuf schema.

## Understanding the Export File Structure

The exported JSON follows the schema defined in **`codec/schema.proto`** (exposed via [`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py)). The file structure contains three top-level keys:

- **`metadata`** – Export timestamp, version info, and aggregate statistics
- **`nodes`** – Array of node objects with labels, properties (name, source file path, docstring), and unique identifiers
- **`relationships`** – Array of edge objects defining the graph topology via source/target node IDs, relationship types (e.g., `CALLS`, `IMPORTS`), and property dictionaries

This structured format ensures compatibility with graph visualization tools and external machine learning pipelines.

## Loading and Verifying Exported Graphs

To inspect a previously exported file without re-connecting to Memgraph, use the **`GraphLoader`** utilities in [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) or the provided example script.

Load programmatically:

```python
from codebase_rag.graph_loader import load_graph

graph = load_graph("graph_export.json")
summary = graph.summary()
print(f"Nodes: {summary['total_nodes']}")
print(f"Relationships: {summary['total_relationships']}")

```

Alternatively, use the CLI example located at **[`examples/graph_export_example.py`](https://github.com/vitali87/code-graph-rag/blob/main/examples/graph_export_example.py)**:

```bash
python examples/graph_export_example.py path/to/graph_export.json

```

This script demonstrates loading via `load_graph` and printing detailed statistics, allowing you to verify export integrity before archiving or sharing the file.

## Summary

- **Initialize** a `MemgraphIngestor` via `connect_memgraph()` from [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py) to buffer graph data during ingestion.
- **Populate** the ingestor by running the code analysis pipeline against your target repository.
- **Export** by calling `export_graph_to_file()`, implemented at line 1403 in [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py), which serializes the graph to JSON.
- **Format** follows the protobuf schema in `codec/schema.proto`, preserving node labels, properties, and relationship metadata.
- **Verify** exports using `codebase_rag.graph_loader.load_graph()` or the reference implementation in [`examples/graph_export_example.py`](https://github.com/vitali87/code-graph-rag/blob/main/examples/graph_export_example.py).

## Frequently Asked Questions

### What file format does Code-Graph-RAG use for knowledge graph exports?

Code-Graph-RAG exports to **JSON** that conforms to the protobuf schema defined in `codec/schema.proto`. The file contains structured arrays for metadata, nodes, and relationships, making it compatible with graph databases, analysis tools, and backup systems.

### Can I export the graph without re-running the entire ingestion pipeline?

You can only export data currently held by a `MemgraphIngestor` instance. If you have an existing Memgraph database from a previous run but no longer have the ingestor object in memory, you must re-initialize the connection and re-populate the ingestor (or query Memgraph directly) before calling `export_graph_to_file()`.

### Where is the export_graph_to_file function defined?

The **`export_graph_to_file()`** function is defined in **[`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py)** around line 1403. It acts as a thin wrapper that invokes `export_graph_to_dict()` on the `MemgraphIngestor` class (found in [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py)) and handles JSON serialization and file I/O.

### How do I validate that my export completed successfully?

The function returns `True` on success and `False` on failure, logging exceptions when errors occur. For structural validation, use the **[`examples/graph_export_example.py`](https://github.com/vitali87/code-graph-rag/blob/main/examples/graph_export_example.py)** script or import `load_graph` from [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) to parse the JSON and inspect node/relationship counts against your expected totals.