How to Export the Code‑Graph‑RAG Knowledge Graph: A Complete Guide
Export your entire Code‑Graph‑RAG knowledge graph from Memgraph to a portable JSON file using either Python helper functions or the built‑in CLI command.
Code‑Graph‑RAG persists its knowledge graph inside a transient Memgraph instance. To move this data into other tools—visualizers, static analyzers, or alternative graph databases—you need to serialize the graph to disk. This article walks through the export architecture, the GraphData schema, and both programmatic and command‑line approaches as implemented in the vitali87/code‑graph‑rag repository.
Export Architecture Overview
The export flow relies on three core components defined in the codebase:
MemgraphIngestor– connects to Memgraph and batches Cypher queries.export_graph_to_dict()– serializes the live graph into aGraphDatadictionary.export_graph_to_file()– writes pretty‑printed JSON and prints a summary.
These are orchestrated through the public entry point in codebase_rag/main.py, with a parallel CLI wrapper in codebase_rag/graph_cli.py that exposes the cgr graph export command.
Export Methods for Code‑Graph‑RAG Knowledge Graphs
Option 1: Python API (Programmatic Export)
For custom pipelines or Jupyter notebooks, import the connection helper and export function directly:
from pathlib import Path
from codebase_rag.main import connect_memgraph, export_graph_to_file
# 1️⃣ Create the Memgraph ingestor (batch size from settings)
ingestor = connect_memgraph(batch_size=1000)
# 2️⃣ Define your output path
output_path = Path("my_graph_export.json")
# 3️⃣ Execute the export
if export_graph_to_file(ingestor, str(output_path)):
print(f"✅ Graph exported successfully to {output_path.resolve()}")
else:
print("❌ Export failed – see console logs for details")
The connect_memgraph() helper reads host, port, and credentials from codebase_rag/constants.py, then instantiates a MemgraphIngestor with your specified batch size. The export_graph_to_file() function delegates to _write_graph_json() in codebase_rag/main.py for actual disk I/O.
Option 2: Command‑Line Interface (CLI)
For CI/CD pipelines or one‑off terminal usage, use the cgr CLI:
cgr graph export /path/to/export.json
Available flags:
--project <NAME>– override the auto‑detected project name in metadata.--repo-path <DIR>– source directory to tag the export with (defaults to$PWD).
On success, the CLI prints a statistics table showing total nodes and relationships extracted from your Code‑Graph‑RAG knowledge graph.
Understanding the Exported JSON Format
The output follows the GraphData schema defined in codebase_rag/types_defs.py:
{
"metadata": {
"exported_at": "2024-01-01T12:34:56Z",
"total_nodes": 12345,
"total_relationships": 67890
},
"nodes": [
{
"id": "n1",
"label": "Function",
"properties": {
"name": "my_function",
"path": "src/module.py"
}
}
],
"relationships": [
{
"id": "r1",
"type": "CALLS",
"source": "n1",
"target": "n2",
"properties": {}
}
]
}
| Field | Description |
|---|---|
metadata |
ISO‑8601 timestamp and aggregate counts |
nodes |
Graph entities (Function, Class, Module) with unique IDs and property bags |
relationships |
Directed edges (CALLS, IMPORTS, EXTENDS, etc.) linking source/target node IDs |
Because this is plain JSON, you can import the Code‑Graph‑RAG knowledge graph into NetworkX, Neo4j, or D3.js visualizations without transformation.
Practical Export Examples
Standalone Export Script
Create export_graph_demo.py for reusable exports:
from codebase_rag.main import connect_memgraph, export_graph_to_file
def main():
ingestor = connect_memgraph(batch_size=500)
success = export_graph_to_file(ingestor, "graph_dump.json")
print("Exported successfully" if success else "Export failed")
if __name__ == "__main__":
main()
Run with:
python export_graph_demo.py
CI Pipeline Automation
Add this GitHub Actions workflow to export on every push:
name: Export Graph
on:
push:
branches: [main]
jobs:
export:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: pip install -e .
- run: cgr graph export ./graph_export.json
- uses: actions/upload-artifact@v4
with:
name: codegraph
path: ./graph_export.json
Reimporting into NetworkX
Process the exported Code‑Graph‑RAG knowledge graph with Python graph libraries:
import json
import networkx as nx
with open("graph_export.json") as f:
data = json.load(f)
G = nx.DiGraph()
for node in data["nodes"]:
G.add_node(node["id"], **node["properties"])
for rel in data["relationships"]:
G.add_edge(rel["source"], rel["target"], type=rel["type"])
print(f"Loaded {G.number_of_nodes()} nodes, {G.number_of_edges()} edges")
Key Source Files in Code‑Graph‑RAG
| File | Purpose |
|---|---|
codebase_rag/main.py |
Core export helpers (_write_graph_json, export_graph_to_file) |
codebase_rag/graph_cli.py |
CLI command group (cgr graph export …) |
codebase_rag/services/graph_service.py |
MemgraphIngestor.export_graph_to_dict() implementation |
codebase_rag/types_defs.py |
GraphData Pydantic schema for serialization |
examples/graph_export_example.py |
End‑to‑end runnable example |
Summary
- Two export paths: Python API via
export_graph_to_file()or CLI viacgr graph export. - Standardized output: JSON matching the
GraphDataschema withmetadata,nodes, andrelationshipskeys. - Configurable batching: Tune
batch_sizeinconnect_memgraph()for memory‑constrained environments. - Tool‑agnostic: Consume exports in NetworkX, Neo4j, visualization suites, or custom analytics.
Frequently Asked Questions
What graph database does Code‑Graph‑RAG use?
Code‑Graph‑RAG uses Memgraph as its primary graph storage engine. The MemgraphIngestor class in codebase_rag/services/graph_service.py handles all Cypher query execution and connection pooling.
Can I export partial graphs or filter by node type?
The current implementation in export_graph_to_dict() performs a full‑graph dump. For filtered exports, you would need to extend MemgraphIngestor with Cypher WHERE clauses or post‑process the resulting JSON.
Is the exported JSON compatible with Neo4j?
Yes. The nodes/relationships structure maps cleanly to Neo4j's property graph model. Use neo4j-admin import or the APOC apoc.import.json procedure with minor field renaming.
How do I automate exports on a schedule?
Combine the CLI with cron or GitHub Actions. Ensure your Memgraph instance is accessible from the automation environment and that codebase_rag/constants.py contains valid connection parameters.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →