How to Export the Code‑Graph‑RAG Knowledge Graph: A Complete Guide

Export your entire Code‑Graph‑RAG knowledge graph from Memgraph to a portable JSON file using either Python helper functions or the built‑in CLI command.

Code‑Graph‑RAG persists its knowledge graph inside a transient Memgraph instance. To move this data into other tools—visualizers, static analyzers, or alternative graph databases—you need to serialize the graph to disk. This article walks through the export architecture, the GraphData schema, and both programmatic and command‑line approaches as implemented in the vitali87/code‑graph‑rag repository.


Export Architecture Overview

The export flow relies on three core components defined in the codebase:

  1. MemgraphIngestor – connects to Memgraph and batches Cypher queries.
  2. export_graph_to_dict() – serializes the live graph into a GraphData dictionary.
  3. export_graph_to_file() – writes pretty‑printed JSON and prints a summary.

These are orchestrated through the public entry point in codebase_rag/main.py, with a parallel CLI wrapper in codebase_rag/graph_cli.py that exposes the cgr graph export command.


Export Methods for Code‑Graph‑RAG Knowledge Graphs

Option 1: Python API (Programmatic Export)

For custom pipelines or Jupyter notebooks, import the connection helper and export function directly:

from pathlib import Path
from codebase_rag.main import connect_memgraph, export_graph_to_file

# 1️⃣  Create the Memgraph ingestor (batch size from settings)

ingestor = connect_memgraph(batch_size=1000)

# 2️⃣  Define your output path

output_path = Path("my_graph_export.json")

# 3️⃣  Execute the export

if export_graph_to_file(ingestor, str(output_path)):
    print(f"✅ Graph exported successfully to {output_path.resolve()}")
else:
    print("❌ Export failed – see console logs for details")

The connect_memgraph() helper reads host, port, and credentials from codebase_rag/constants.py, then instantiates a MemgraphIngestor with your specified batch size. The export_graph_to_file() function delegates to _write_graph_json() in codebase_rag/main.py for actual disk I/O.

Option 2: Command‑Line Interface (CLI)

For CI/CD pipelines or one‑off terminal usage, use the cgr CLI:

cgr graph export /path/to/export.json

Available flags:

  • --project <NAME> – override the auto‑detected project name in metadata.
  • --repo-path <DIR> – source directory to tag the export with (defaults to $PWD).

On success, the CLI prints a statistics table showing total nodes and relationships extracted from your Code‑Graph‑RAG knowledge graph.


Understanding the Exported JSON Format

The output follows the GraphData schema defined in codebase_rag/types_defs.py:

{
  "metadata": {
    "exported_at": "2024-01-01T12:34:56Z",
    "total_nodes": 12345,
    "total_relationships": 67890
  },
  "nodes": [
    {
      "id": "n1",
      "label": "Function",
      "properties": {
        "name": "my_function",
        "path": "src/module.py"
      }
    }
  ],
  "relationships": [
    {
      "id": "r1",
      "type": "CALLS",
      "source": "n1",
      "target": "n2",
      "properties": {}
    }
  ]
}
Field Description
metadata ISO‑8601 timestamp and aggregate counts
nodes Graph entities (Function, Class, Module) with unique IDs and property bags
relationships Directed edges (CALLS, IMPORTS, EXTENDS, etc.) linking source/target node IDs

Because this is plain JSON, you can import the Code‑Graph‑RAG knowledge graph into NetworkX, Neo4j, or D3.js visualizations without transformation.


Practical Export Examples

Standalone Export Script

Create export_graph_demo.py for reusable exports:

from codebase_rag.main import connect_memgraph, export_graph_to_file

def main():
    ingestor = connect_memgraph(batch_size=500)
    success = export_graph_to_file(ingestor, "graph_dump.json")
    print("Exported successfully" if success else "Export failed")

if __name__ == "__main__":
    main()

Run with:

python export_graph_demo.py

CI Pipeline Automation

Add this GitHub Actions workflow to export on every push:

name: Export Graph
on:
  push:
    branches: [main]

jobs:
  export:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: pip install -e .
      - run: cgr graph export ./graph_export.json
      - uses: actions/upload-artifact@v4
        with:
          name: codegraph
          path: ./graph_export.json

Reimporting into NetworkX

Process the exported Code‑Graph‑RAG knowledge graph with Python graph libraries:

import json
import networkx as nx

with open("graph_export.json") as f:
    data = json.load(f)

G = nx.DiGraph()
for node in data["nodes"]:
    G.add_node(node["id"], **node["properties"])

for rel in data["relationships"]:
    G.add_edge(rel["source"], rel["target"], type=rel["type"])

print(f"Loaded {G.number_of_nodes()} nodes, {G.number_of_edges()} edges")

Key Source Files in Code‑Graph‑RAG

File Purpose
codebase_rag/main.py Core export helpers (_write_graph_json, export_graph_to_file)
codebase_rag/graph_cli.py CLI command group (cgr graph export …)
codebase_rag/services/graph_service.py MemgraphIngestor.export_graph_to_dict() implementation
codebase_rag/types_defs.py GraphData Pydantic schema for serialization
examples/graph_export_example.py End‑to‑end runnable example

Summary

  • Two export paths: Python API via export_graph_to_file() or CLI via cgr graph export.
  • Standardized output: JSON matching the GraphData schema with metadata, nodes, and relationships keys.
  • Configurable batching: Tune batch_size in connect_memgraph() for memory‑constrained environments.
  • Tool‑agnostic: Consume exports in NetworkX, Neo4j, visualization suites, or custom analytics.

Frequently Asked Questions

What graph database does Code‑Graph‑RAG use?

Code‑Graph‑RAG uses Memgraph as its primary graph storage engine. The MemgraphIngestor class in codebase_rag/services/graph_service.py handles all Cypher query execution and connection pooling.

Can I export partial graphs or filter by node type?

The current implementation in export_graph_to_dict() performs a full‑graph dump. For filtered exports, you would need to extend MemgraphIngestor with Cypher WHERE clauses or post‑process the resulting JSON.

Is the exported JSON compatible with Neo4j?

Yes. The nodes/relationships structure maps cleanly to Neo4j's property graph model. Use neo4j-admin import or the APOC apoc.import.json procedure with minor field renaming.

How do I automate exports on a schedule?

Combine the CLI with cron or GitHub Actions. Ensure your Memgraph instance is accessible from the automation environment and that codebase_rag/constants.py contains valid connection parameters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →