How to Export the Knowledge Graph to Cypher or Other Formats in code-graph-rag

Export a code-graph-rag knowledge graph using Cypher queries for direct database transfer or JSON serialization for portable storage and re-loading.

The vitali87/code-graph-rag repository stores parsed codebases as property graphs in Memgraph. Two built-in export mechanisms let you extract this data: Cypher-based export for replaying into other graph databases, and JSON-based export for archival, versioning, or offline analysis. This guide covers both approaches with working code examples drawn directly from the source.

Cypher-Based Export: Raw Node and Relationship Queries

For direct database-to-database transfer, use the predefined Cypher queries in codebase_rag/cypher_queries.py. These queries return complete, unfiltered dumps of your graph structure.

The Export Queries

Two constants capture the full graph state:

  • CYPHER_EXPORT_NODES — returns every node's internal ID, labels, and properties
  • CYPHER_EXPORT_RELATIONSHIPS — returns every edge's source ID, target ID, relationship type, and properties

Both queries use Memgraph's id() function to generate stable numeric identifiers that correlate nodes with their relationships.


# codebase_rag/cypher_queries.py (lines 48-55)

CYPHER_EXPORT_NODES = """
MATCH (n)
RETURN id(n) AS node_id, labels(n) AS labels, properties(n) AS properties
"""

CYPHER_EXPORT_RELATIONSHIPS = """
MATCH ()-[r]->()
RETURN id(startNode(r)) AS source_id, 
       id(endNode(r)) AS target_id, 
       type(r) AS rel_type, 
       properties(r) AS properties
"""

Executing Cypher Exports

Run these queries through any Cypher client or the repository's MemgraphIngestor:

from codebase_rag.main import connect_memgraph
from codebase_rag.cypher_queries import CYPHER_EXPORT_NODES, CYPHER_EXPORT_RELATIONSHIPS

ingestor = connect_memgraph(batch_size=5000)

# Pull complete node set

nodes = ingestor.fetch_all(CYPHER_EXPORT_NODES)

# Pull complete relationship set

relationships = ingestor.fetch_all(CYPHER_EXPORT_RELATIONSHIPS)

print(f"Exported {len(nodes)} nodes and {len(relationships)} relationships")

The fetch_all method executes the query and returns results as a list of dictionaries. This format works with Neo4j, Memgraph, or any Cypher-compatible system.

JSON-Based Export: Portable Graph Serialization

For archiving, version control, or offline processing, use the high-level export_graph_to_file utility in codebase_rag/main.py.

The Export Flow

The JSON export process involves three layers:

  1. MemgraphIngestor.export_graph_to_dict() — queries live Memgraph and builds a Python dictionary
  2. export_graph_to_file() — handles file I/O and metadata injection
  3. codebase_rag/graph_loader.py — provides load_graph() for re-hydration

Exporting to JSON File

from pathlib import Path
from codebase_rag.main import connect_memgraph, export_graph_to_file

# Establish database connection

ingestor = connect_memgraph(batch_size=5000)

# Export to JSON with automatic summary

output_path = Path("exports/my_graph.json")
success = export_graph_to_file(ingestor, str(output_path))

# Output includes node/relationship counts and timestamp

if success:
    print(f"Graph exported: {output_path.resolve()}")

The resulting JSON follows the schema defined in codec/schema.proto and contains:

  • nodes: array of node objects with IDs, labels, and properties
  • relationships: array of edge objects with source/target mapping
  • metadata: export timestamp, total counts, and version info

Loading Exported Graphs

Re-hydrate a JSON export using load_graph:

from codebase_rag.graph_loader import load_graph
from codebase_rag.constants import NodeLabel

# Reload serialized graph

graph = load_graph("exports/my_graph.json")

# Query by label, property, or relationship type

functions = graph.find_nodes_by_label(NodeLabel.FUNCTION)
print(f"Graph contains {len(functions)} function definitions")

Example: Complete Export and Analysis Workflow

The repository includes examples/graph_export_example.py demonstrating real-world usage. Key operations from lines [66-70]:


# From examples/graph_export_example.py

graph = load_graph("graph_export.json")

# Print structural summary

print(f"Total nodes: {len(graph.nodes)}")
print(f"Total edges: {len(graph.relationships)}")

# Sample specific node types for inspection

sample_funcs = list(graph.find_nodes_by_label(NodeLabel.FUNCTION))[:5]
for func in sample_funcs:
    print(f"  - {func.get('name')} in {func.get('file_path')}")

Key Source Files and Implementation Details

Component File Line Reference Purpose
JSON export orchestration codebase_rag/main.py [1403-1405] export_graph_to_file() wrapper
Cypher query definitions codebase_rag/cypher_queries.py [48-55] CYPHER_EXPORT_NODES, CYPHER_EXPORT_RELATIONSHIPS
Core export logic codebase_rag/services/graph_service.py Full class MemgraphIngestor.export_graph_to_dict()
JSON re-loading codebase_rag/graph_loader.py Full module load_graph() and GraphLoader class
Working example examples/graph_export_example.py [66-70] End-to-end export analysis

Performance and Storage Considerations

  • Cypher export streams results directly from Memgraph; memory usage depends on result set size
  • JSON export builds complete in-memory representation before serialization—use for graphs under ~1M elements
  • Batch sizing in connect_memgraph(batch_size) affects transaction boundaries but not export completeness
  • The id(n) values from Cypher exports are database-specific; re-import to different Memgraph instances generates new internal IDs

Summary

  • Use Cypher queries (CYPHER_EXPORT_NODES, CYPHER_EXPORT_RELATIONSHIPS) for direct database replication and tool interoperability
  • Use export_graph_to_file for portable, versionable JSON archives with embedded metadata
  • Re-load JSON exports via load_graph to restore a queryable GraphLoader instance without database connection
  • All export paths rely on MemgraphIngestor in codebase_rag/services/graph_service.py for database communication

Frequently Asked Questions

Can I export directly to Neo4j-compatible Cypher?

Yes. The CYPHER_EXPORT_NODES and CYPHER_EXPORT_RELATIONSHIPS queries use standard openCypher syntax. Results can be fed into Neo4j's UNWIND batch operations or neo4j-admin import after minor syntax adjustment for CREATE/MERGE statements.

Does the JSON export include embeddings or vector data?

Yes. Node properties captured by export_graph_to_dict() include all stored attributes—code text, parsed metadata, and any vector embeddings computed during ingestion. The export is schema-agnostic and captures the complete property set.

How do I automate exports on a schedule?

Import export_graph_to_file and connect_memgraph into your scheduling script. The functions accept string paths and require no interactive input, making them compatible with cron, GitHub Actions, or Airflow workflows.

What limits exist on graph size for JSON export?

The JSON export loads the entire graph into memory as a Python dictionary before serialization. For graphs exceeding available RAM, use the Cypher-based approach with pagination (modify the queries with SKIP/LIMIT) or stream directly to a line-delimited JSON format.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →