How to Export the Knowledge Graph to Cypher or Other Formats in code-graph-rag
Export a code-graph-rag knowledge graph using Cypher queries for direct database transfer or JSON serialization for portable storage and re-loading.
The vitali87/code-graph-rag repository stores parsed codebases as property graphs in Memgraph. Two built-in export mechanisms let you extract this data: Cypher-based export for replaying into other graph databases, and JSON-based export for archival, versioning, or offline analysis. This guide covers both approaches with working code examples drawn directly from the source.
Cypher-Based Export: Raw Node and Relationship Queries
For direct database-to-database transfer, use the predefined Cypher queries in codebase_rag/cypher_queries.py. These queries return complete, unfiltered dumps of your graph structure.
The Export Queries
Two constants capture the full graph state:
CYPHER_EXPORT_NODES— returns every node's internal ID, labels, and propertiesCYPHER_EXPORT_RELATIONSHIPS— returns every edge's source ID, target ID, relationship type, and properties
Both queries use Memgraph's id() function to generate stable numeric identifiers that correlate nodes with their relationships.
# codebase_rag/cypher_queries.py (lines 48-55)
CYPHER_EXPORT_NODES = """
MATCH (n)
RETURN id(n) AS node_id, labels(n) AS labels, properties(n) AS properties
"""
CYPHER_EXPORT_RELATIONSHIPS = """
MATCH ()-[r]->()
RETURN id(startNode(r)) AS source_id,
id(endNode(r)) AS target_id,
type(r) AS rel_type,
properties(r) AS properties
"""
Executing Cypher Exports
Run these queries through any Cypher client or the repository's MemgraphIngestor:
from codebase_rag.main import connect_memgraph
from codebase_rag.cypher_queries import CYPHER_EXPORT_NODES, CYPHER_EXPORT_RELATIONSHIPS
ingestor = connect_memgraph(batch_size=5000)
# Pull complete node set
nodes = ingestor.fetch_all(CYPHER_EXPORT_NODES)
# Pull complete relationship set
relationships = ingestor.fetch_all(CYPHER_EXPORT_RELATIONSHIPS)
print(f"Exported {len(nodes)} nodes and {len(relationships)} relationships")
The fetch_all method executes the query and returns results as a list of dictionaries. This format works with Neo4j, Memgraph, or any Cypher-compatible system.
JSON-Based Export: Portable Graph Serialization
For archiving, version control, or offline processing, use the high-level export_graph_to_file utility in codebase_rag/main.py.
The Export Flow
The JSON export process involves three layers:
MemgraphIngestor.export_graph_to_dict()— queries live Memgraph and builds a Python dictionaryexport_graph_to_file()— handles file I/O and metadata injectioncodebase_rag/graph_loader.py— providesload_graph()for re-hydration
Exporting to JSON File
from pathlib import Path
from codebase_rag.main import connect_memgraph, export_graph_to_file
# Establish database connection
ingestor = connect_memgraph(batch_size=5000)
# Export to JSON with automatic summary
output_path = Path("exports/my_graph.json")
success = export_graph_to_file(ingestor, str(output_path))
# Output includes node/relationship counts and timestamp
if success:
print(f"Graph exported: {output_path.resolve()}")
The resulting JSON follows the schema defined in codec/schema.proto and contains:
nodes: array of node objects with IDs, labels, and propertiesrelationships: array of edge objects with source/target mappingmetadata: export timestamp, total counts, and version info
Loading Exported Graphs
Re-hydrate a JSON export using load_graph:
from codebase_rag.graph_loader import load_graph
from codebase_rag.constants import NodeLabel
# Reload serialized graph
graph = load_graph("exports/my_graph.json")
# Query by label, property, or relationship type
functions = graph.find_nodes_by_label(NodeLabel.FUNCTION)
print(f"Graph contains {len(functions)} function definitions")
Example: Complete Export and Analysis Workflow
The repository includes examples/graph_export_example.py demonstrating real-world usage. Key operations from lines [66-70]:
# From examples/graph_export_example.py
graph = load_graph("graph_export.json")
# Print structural summary
print(f"Total nodes: {len(graph.nodes)}")
print(f"Total edges: {len(graph.relationships)}")
# Sample specific node types for inspection
sample_funcs = list(graph.find_nodes_by_label(NodeLabel.FUNCTION))[:5]
for func in sample_funcs:
print(f" - {func.get('name')} in {func.get('file_path')}")
Key Source Files and Implementation Details
| Component | File | Line Reference | Purpose |
|---|---|---|---|
| JSON export orchestration | codebase_rag/main.py |
[1403-1405] | export_graph_to_file() wrapper |
| Cypher query definitions | codebase_rag/cypher_queries.py |
[48-55] | CYPHER_EXPORT_NODES, CYPHER_EXPORT_RELATIONSHIPS |
| Core export logic | codebase_rag/services/graph_service.py |
Full class | MemgraphIngestor.export_graph_to_dict() |
| JSON re-loading | codebase_rag/graph_loader.py |
Full module | load_graph() and GraphLoader class |
| Working example | examples/graph_export_example.py |
[66-70] | End-to-end export analysis |
Performance and Storage Considerations
- Cypher export streams results directly from Memgraph; memory usage depends on result set size
- JSON export builds complete in-memory representation before serialization—use for graphs under ~1M elements
- Batch sizing in
connect_memgraph(batch_size)affects transaction boundaries but not export completeness - The
id(n)values from Cypher exports are database-specific; re-import to different Memgraph instances generates new internal IDs
Summary
- Use Cypher queries (
CYPHER_EXPORT_NODES,CYPHER_EXPORT_RELATIONSHIPS) for direct database replication and tool interoperability - Use
export_graph_to_filefor portable, versionable JSON archives with embedded metadata - Re-load JSON exports via
load_graphto restore a queryableGraphLoaderinstance without database connection - All export paths rely on
MemgraphIngestorincodebase_rag/services/graph_service.pyfor database communication
Frequently Asked Questions
Can I export directly to Neo4j-compatible Cypher?
Yes. The CYPHER_EXPORT_NODES and CYPHER_EXPORT_RELATIONSHIPS queries use standard openCypher syntax. Results can be fed into Neo4j's UNWIND batch operations or neo4j-admin import after minor syntax adjustment for CREATE/MERGE statements.
Does the JSON export include embeddings or vector data?
Yes. Node properties captured by export_graph_to_dict() include all stored attributes—code text, parsed metadata, and any vector embeddings computed during ingestion. The export is schema-agnostic and captures the complete property set.
How do I automate exports on a schedule?
Import export_graph_to_file and connect_memgraph into your scheduling script. The functions accept string paths and require no interactive input, making them compatible with cron, GitHub Actions, or Airflow workflows.
What limits exist on graph size for JSON export?
The JSON export loads the entire graph into memory as a Python dictionary before serialization. For graphs exceeding available RAM, use the Cypher-based approach with pagination (modify the queries with SKIP/LIMIT) or stream directly to a line-delimited JSON format.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →