How to Export Knowledge Graphs to RDF, Parquet, and Cypher Formats in Semantica

Semantica provides a modular export framework with dedicated exporter classes—RDFExporter, ParquetExporter, and LPGExporter—that serialize knowledge graphs into Turtle, columnar Parquet files, or Cypher scripts for Neo4j ingestion.

Exporting knowledge graphs to RDF, Parquet, and Cypher formats is a core capability of the Semantica AGI framework, enabling seamless interoperability with semantic web tools, analytics pipelines, and graph databases. The library implements a clean separation of concerns where format-specific logic is isolated in dedicated exporter classes while a unified dispatch system handles routing. Whether you need to publish linked data, feed a data lake, or populate a Neo4j instance, Semantica's export architecture provides validated, production-ready serializers.

Export Architecture Overview

The export system centers on three primary exporter classes located in the semantica/export/ module. Each class encapsulates validation, serialization, and I/O logic for its respective format.

  • RDFExporter (semantica/export/rdf_exporter.py, lines 910-965): Handles namespace management, RDF validation, and serialization to Turtle, RDF/XML, JSON-LD, N-Triples, or N3.
  • ParquetExporter (semantica/export/parquet_exporter.py, lines 93-115): Manages Apache Arrow schemas for entities and relationships with support for compression codecs.
  • LPGExporter (semantica/export/lpg_exporter.py, lines 32-50): Generates Cypher CREATE statements for labeled property graphs along with optional indexes and constraints.

All exporters integrate with the MCP tool export_graph defined in mcp/tools/export.py (lines 24-88), which examines the format argument and dispatches to the appropriate serializer. For programmatic use, the high-level convenience function export_knowledge_graph in semantica/export/methods.py (lines 865-902) abstracts this dispatcher.

Exporting to RDF Format

The RDFExporter class converts knowledge graphs into standard RDF serializations. According to the source code in semantica/export/rdf_exporter.py, the class manages namespaces through a NamespaceManager and validates output via RDFValidator before serialization.

The exporter supports multiple RDF syntaxes:

  • Turtle (turtle)
  • RDF/XML (xml)
  • JSON-LD (json-ld)
  • N-Triples (nt)
  • Notation3 (n3)
from semantica.export import RDFExporter

kg = {
    "entities": [
        {"id": "E1", "type": "Person", "name": "Alice"},
        {"id": "E2", "type": "City", "name": "Wonderland"},
    ],
    "relationships": [
        {"source_id": "E1", "target_id": "E2", "type": "LIVES_IN"}
    ],
}

rdf_exporter = RDFExporter()
turtle_str = rdf_exporter.export_to_rdf(kg, format="turtle")
print(turtle_str)

The method returns a string representation that can be persisted to disk or served via API. The exporter automatically handles URI generation and prefix management through the internal NamespaceManager.

Exporting to Parquet Format

For analytics workflows, ParquetExporter provides columnar storage optimized for Arrow-compatible query engines. Located in semantica/export/parquet_exporter.py, this class defines explicit Arrow schemas for both entity and relationship tables.

Key capabilities include:

  • Compression codecs: Supports snappy, gzip, and other standard algorithms.
  • Multi-file exports: Generates separate entities.parquet and relationships.parquet files within a specified directory.
  • Schema enforcement: Validates KG structure against Arrow schemas before serialization.
from semantica.export import ParquetExporter

parquet_exporter = ParquetExporter(compression="gzip")
parquet_exporter.export_knowledge_graph(kg, "kg_parquet")

This creates kg_parquet/entities.parquet and kg_parquet/relationships.parquet, ready for ingestion into Pandas, Polars, or cloud data warehouses.

Exporting to Cypher Format

The LPGExporter class (Labeled Property Graph) targets Neo4j, Memgraph, and compatible graph databases. As implemented in semantica/export/lpg_exporter.py, it generates Cypher CREATE statements and optionally emits schema constraints.

The exporter sanitizes property keys and handles batching through the batch_size parameter to prevent memory issues with large graphs. It outputs a .cypher script file that can be executed directly via the Neo4j Browser or cypher-shell.

from semantica.export import LPGExporter

lpg_exporter = LPGExporter(batch_size=500)
lpg_exporter.export_knowledge_graph(kg, "kg.cypher")

The resulting file contains CREATE statements for nodes and relationships, plus optional indexes defined during the export process.

Using the Unified Export API

For most use cases, the export_knowledge_graph function in semantica/export/methods.py provides the simplest interface. This high-level dispatcher automatically instantiates the correct exporter based on the format parameter.

from semantica.export.methods import export_knowledge_graph

# RDF Turtle

export_knowledge_graph(kg, "my_graph.ttl", format="turtle")

# Parquet directory

export_knowledge_graph(kg, "my_graph_parquet", format="parquet")

# Cypher script

export_knowledge_graph(kg, "my_graph.cypher", format="cypher")

The function handles progress tracking via get_progress_tracker() and validates inputs before delegation, ensuring consistent error handling across all formats.

Summary

  • Semantica implements format-specific exporters in semantica/export/rdf_exporter.py, parquet_exporter.py, and lpg_exporter.py.
  • RDFExporter supports Turtle, RDF/XML, JSON-LD, N-Triples, and N3 with integrated namespace management.
  • ParquetExporter generates Apache Arrow-compliant columnar files with configurable compression for analytics pipelines.
  • LPGExporter produces executable Cypher scripts for Neo4j and Memgraph, including batching and constraint generation.
  • The export_knowledge_graph function in semantica/export/methods.py provides a unified entry point that routes to the appropriate exporter based on the format argument.
  • All exporters integrate with the MCP tool export_graph for standardized tool calling interfaces.

Frequently Asked Questions

What RDF serializations does Semantica support?

Semantica supports Turtle, RDF/XML, JSON-LD, N-Triples, and Notation3 through the RDFExporter class. The export_to_rdf method accepts a format parameter that specifies the desired syntax, returning a string representation that can be written to disk or transmitted via API.

Can I control compression when exporting to Parquet?

Yes. The ParquetExporter constructor accepts a compression parameter that supports standard codecs including snappy and gzip. The exporter writes separate Parquet files for entities and relationships within the specified output directory, enforcing Arrow schemas for type safety.

How do I import the generated Cypher file into Neo4j?

The LPGExporter generates a plain-text .cypher script containing CREATE statements for nodes and relationships. Execute this file using the Neo4j Browser, cypher-shell, or automated deployment pipelines. The exporter optionally includes indexes and constraints to optimize import performance and data integrity.

Is there a way to export knowledge graphs through the MCP interface?

Yes. The export_graph tool in mcp/tools/export.py provides the MCP entry point. It examines the format argument to instantiate the appropriate exporter class (RDFExporter, ParquetExporter, or LPGExporter) and returns either a string payload for text formats or a file path for binary outputs like Parquet.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →