How Semantica Handles Polyglot Storage for Knowledge Graphs: A Technical Deep Dive
Semantica implements polyglot storage for knowledge graphs by representing all graphs as a unified Python dictionary internally, then providing pluggable exporters for formats like JSON-LD, GraphML, Cypher, RDF/Turtle, and Parquet, while maintaining a separate VectorStore abstraction for embedding persistence.
The semantica-agi/semantica repository provides a flexible framework for building and querying knowledge graphs without locking users into a single storage backend. By treating the knowledge graph as a format-agnostic data structure and supporting multiple export formats, Semantica enables true polyglot storage that adapts to diverse downstream requirements.
The Unified Knowledge Graph Model
At the core of Semantica’s polyglot architecture is a lightweight, normalized representation of the knowledge graph. All graphs are stored internally as a plain Python dictionary with three top-level keys: entities, relationships, and optional metadata. This unified model lives in semantica/kgraph/builder.py, where the KnowledgeGraphBuilder class constructs the dictionary from raw nodes and edges.
This design decouples the in-memory representation from any specific storage engine. Because the graph exists as a standard Python object before persistence, the system can serialize it to any target format without requiring backend-specific data models.
Native Support for Multiple Graph Formats
The polyglot capability centers on the export layer defined in semantica/visualization/export_formats.py. This module contains transformation logic that converts the internal dictionary into a wide variety of industry-standard formats.
Core Export Formats
According to the source code, the following exporters are available:
- JSON / JSON-LD – Generic interchange formats for web APIs and semantic web pipelines (
export_json,export_json_ld). - GraphML – Property-graph format compatible with tools like Gephi and yEd (
export_graphml). - Cypher – Neo4j-compatible CREATE statements for direct import into graph databases (
export_cypher). - RDF/Turtle – Semantic web triples for OWL/RDF-based reasoning (
export_turtle). - Parquet / CSV – Columnar storage for large-scale analytics and data science workflows (
export_parquet,export_csv).
The Export Registry
All exporters are registered in semantica/visualization/export_methods.py. This registry maps format names to their corresponding functions, allowing the system to dynamically invoke the correct serializer based on runtime configuration. Adding a new storage backend requires only implementing a conversion function and registering it in this file—no changes to core graph logic are necessary.
Decoupling Vector Storage from Graph Storage
Semantica separates vector embedding storage from the knowledge graph itself through the VectorStore abstraction. The semantica/vector_store/registry.py file defines a central registry for backend implementations, including:
- FAISS (
faiss_store.py) - Qdrant (
qdrant_store.py) - Weaviate (
weaviate_store.py) - pgvector (
pgvector_store.py) - SQLite-vec (
sqlite_vec_store.py) - Milvus (
milvus_store.py) - Pinecone (
pinecone_store.py)
This architecture allows the knowledge graph to be persisted in graph-optimized formats (like GraphML or Cypher) while embeddings reside in high-performance vector databases. The loose coupling means you can store your graph in Neo4j (via Cypher export) while running similarity search in Qdrant, achieving true polyglot persistence.
Context Retrieval with Optional Knowledge Graph Integration
The retrieval pipeline demonstrates how polyglot storage integrates with runtime operations. In semantica/vector_store/context_retriever.py, the ContextRetriever class accepts an optional knowledge_graph parameter during initialization.
When a knowledge graph is supplied, the retriever enables KG-specific algorithms such as graph-based reranking and entity-aware filtering. If knowledge_graph=None, the pipeline falls back to pure vector search. This conditional design illustrates how Semantica accommodates both graph-augmented and vector-only workflows without mandating a specific storage backend.
Normalization and Visualization
Before rendering, the knowledge graph undergoes normalization via the _convert_knowledge_graph method in semantica/visualization/kg_visualizer.py. This ensures consistent entity and relationship schemas regardless of the source storage format. The normalization logic is validated by unit tests in tests/visualization/test_kg_visualizer_normalize_graph.py, which confirm deterministic conversion across diverse input structures.
The visualization layer consumes the normalized dictionary and produces interactive dashboards using Plotly, operating independently of whether the underlying data came from Parquet, RDF, or a live graph database.
Practical Implementation: Building and Exporting Knowledge Graphs
The following example demonstrates the complete workflow: building a knowledge graph, exporting it to multiple polyglot formats, and attaching it to a vector retrieval pipeline.
# 1️⃣ Build a knowledge graph from raw entities & relationships
from semantica.kgraph.builder import KnowledgeGraphBuilder
entities = [
{"id": "e1", "type": "Person", "name": "Alice"},
{"id": "e2", "type": "Company", "name": "Acme Corp"},
]
relationships = [
{"source": "e1", "target": "e2", "type": "EMPLOYED_BY"},
]
kg = KnowledgeGraphBuilder().build(entities + relationships)
# 2️⃣ Export the KG to several formats in one call
from semantica.visualization.export_methods import export_knowledge_graph
export_knowledge_graph(
kg,
output_path="out/",
formats=["json", "graphml", "turtle", "parquet"],
)
# 3️⃣ Initialise a vector store (e.g., Qdrant) and attach the KG
from semantica.vector_store.qdrant_store import QdrantStore
from semantica.vector_store.context_retriever import ContextRetriever
vec_store = QdrantStore(collection_name="my_collection")
retriever = ContextRetriever(vector_store=vec_store, knowledge_graph=kg)
# 4️⃣ Retrieve context – KG-aware reranking will be applied automatically
results = retriever(query="What company does Alice work for?")
print(results)
Summary
- Unified Representation: Semantica stores all knowledge graphs as a standard Python dictionary (
semantica/kgraph/builder.py), abstracting away storage specifics. - Polyglot Exports: The framework supports JSON-LD, GraphML, Cypher, RDF/Turtle, Parquet, and CSV through pluggable exporters (
export_formats.py). - Decoupled Storage: Vector embeddings and graph structures are stored independently, with support for FAISS, Qdrant, Weaviate, pgvector, and other backends (
vector_store/registry.py). - Runtime Flexibility: The
ContextRetrieveroptionally accepts a knowledge graph (context_retriever.py), enabling graph-aware retrieval without mandating graph storage. - Extensible Design: New storage formats are added by registering exporters in
export_methods.py, while new vector backends implement theVectorStoreinterface.
Frequently Asked Questions
What is polyglot storage in the context of knowledge graphs?
Polyglot storage refers to the ability to persist the same knowledge graph in multiple backend formats simultaneously—such as graph databases, RDF triple stores, columnar files, and document stores—without modifying application logic. Semantica achieves this by maintaining a format-agnostic internal representation and providing serializers for each target storage engine.
Which export formats does Semantica support?
According to the source code in semantica/visualization/export_formats.py, Semantica natively exports to JSON, JSON-LD, GraphML, Cypher (for Neo4j), RDF/Turtle, Parquet, and CSV. Each format targets specific use cases: GraphML for visualization tools, Cypher for graph databases, Turtle for semantic web applications, and Parquet for analytical processing.
How does Semantica separate vector storage from graph storage?
The framework uses distinct abstractions for each concern. Vector embeddings are managed through the VectorStore interface (semantica/vector_store/registry.py), with implementations for FAISS, Qdrant, Weaviate, and others. The knowledge graph itself exists as a separate Python dictionary that is serialized via the export utilities. This separation allows users to store embeddings in high-performance vector databases while keeping the graph topology in graph-optimized formats.
Can I use Semantica with Neo4j?
Yes. By exporting your knowledge graph to Cypher format using export_cypher from semantica/visualization/export_formats.py, you generate Neo4j-compatible CREATE statements. You can then import these statements into a Neo4j instance while continuing to use a separate vector store (like Qdrant or Pinecone) for embedding-based retrieval, creating a hybrid polyglot architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →