# How Semantica Handles Polyglot Storage for Knowledge Graphs: A Technical Deep Dive

> Discover how Semantica manages polyglot storage for knowledge graphs. Learn about its unified dictionary representation and pluggable exporters for various formats and vector embedding persistence.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: deep-dive
- Published: 2026-09-09

---

**Semantica implements polyglot storage for knowledge graphs by representing all graphs as a unified Python dictionary internally, then providing pluggable exporters for formats like JSON-LD, GraphML, Cypher, RDF/Turtle, and Parquet, while maintaining a separate VectorStore abstraction for embedding persistence.**

The `semantica-agi/semantica` repository provides a flexible framework for building and querying knowledge graphs without locking users into a single storage backend. By treating the knowledge graph as a format-agnostic data structure and supporting multiple export formats, Semantica enables true polyglot storage that adapts to diverse downstream requirements.

## The Unified Knowledge Graph Model

At the core of Semantica’s polyglot architecture is a lightweight, normalized representation of the knowledge graph. All graphs are stored internally as a plain Python dictionary with three top-level keys: `entities`, `relationships`, and optional `metadata`. This unified model lives in [`semantica/kgraph/builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kgraph/builder.py), where the `KnowledgeGraphBuilder` class constructs the dictionary from raw nodes and edges.

This design decouples the in-memory representation from any specific storage engine. Because the graph exists as a standard Python object before persistence, the system can serialize it to any target format without requiring backend-specific data models.

## Native Support for Multiple Graph Formats

The polyglot capability centers on the export layer defined in [`semantica/visualization/export_formats.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/export_formats.py). This module contains transformation logic that converts the internal dictionary into a wide variety of industry-standard formats.

### Core Export Formats

According to the source code, the following exporters are available:

- **JSON / JSON-LD** – Generic interchange formats for web APIs and semantic web pipelines (`export_json`, `export_json_ld`).
- **GraphML** – Property-graph format compatible with tools like Gephi and yEd (`export_graphml`).
- **Cypher** – Neo4j-compatible CREATE statements for direct import into graph databases (`export_cypher`).
- **RDF/Turtle** – Semantic web triples for OWL/RDF-based reasoning (`export_turtle`).
- **Parquet / CSV** – Columnar storage for large-scale analytics and data science workflows (`export_parquet`, `export_csv`).

### The Export Registry

All exporters are registered in [`semantica/visualization/export_methods.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/export_methods.py). This registry maps format names to their corresponding functions, allowing the system to dynamically invoke the correct serializer based on runtime configuration. Adding a new storage backend requires only implementing a conversion function and registering it in this file—no changes to core graph logic are necessary.

## Decoupling Vector Storage from Graph Storage

Semantica separates vector embedding storage from the knowledge graph itself through the `VectorStore` abstraction. The [`semantica/vector_store/registry.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/registry.py) file defines a central registry for backend implementations, including:

- **FAISS** ([`faiss_store.py`](https://github.com/semantica-agi/semantica/blob/main/faiss_store.py))
- **Qdrant** ([`qdrant_store.py`](https://github.com/semantica-agi/semantica/blob/main/qdrant_store.py))
- **Weaviate** ([`weaviate_store.py`](https://github.com/semantica-agi/semantica/blob/main/weaviate_store.py))
- **pgvector** ([`pgvector_store.py`](https://github.com/semantica-agi/semantica/blob/main/pgvector_store.py))
- **SQLite-vec** ([`sqlite_vec_store.py`](https://github.com/semantica-agi/semantica/blob/main/sqlite_vec_store.py))
- **Milvus** ([`milvus_store.py`](https://github.com/semantica-agi/semantica/blob/main/milvus_store.py))
- **Pinecone** ([`pinecone_store.py`](https://github.com/semantica-agi/semantica/blob/main/pinecone_store.py))

This architecture allows the knowledge graph to be persisted in graph-optimized formats (like GraphML or Cypher) while embeddings reside in high-performance vector databases. The loose coupling means you can store your graph in Neo4j (via Cypher export) while running similarity search in Qdrant, achieving true polyglot persistence.

## Context Retrieval with Optional Knowledge Graph Integration

The retrieval pipeline demonstrates how polyglot storage integrates with runtime operations. In [`semantica/vector_store/context_retriever.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/context_retriever.py), the `ContextRetriever` class accepts an optional `knowledge_graph` parameter during initialization.

When a knowledge graph is supplied, the retriever enables KG-specific algorithms such as graph-based reranking and entity-aware filtering. If `knowledge_graph=None`, the pipeline falls back to pure vector search. This conditional design illustrates how Semantica accommodates both graph-augmented and vector-only workflows without mandating a specific storage backend.

## Normalization and Visualization

Before rendering, the knowledge graph undergoes normalization via the `_convert_knowledge_graph` method in [`semantica/visualization/kg_visualizer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/kg_visualizer.py). This ensures consistent entity and relationship schemas regardless of the source storage format. The normalization logic is validated by unit tests in [`tests/visualization/test_kg_visualizer_normalize_graph.py`](https://github.com/semantica-agi/semantica/blob/main/tests/visualization/test_kg_visualizer_normalize_graph.py), which confirm deterministic conversion across diverse input structures.

The visualization layer consumes the normalized dictionary and produces interactive dashboards using Plotly, operating independently of whether the underlying data came from Parquet, RDF, or a live graph database.

## Practical Implementation: Building and Exporting Knowledge Graphs

The following example demonstrates the complete workflow: building a knowledge graph, exporting it to multiple polyglot formats, and attaching it to a vector retrieval pipeline.

```python

# 1️⃣ Build a knowledge graph from raw entities & relationships

from semantica.kgraph.builder import KnowledgeGraphBuilder

entities = [
    {"id": "e1", "type": "Person", "name": "Alice"},
    {"id": "e2", "type": "Company", "name": "Acme Corp"},
]
relationships = [
    {"source": "e1", "target": "e2", "type": "EMPLOYED_BY"},
]

kg = KnowledgeGraphBuilder().build(entities + relationships)

# 2️⃣ Export the KG to several formats in one call

from semantica.visualization.export_methods import export_knowledge_graph

export_knowledge_graph(
    kg,
    output_path="out/",
    formats=["json", "graphml", "turtle", "parquet"],
)

# 3️⃣ Initialise a vector store (e.g., Qdrant) and attach the KG

from semantica.vector_store.qdrant_store import QdrantStore
from semantica.vector_store.context_retriever import ContextRetriever

vec_store = QdrantStore(collection_name="my_collection")
retriever = ContextRetriever(vector_store=vec_store, knowledge_graph=kg)

# 4️⃣ Retrieve context – KG-aware reranking will be applied automatically

results = retriever(query="What company does Alice work for?")
print(results)

```

## Summary

- **Unified Representation**: Semantica stores all knowledge graphs as a standard Python dictionary ([`semantica/kgraph/builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kgraph/builder.py)), abstracting away storage specifics.
- **Polyglot Exports**: The framework supports JSON-LD, GraphML, Cypher, RDF/Turtle, Parquet, and CSV through pluggable exporters ([`export_formats.py`](https://github.com/semantica-agi/semantica/blob/main/export_formats.py)).
- **Decoupled Storage**: Vector embeddings and graph structures are stored independently, with support for FAISS, Qdrant, Weaviate, pgvector, and other backends ([`vector_store/registry.py`](https://github.com/semantica-agi/semantica/blob/main/vector_store/registry.py)).
- **Runtime Flexibility**: The `ContextRetriever` optionally accepts a knowledge graph ([`context_retriever.py`](https://github.com/semantica-agi/semantica/blob/main/context_retriever.py)), enabling graph-aware retrieval without mandating graph storage.
- **Extensible Design**: New storage formats are added by registering exporters in [`export_methods.py`](https://github.com/semantica-agi/semantica/blob/main/export_methods.py), while new vector backends implement the `VectorStore` interface.

## Frequently Asked Questions

### What is polyglot storage in the context of knowledge graphs?

Polyglot storage refers to the ability to persist the same knowledge graph in multiple backend formats simultaneously—such as graph databases, RDF triple stores, columnar files, and document stores—without modifying application logic. Semantica achieves this by maintaining a format-agnostic internal representation and providing serializers for each target storage engine.

### Which export formats does Semantica support?

According to the source code in [`semantica/visualization/export_formats.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/export_formats.py), Semantica natively exports to JSON, JSON-LD, GraphML, Cypher (for Neo4j), RDF/Turtle, Parquet, and CSV. Each format targets specific use cases: GraphML for visualization tools, Cypher for graph databases, Turtle for semantic web applications, and Parquet for analytical processing.

### How does Semantica separate vector storage from graph storage?

The framework uses distinct abstractions for each concern. Vector embeddings are managed through the `VectorStore` interface ([`semantica/vector_store/registry.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/registry.py)), with implementations for FAISS, Qdrant, Weaviate, and others. The knowledge graph itself exists as a separate Python dictionary that is serialized via the export utilities. This separation allows users to store embeddings in high-performance vector databases while keeping the graph topology in graph-optimized formats.

### Can I use Semantica with Neo4j?

Yes. By exporting your knowledge graph to Cypher format using `export_cypher` from [`semantica/visualization/export_formats.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/export_formats.py), you generate Neo4j-compatible CREATE statements. You can then import these statements into a Neo4j instance while continuing to use a separate vector store (like Qdrant or Pinecone) for embedding-based retrieval, creating a hybrid polyglot architecture.