How Semantica Supports Polyglot Storage for Vector and Graph Databases
Semantica enables polyglot storage for vector and graph databases through a unified abstraction layer with pluggable backend registries, allowing seamless swapping of databases without code changes.
The semantica-agi/semantica repository provides an extensible AI orchestration framework designed around polyglot storage for vector and graph databases. Rather than locking developers into a single vendor, the architecture abstracts storage implementations behind uniform Pythonic interfaces. This design allows teams to mix and match vector databases like Qdrant or Weaviate with graph databases like Neo4j or TigerGraph using identical application code.
Unified Store Interfaces
At the core of Semantica’s polyglot capability are abstract base classes that define generic storage contracts. These interfaces operate on standard Python data structures—dictionaries, lists, and np.ndarray objects—ensuring that application logic remains decoupled from vendor-specific SDKs.
VectorStore Abstract Interface
The VectorStore interface, defined in semantica/vector_store/vector_store.py, specifies the contract for all vector database drivers. Concrete implementations must provide the following methods:
upsert(ids, vectors, metadata)– Insert or update vectors with associated metadatadelete(ids)– Remove vectors by identifiersearch(query_vector, top_k, filter)– Perform similarity search with optional filteringbatch_upsert()– Efficient bulk insertion for large datasetsclear()– Remove all vectors from the namespace
GraphStore Abstract Interface
Similarly, the GraphStore interface in semantica/graph_store/graph_store.py abstracts graph database operations:
add_node(node_id, labels, properties)– Create nodes with typed labelsadd_edge(source, target, rel_type, properties)– Create relationships between nodesremove_node(node_id)– Delete nodes and their edgesremove_edge(source, target, rel_type)– Remove specific relationshipsquery(cypher_query, parameters)– Execute graph queriesstats()– Return database statistics
Backend Registry and Polyglot Capability
Semantica achieves runtime flexibility through a dictionary-based plugin pattern implemented in semantica/vector_store/registry.py and semantica/graph_store/registry.py.
Dynamic Registration Pattern
New backends register themselves using the method_registry.register() function:
from semantica.vector_store.registry import method_registry
from semantica.vector_store.weaviate_store import WeaviateStore
method_registry.register(
task="store",
method_name="weaviate",
method_func=WeaviateStore,
description="Weaviate vector DB driver"
)
The registry maintains a mapping of {task: {method_name: Callable}}, allowing the system to discover drivers dynamically without modifying core framework code.
Factory Pattern for Runtime Selection
The VectorStoreFactory and GraphStoreFactory classes instantiate the correct backend at runtime based on configuration strings:
from semantica.vector_store import VectorStoreFactory
from semantica.graph_store import GraphStoreFactory
# Instantiate different backends using identical factory interfaces
qdrant_store = VectorStoreFactory.create("qdrant", url="http://localhost:6333")
pinecone_store = VectorStoreFactory.create("pinecone", api_key="...")
neo4j_graph = GraphStoreFactory.create("neo4j", uri="bolt://localhost:7687")
Supported vector backends include Weaviate, Qdrant, Pinecone, Milvus, and PGVector. Supported graph backends include Neo4j and TigerGraph.
Provenance and Metadata Layer
To maintain audit trails across heterogeneous storage systems, Semantica wraps store operations with provenance tracking. The VectorStoreProvenance class in semantica/vector_store/vector_store_provenance.py automatically attaches UUIDs, timestamps, and source tags to every upsert and delete operation. Similarly, GraphStoreProvenance in semantica/graph_store/graph_store_provenance.py records metadata for node and edge creation, including originating decision IDs.
This layer ensures that analytics pipelines receive consistent provenance information regardless of whether vectors reside in Qdrant or graphs in Neo4j.
Practical Polyglot Workflow
The following example demonstrates mixing a Qdrant vector store with a Neo4j graph database:
from semantica.vector_store import VectorStoreFactory
from semantica.graph_store import GraphStoreFactory
# Initialize different backend technologies
vs = VectorStoreFactory.create("qdrant", url="http://localhost:6333")
gs = GraphStoreFactory.create("neo4j", uri="bolt://localhost:7687")
# Store vectors
vs.upsert(
ids=["doc1"],
vectors=[[0.1, 0.2, 0.3]],
metadata=[{"title": "Introduction"}]
)
# Link to graph entities
gs.add_node(node_id="doc1", labels=["Document"], properties={"title": "Introduction"})
gs.add_edge(source="doc1", target="topic123", rel_type="MENTIONS")
# Hybrid search (vector + graph context)
results = vs.search(
query_vector=[0.1, 0.2, 0.3],
top_k=5,
filter={"graph_node": "topic123"}
)
Changing the vector backend to Pinecone or the graph backend to TigerGraph requires only modifying the initialization string passed to the factory.
Extending Semantica with Custom Backends
Adding support for new databases requires three steps:
- Inherit from
BaseVectorStoreorBaseGraphStorein the respective interface files - Implement all required abstract methods (
upsert,searchfor vectors;add_node,queryfor graphs) - Register the class with the appropriate registry module
Because registration occurs at runtime, custom drivers become immediately available to all existing pipelines without framework modifications.
Summary
- Semantica provides unified
VectorStoreandGraphStoreinterfaces that abstract vendor-specific implementations. - The backend registry in
registry.pyfiles enables dynamic discovery of drivers using a dictionary-based plugin pattern. - Factory classes instantiate concrete backends at runtime, supporting databases like Qdrant, Weaviate, Neo4j, and TigerGraph.
- Provenance wrappers automatically attach metadata and audit trails across all storage operations.
- Applications can mix vector and graph technologies freely, swapping backends by changing configuration strings rather than application code.
Frequently Asked Questions
What is polyglot storage in the context of Semantica?
Polyglot storage refers to the ability to use multiple database technologies simultaneously within the same application. In Semantica, this means combining vector databases like Qdrant or Pinecone with graph databases like Neo4j through a single, unified API without vendor lock-in.
Which vector and graph databases does Semantica currently support?
According to the source code, Semantica supports vector databases including Weaviate, Qdrant, Pinecone, Milvus, and PGVector. For graph storage, it supports Neo4j and TigerGraph. The registry pattern allows additional backends to be added without core code changes.
How does the provenance layer work across different storage backends?
The provenance layer uses wrapper classes—VectorStoreProvenance and GraphStoreProvenance—that intercept storage method calls and inject metadata such as UUIDs, timestamps, and source identifiers. This ensures consistent audit trails regardless of whether the underlying store is vector-based or graph-based.
Can I use different vector and graph databases in the same application?
Yes. Semantica’s architecture explicitly supports mixing storage technologies. You can initialize a Qdrant vector store and a Neo4j graph store in the same Python process using their respective factories, link data between them using shared identifiers, and perform hybrid queries that leverage both storage types.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →