How Semantica Supports Polyglot Storage for Vector and Graph Databases

Semantica enables polyglot storage for vector and graph databases through a unified abstraction layer with pluggable backend registries, allowing seamless swapping of databases without code changes.

The semantica-agi/semantica repository provides an extensible AI orchestration framework designed around polyglot storage for vector and graph databases. Rather than locking developers into a single vendor, the architecture abstracts storage implementations behind uniform Pythonic interfaces. This design allows teams to mix and match vector databases like Qdrant or Weaviate with graph databases like Neo4j or TigerGraph using identical application code.

Unified Store Interfaces

At the core of Semantica’s polyglot capability are abstract base classes that define generic storage contracts. These interfaces operate on standard Python data structures—dictionaries, lists, and np.ndarray objects—ensuring that application logic remains decoupled from vendor-specific SDKs.

VectorStore Abstract Interface

The VectorStore interface, defined in semantica/vector_store/vector_store.py, specifies the contract for all vector database drivers. Concrete implementations must provide the following methods:

  • upsert(ids, vectors, metadata) – Insert or update vectors with associated metadata
  • delete(ids) – Remove vectors by identifier
  • search(query_vector, top_k, filter) – Perform similarity search with optional filtering
  • batch_upsert() – Efficient bulk insertion for large datasets
  • clear() – Remove all vectors from the namespace

GraphStore Abstract Interface

Similarly, the GraphStore interface in semantica/graph_store/graph_store.py abstracts graph database operations:

  • add_node(node_id, labels, properties) – Create nodes with typed labels
  • add_edge(source, target, rel_type, properties) – Create relationships between nodes
  • remove_node(node_id) – Delete nodes and their edges
  • remove_edge(source, target, rel_type) – Remove specific relationships
  • query(cypher_query, parameters) – Execute graph queries
  • stats() – Return database statistics

Backend Registry and Polyglot Capability

Semantica achieves runtime flexibility through a dictionary-based plugin pattern implemented in semantica/vector_store/registry.py and semantica/graph_store/registry.py.

Dynamic Registration Pattern

New backends register themselves using the method_registry.register() function:

from semantica.vector_store.registry import method_registry
from semantica.vector_store.weaviate_store import WeaviateStore

method_registry.register(
    task="store",
    method_name="weaviate",
    method_func=WeaviateStore,
    description="Weaviate vector DB driver"
)

The registry maintains a mapping of {task: {method_name: Callable}}, allowing the system to discover drivers dynamically without modifying core framework code.

Factory Pattern for Runtime Selection

The VectorStoreFactory and GraphStoreFactory classes instantiate the correct backend at runtime based on configuration strings:

from semantica.vector_store import VectorStoreFactory
from semantica.graph_store import GraphStoreFactory

# Instantiate different backends using identical factory interfaces

qdrant_store = VectorStoreFactory.create("qdrant", url="http://localhost:6333")
pinecone_store = VectorStoreFactory.create("pinecone", api_key="...")
neo4j_graph = GraphStoreFactory.create("neo4j", uri="bolt://localhost:7687")

Supported vector backends include Weaviate, Qdrant, Pinecone, Milvus, and PGVector. Supported graph backends include Neo4j and TigerGraph.

Provenance and Metadata Layer

To maintain audit trails across heterogeneous storage systems, Semantica wraps store operations with provenance tracking. The VectorStoreProvenance class in semantica/vector_store/vector_store_provenance.py automatically attaches UUIDs, timestamps, and source tags to every upsert and delete operation. Similarly, GraphStoreProvenance in semantica/graph_store/graph_store_provenance.py records metadata for node and edge creation, including originating decision IDs.

This layer ensures that analytics pipelines receive consistent provenance information regardless of whether vectors reside in Qdrant or graphs in Neo4j.

Practical Polyglot Workflow

The following example demonstrates mixing a Qdrant vector store with a Neo4j graph database:

from semantica.vector_store import VectorStoreFactory
from semantica.graph_store import GraphStoreFactory

# Initialize different backend technologies

vs = VectorStoreFactory.create("qdrant", url="http://localhost:6333")
gs = GraphStoreFactory.create("neo4j", uri="bolt://localhost:7687")

# Store vectors

vs.upsert(
    ids=["doc1"], 
    vectors=[[0.1, 0.2, 0.3]], 
    metadata=[{"title": "Introduction"}]
)

# Link to graph entities

gs.add_node(node_id="doc1", labels=["Document"], properties={"title": "Introduction"})
gs.add_edge(source="doc1", target="topic123", rel_type="MENTIONS")

# Hybrid search (vector + graph context)

results = vs.search(
    query_vector=[0.1, 0.2, 0.3], 
    top_k=5,
    filter={"graph_node": "topic123"}
)

Changing the vector backend to Pinecone or the graph backend to TigerGraph requires only modifying the initialization string passed to the factory.

Extending Semantica with Custom Backends

Adding support for new databases requires three steps:

  1. Inherit from BaseVectorStore or BaseGraphStore in the respective interface files
  2. Implement all required abstract methods (upsert, search for vectors; add_node, query for graphs)
  3. Register the class with the appropriate registry module

Because registration occurs at runtime, custom drivers become immediately available to all existing pipelines without framework modifications.

Summary

  • Semantica provides unified VectorStore and GraphStore interfaces that abstract vendor-specific implementations.
  • The backend registry in registry.py files enables dynamic discovery of drivers using a dictionary-based plugin pattern.
  • Factory classes instantiate concrete backends at runtime, supporting databases like Qdrant, Weaviate, Neo4j, and TigerGraph.
  • Provenance wrappers automatically attach metadata and audit trails across all storage operations.
  • Applications can mix vector and graph technologies freely, swapping backends by changing configuration strings rather than application code.

Frequently Asked Questions

What is polyglot storage in the context of Semantica?

Polyglot storage refers to the ability to use multiple database technologies simultaneously within the same application. In Semantica, this means combining vector databases like Qdrant or Pinecone with graph databases like Neo4j through a single, unified API without vendor lock-in.

Which vector and graph databases does Semantica currently support?

According to the source code, Semantica supports vector databases including Weaviate, Qdrant, Pinecone, Milvus, and PGVector. For graph storage, it supports Neo4j and TigerGraph. The registry pattern allows additional backends to be added without core code changes.

How does the provenance layer work across different storage backends?

The provenance layer uses wrapper classes—VectorStoreProvenance and GraphStoreProvenance—that intercept storage method calls and inject metadata such as UUIDs, timestamps, and source identifiers. This ensures consistent audit trails regardless of whether the underlying store is vector-based or graph-based.

Can I use different vector and graph databases in the same application?

Yes. Semantica’s architecture explicitly supports mixing storage technologies. You can initialize a Qdrant vector store and a Neo4j graph store in the same Python process using their respective factories, link data between them using shared identifiers, and perform hybrid queries that leverage both storage types.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →