How to Configure Different Vector Backends (FAISS, Qdrant, etc.) in Semantica
Semantica configures vector backends through the VectorStoreConfig singleton and the VectorStore façade, allowing you to switch between FAISS, Qdrant, Weaviate, and others via environment variables, constructor arguments, or runtime overrides.
The semantica-agi/semantica repository provides a backend-agnostic vector storage subsystem that abstracts index operations behind a unified interface. Whether you need local development with FAISS or production-scale retrieval with Qdrant or Pinecone, you can configure different vector backends without modifying application logic. This flexibility is achieved through a configuration-driven architecture that delegates all vector operations to protocol-compliant backend implementations.
Architecture of the Vector Store System
Semantica's vector storage architecture separates configuration, abstraction, and implementation into three distinct layers. Understanding these layers helps you configure backends effectively across different deployment scenarios.
Configuration Layer: VectorStoreConfig
The VectorStoreConfig class in semantica/vector_store/config.py serves as the central configuration manager. It reads environment variables such as VECTOR_STORE_DEFAULT_BACKEND and VECTOR_STORE_QDRANT_URL, merges them with optional configuration files, and provides sensible defaults (typically faiss). This singleton pattern ensures consistent backend selection across the entire process unless explicitly overridden.
Abstraction Layer: The VectorStore Façade
The VectorStore class in semantica/vector_store/vector_store.py exposes a uniform API including methods like add_vectors, search, count, and filter_by_metadata. When instantiated, it queries VectorStoreConfig for the default backend or accepts a backend= argument, then lazily initializes the corresponding concrete implementation. All public methods delegate to this internal backend instance, allowing the rest of Semantica—including knowledge graph pipelines and CLI utilities—to operate agnostically.
Implementation Layer: Backend Protocols
Each supported backend resides in its own module and implements the VectorBackend protocol. Concrete implementations include FAISSStore in semantica/vector_store/faiss_store.py, QdrantStore in semantica/vector_store/qdrant_store.py, and similar classes for Weaviate, Pinecone, Milvus, SQLite-Vec, and PgVector. These classes handle backend-specific connection management, index types, and search algorithms while presenting a consistent interface to the façade.
Configuration Methods
You can configure vector backends at three levels: globally via environment variables, per-instance via constructor arguments, or dynamically at runtime.
Global Default via Environment Variables
Set process-wide defaults before importing Semantica components. The system variable VECTOR_STORE_DEFAULT_BACKEND controls which backend initializes when no explicit argument is provided.
export VECTOR_STORE_DEFAULT_BACKEND=qdrant
export VECTOR_STORE_QDRANT_URL="http://localhost:6333"
from semantica.vector_store.vector_store import VectorStore
# Automatically uses Qdrant based on environment configuration
store = VectorStore(dimension=384)
store.add_vectors(vectors)
results = store.search(query_vector, top_k=5)
Per-Instance Backend Selection
Override the global default for specific store instances using the backend parameter. This is useful when different components of your application require different storage characteristics.
from semantica.vector_store.vector_store import VectorStore
# Explicitly request FAISS for this instance only
faiss_store = VectorStore(
backend="faiss",
dimension=768,
config={"faiss_index_type": "IVF"}
)
Backend-Specific Configuration Options
Pass connection parameters, authentication credentials, and index settings through the config dictionary. Each backend recognizes its own set of keys.
# Qdrant configuration with custom collection
qdrant_store = VectorStore(
backend="qdrant",
dimension=256,
config={
"qdrant_url": "http://qdrant:6333",
"collection_name": "my_vectors"
}
)
# Pinecone configuration with API authentication
pinecone_store = VectorStore(
backend="pinecone",
dimension=512,
config={
"pinecone_api_key": "YOUR_KEY",
"index_name": "my-index"
}
)
Practical Implementation Examples
Using LangChain Integration
The LangChain wrapper in integrations/langchain/vectorstore.py accepts the same backend arguments, ensuring your integration code remains portable across storage systems.
from semantica.integrations.langchain.vectorstore import SemanticaVectorStore
langchain_store = SemanticaVectorStore(
backend="weaviate",
dimension=256,
config={"weaviate_url": "http://localhost:8080"}
)
retriever = langchain_store.as_retriever(search_kwargs={"k": 4})
Runtime Backend Switching
For advanced scenarios, you can swap backends after instantiation by modifying the backend attribute and forcing re-initialization. This bypasses the typical configuration flow but requires manual state management.
store = VectorStore(dimension=128) # Initially uses default (e.g., FAISS)
store.add_vectors(vectors)
# Switch to Qdrant dynamically
store.backend = "qdrant"
store._backend_store = None # Force re-initialisation
store._ensure_backend() # Internal method to create QdrantStore
store.add_vectors(more_vectors)
Supported Backends and Configuration
Semantica supports multiple production-grade and embedded vector databases:
- FAISS (
faiss): Local, high-performance similarity search with configurable index types (Flat, IVF, HNSW) viafaiss_index_type. - Qdrant (
qdrant): Cloud-native vector database configured viaqdrant_urlandcollection_name. - Weaviate (
weaviate): GraphQL-enabled vector search engine usingweaviate_urland authentication headers. - Pinecone (
pinecone): Managed cloud service requiringpinecone_api_keyandindex_name. - Milvus (
milvus): Distributed vector database for enterprise scale. - SQLite-Vec (
sqlite-vec): Serverless, embedded option for edge deployments. - PgVector (
pgvector): PostgreSQL extension for ACID-compliant vector storage.
Each backend implementation resides in semantica/vector_store/{backend}_store.py and respects the VectorBackend protocol methods: add, search, delete, and count.
Summary
- Centralized configuration in
semantica/vector_store/config.pyuses environment variables likeVECTOR_STORE_DEFAULT_BACKENDto set process-wide defaults. - The
VectorStorefaçade insemantica/vector_store/vector_store.pyprovides a unified API that delegates to backend-specific implementations. - Per-instance overrides via the
backend=argument allow mixed storage strategies within the same application. - Backend-specific settings pass through the
configdictionary, enabling authentication, connection URLs, and index tuning. - LangChain integrations and CLI tools automatically respect these configuration mechanisms without code changes.
Frequently Asked Questions
What is the default vector backend in Semantica?
FAISS is the default backend when no environment variable or constructor argument is specified. The VectorStoreConfig singleton returns "faiss" as the fallback value when VECTOR_STORE_DEFAULT_BACKEND is unset, ensuring immediate functionality without external dependencies.
Can I switch vector backends without changing my application code?
Yes. Set the VECTOR_STORE_DEFAULT_BACKEND environment variable to your preferred backend (e.g., qdrant, weaviate) and provide necessary connection settings. The VectorStore façade automatically instantiates the correct implementation, keeping your application code unchanged.
How do I configure authentication for cloud-based backends like Pinecone?
Pass credentials through the config dictionary. For Pinecone, include "pinecone_api_key" and "index_name" in the config dict when instantiating VectorStore. Similarly, use "qdrant_url" for Qdrant or authentication headers for Weaviate. These values are forwarded directly to the backend constructor in the respective *_store.py module.
Is it possible to use multiple vector backends simultaneously in the same application?
Yes. Instantiate separate VectorStore objects with different backend= arguments. Each instance maintains its own connection to the specified backend, allowing you to use FAISS for local caching and Qdrant for persistent storage within the same process. The global configuration only affects instances where the backend is not explicitly specified.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →