How to Integrate Semantica with Existing LLM and Vector Store Investments
Semantica’s modular pipeline architecture allows seamless integration with existing LLM and vector store investments by implementing the ProviderBase class for custom language models and configuring the VectorStore abstraction to connect to external vector databases.
The semantica-agi/semantica repository implements a knowledge-graph enrichment platform designed as a series of composable stages. Because each component—from ingestion to storage—operates behind well-defined interfaces, you can integrate your current AI infrastructure without rebuilding your data pipelines.
Understanding Semantica’s Modular Pipeline
Semantica separates data processing into discrete stages defined in ARCHITECTURE.md. The pipeline flows from Sources through Parse, Extract, KG Construction, and finally to Storage, where the VectorStore interface and graph stores reside.
This design means your existing investments sit at two primary extension points:
- LLM Integration: The
Extractstage uses aProviderabstraction defined insemantica/llm/provider.pyto handle all LLM interactions - Vector Storage: The
Storagestage relies on theVectorStorebase class insemantica/vector_store/vector_store.pyto persist embeddings
Both interfaces accept custom implementations through configuration or subclassing, allowing Semantica to orchestrate your existing infrastructure rather than replace it.
Integrating Custom LLMs via the Provider Interface
To route semantic extraction through your own LLM deployment, you create a provider class that adheres to the interface defined in semantica/llm/provider.py.
Creating a Custom LLM Provider
Subclass ProviderBase and implement three required methods: create_client(), generate(), and parse_response(). This encapsulates your SDK initialization, inference logic, and response formatting.
import os
import json
from semantica.llm.provider import ProviderBase, provider_registry
class MyCustomProvider(ProviderBase):
def create_client(self):
# Initialise your existing SDK or HTTP client here
return MyClient(api_key=os.getenv("MY_API_KEY"))
def generate(self, prompt: str, **kwargs):
resp = self.client.complete(prompt, **kwargs)
return resp.text
def parse_response(self, raw: str):
# Convert raw LLM output into the expected JSON schema
return json.loads(raw)
# Register the provider for discovery
provider_registry.register("my-custom-model", MyCustomProvider)
The generate() method handles the raw prompt-to-text inference, while parse_response() ensures the output conforms to Semantica’s expected entity and relation schemas.
Using Your Provider in Extraction Workflows
Once registered, reference your provider by name in any extraction call. The extract_entities_llm() function and similar methods accept a provider parameter that routes requests to your implementation:
from semantica.semantic_extract import extract_entities_llm
entities = extract_entities_llm(
text="Semantica supports modular LLM integration",
provider="my-custom-model" # routes to your custom class
)
This approach works for Azure OpenAI, Anthropic, Groq, or any private model hosting infrastructure that exposes a Python SDK or HTTP endpoint.
Connecting to Existing Vector Store Deployments
Semantica treats vector storage as an abstraction layer. The VectorStore class in semantica/vector_store/vector_store.py acts as a factory that instantiates concrete backends based on configuration parameters.
Configuring the VectorStore Abstraction
To point Semantica at an existing vector database, instantiate VectorStore with the appropriate backend identifier and connection credentials. The factory automatically returns the correct concrete implementation (e.g., QdrantStore, MilvusStore, PgVectorStore).
from semantica.vector_store import VectorStore
# Connect to an existing Qdrant instance
vector_store = VectorStore(
backend="qdrant", # matches semantica.vector_store.qdrant_store.QdrantStore
url="https://my-qdrant.example.com",
api_key=os.getenv("QDRANT_API_KEY"),
dimension=768, # must match your embedding model output
)
# Store embeddings with metadata
vector_store.upsert(
ids=["doc-1"],
vectors=[[0.12, -0.34, 0.56, ...]], # 768-dimensional embedding
metadata=[{"title": "Integration Guide", "source": "docs"}],
)
Supported Backend Implementations
The vector store layer supports multiple backends through concrete implementations:
- FAISS: Local file-based storage (
semantica/vector_store/faiss_store.py) - Qdrant: Cloud or self-hosted Qdrant clusters (
semantica/vector_store/qdrant_store.py) - Milvus, Weaviate, PgVector, Pinecone: Enterprise vector databases
- SQLite-Vec: Lightweight local option
Each backend inherits from the same base class, ensuring that methods like upsert(), search(), and delete() behave identically regardless of the underlying infrastructure.
End-to-End Integration Example
The following workflow demonstrates registering a custom OpenAI-compatible provider and connecting to an external Qdrant cluster within a single SemanticaPipeline defined in semantica/worker.py.
import os
from semantica.llm.provider import ProviderBase, provider_registry
from semantica.vector_store import VectorStore
from semantica.worker import SemanticaPipeline
# 1️⃣ Register custom LLM provider
class MyOpenAIProvider(ProviderBase):
def create_client(self):
from openai import OpenAI
return OpenAI(api_key=os.getenv("OPENAI_API_KEY"), base_url=os.getenv("CUSTOM_BASE_URL"))
def generate(self, prompt: str, **kwargs):
return self.client.chat.completions.create(
model="gpt-4", messages=[{"role": "user", "content": prompt}], **kwargs
).choices[0].message.content
def parse_response(self, raw: str):
import json
return json.loads(raw)
provider_registry.register("my-openai", MyOpenAIProvider)
# 2️⃣ Connect to existing vector store
vector_store = VectorStore(
backend="qdrant",
url=os.getenv("QDRANT_URL"),
api_key=os.getenv("QDRANT_API_KEY"),
dimension=1536, # e.g., text-embedding-3-large
)
# 3️⃣ Build pipeline with custom components
pipeline = SemanticaPipeline(
llm_provider="my-openai",
vector_store=vector_store,
)
# 4️⃣ Process a document through the full pipeline
pipeline.run("examples/technical_docs.pdf")
# 5️⃣ Query using hybrid KG-vector search
results = pipeline.context.graph_store.query(
"Find entities related to 'vector integration' with similar embeddings",
vector_query={"text": "vector integration", "top_k": 5}
)
This pattern preserves your existing LLM infrastructure investments while allowing Semantica to handle knowledge graph construction, conflict resolution, and semantic enrichment.
Summary
- Semantica’s architecture separates concerns into discrete stages, allowing you to swap LLM and storage components without modifying core logic.
- Custom LLM integration requires subclassing
ProviderBaseinsemantica/llm/provider.py, implementingcreate_client(),generate(), andparse_response(), then registering viaprovider_registry. - Vector store integration uses the
VectorStorefactory insemantica/vector_store/vector_store.pyto connect to existing FAISS, Qdrant, Milvus, or PgVector deployments by changing configuration parameters. - The
SemanticaPipelineclass insemantica/worker.pyorchestrates these components, enabling hybrid knowledge-graph and vector searches across your existing infrastructure.
Frequently Asked Questions
Can I use Azure OpenAI or Anthropic with Semantica?
Yes. Create a provider class that wraps your Azure OpenAI or Anthropic SDK client in the create_client() method, then implement generate() to call the appropriate API endpoints. Register the class via provider_registry.register() and reference it in your pipeline configuration.
Does Semantica support proprietary vector stores like Pinecone?
Yes. The VectorStore abstraction supports Pinecone and other proprietary backends. Instantiate VectorStore with backend="pinecone" and provide your API key and environment variables. The interface remains consistent with open-source alternatives like Qdrant or FAISS.
How do I migrate from the default FAISS to an existing Qdrant cluster?
Change your VectorStore instantiation from backend="faiss" to backend="qdrant" and supply the connection URL and API key. Ensure the dimension parameter matches your embedding model’s output size. Re-run your ingestion pipeline; Semantica will populate the existing Qdrant cluster with new embeddings while maintaining the same knowledge graph structure.
Is it possible to use different LLM providers for different extraction tasks?
Yes. The provider parameter in extraction functions like extract_entities_llm() accepts string identifiers registered in the provider registry. You can register multiple providers (e.g., "fast-model" for Groq and "powerful-model" for GPT-4) and specify which to use per call or per pipeline stage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →