How to Use the Derisk Embedding API for Vector Storage and Similarity Search
The Derisk embedding API bridges language model encoders and vector store backends through the VectorStoreConnector, enabling seamless document indexing via DefaultEmbeddingFactory and semantic retrieval using similar_search with configurable scoring.
The derisk-ai/openderisk repository provides a modular embedding API designed to unify vector storage operations for retrieval-augmented generation (RAG) pipelines. This system abstracts database-specific implementations behind a consistent connector interface, allowing you to index text chunks and execute similarity searches across Chroma, PGVector, Milvus, and other backends without modifying core application logic.
Core Architecture of the Embedding API
The embedding API follows a factory-to-connector pattern that separates embedding generation from storage concerns.
Embedding Factories
The DefaultEmbeddingFactory class in derisk/rag/embedding/embedding_factory.py resolves concrete Embeddings implementations based on model names or API endpoints. It supports:
- Local HuggingFace models via
HuggingFaceEmbeddings(sentence-transformers) - Remote OpenAI-compatible endpoints via wrapped API clients
- Custom providers through the abstract
EmbeddingFactorybase class
Vector Store Connector
The VectorStoreConnector class in derisk_serve/rag/connector.py acts as the central dispatcher. It registers concrete store implementations at import time using the @register_resource decorator, then maps VectorStoreConfig objects to specific backends like ChromaStore or PGVectorStore. This design enables the connector to forward load_document, similar_search, and similar_search_with_scores calls to the appropriate underlying database.
Concrete Vector Store Implementations
Actual storage and retrieval logic resides in backend-specific classes:
ChromaStoreinderisk_ext/storage/vector_store/chroma_store.pyPGVectorStoreinderisk_ext/storage/vector_store/pgvector_store.py
These implementations handle low-level CRUD operations, index management, and distance calculations.
Initializing Embedding Models
You must instantiate an embeddings provider before connecting to a vector store.
Local HuggingFace Sentence Transformers
Use DefaultEmbeddingFactory.default() to load a local sentence-transformers model. If you omit the model_name parameter, it falls back to the DEFAULT_MODEL_NAME defined in the factory configuration.
from derisk.rag.embedding import DefaultEmbeddingFactory
embeddings = DefaultEmbeddingFactory.default(
model_name="sentence-transformers/all-mpnet-base-v2"
)
Remote OpenAI-Compatible Endpoints
For cloud-based or self-hosted embedding services, use DefaultEmbeddingFactory.openai() to configure API access:
embeddings = DefaultEmbeddingFactory.openai(
api_url="https://my-embedding-service/v1/embeddings",
api_key="MY_API_KEY",
model_name="text-embedding-3-small",
)
Configuring Vector Storage Backends
The VectorStoreConnector requires a configuration object specific to your chosen backend.
ChromaDB Setup
Import ChromaVectorConfig to define collection names and persistence settings:
from derisk_ext.storage.vector_store.chroma_store import ChromaVectorConfig
vector_cfg = ChromaVectorConfig(name="my_collection")
connector = VectorStoreConnector.from_default(
vector_store_type="Chroma",
embedding_fn=embeddings,
vector_store_config=vector_cfg,
)
PGVector Setup
For PostgreSQL with the pgvector extension, use PGVectorConfig with a connection string:
from derisk_ext.storage.vector_store.pgvector_store import PGVectorConfig
pg_cfg = PGVectorConfig(
name="my_pg_collection",
connection_string="postgresql://user:pwd@localhost:5432/db"
)
connector = VectorStoreConnector.from_default(
vector_store_type="PGVector",
embedding_fn=embeddings,
vector_store_config=pg_cfg,
)
Loading Documents and Performing Similarity Search
Once configured, the connector accepts Chunk objects from derisk.core and manages the indexing and query pipeline.
Synchronous Document Indexing
Convert raw text into Chunk instances, load them into the store, and execute vector similarity search:
from derisk.core import Chunk
from derisk.rag.embedding import DefaultEmbeddingFactory
from derisk_serve.rag.connector import VectorStoreConnector
from derisk_ext.storage.vector_store.chroma_store import ChromaVectorConfig
# ① Create embeddings
embeddings = DefaultEmbeddingFactory.default(
model_name="sentence-transformers/all-mpnet-base-v2"
)
# ② Build connector
vector_cfg = ChromaVectorConfig(name="my_collection")
connector = VectorStoreConnector.from_default(
vector_store_type="Chroma",
embedding_fn=embeddings,
vector_store_config=vector_cfg,
)
# ③ Prepare chunks
texts = [
"Derisk is an AI-augmented risk-analysis platform.",
"Vector stores enable fast similarity search."
]
chunks = [Chunk(content=t, metadata={"source": i}) for i, t in enumerate(texts)]
# ④ Load documents
doc_ids = connector.load_document(chunks)
print("Indexed docs:", doc_ids)
# ⑤ Similarity search
query = "What does Derisk do?"
hits = connector.similar_search(query, top_k=3)
for hit in hits:
print(f"→ {hit.content} (score: {hit.score:.2f})")
Asynchronous Operations for Web Services
For high-concurrency applications, use the async API methods aload_document and asimilar_search_with_scores:
import asyncio
from derisk.core import Chunk
from derisk.rag.embedding import DefaultEmbeddingFactory
from derisk_serve.rag.connector import VectorStoreConnector
async def demo():
embed = DefaultEmbeddingFactory.default(
model_name="sentence-transformers/all-mpnet-base-v2"
)
connector = VectorStoreConnector.from_default(
embedding_fn=embed,
vector_store_config=ChromaVectorConfig(name="async_demo")
)
# Async load
await connector.aload_document([
Chunk(content="async doc 1"),
Chunk(content="async doc 2")
])
# Async search with scores and threshold
results = await connector.asimilar_search_with_scores(
"async query", top_k=4, score_threshold=0.1
)
for r in results:
print(r.content, r.score)
asyncio.run(demo())
Advanced Search Parameters
The similar_search_with_scores method accepts additional filtering parameters:
- top_k: Maximum number of results to return (default varies by backend)
- score_threshold: Minimum similarity score (0.0 to 1.0) to include in results
- filter: Backend-specific metadata filters (e.g.,
{"source": 1})
These parameters are forwarded to the underlying store implementation in chroma_store.py or pgvector_store.py.
Summary
- The Derisk embedding API unifies embedding generation and vector storage through the
VectorStoreConnectorclass inderisk_serve/rag/connector.py. - Use DefaultEmbeddingFactory to instantiate local HuggingFace models or remote OpenAI-compatible endpoints.
- Vector stores are configured via VectorStoreConfig subclasses like
ChromaVectorConfigorPGVectorConfig. - Documents must be wrapped in Chunk objects before calling
load_documentoraload_document. - Perform retrieval using similar_search for basic results or similar_search_with_scores when you need relevance thresholds.
Frequently Asked Questions
What embedding models are supported by the Derisk embedding API?
The API supports any sentence-transformers model via HuggingFaceEmbeddings and any OpenAI-compatible REST endpoint via the openai factory method. You can extend support by subclassing EmbeddingFactory in derisk/rag/embedding/embedding_factory.py to add custom providers.
How does VectorStoreConnector select the vector store backend?
The connector uses the @register_resource decorator to build a registry of store classes at import time. When you call from_default() with a vector_store_type string (e.g., "Chroma" or "PGVector"), it looks up the registered class in derisk_serve/rag/connector.py and instantiates it with your VectorStoreConfig.
Can I use async operations for high-throughput applications?
Yes. The connector exposes aload_document and asimilar_search_with_scores methods that delegate to async implementations in the underlying store classes. Use these in FastAPI or other async web frameworks to avoid blocking the event loop during embedding generation or database queries.
What is the difference between similar_search and similar_search_with_scores?
similar_search returns a list of Chunk objects ranked by relevance, while similar_search_with_scores returns tuples of (chunk, score) allowing you to filter results by the score_threshold parameter. Both methods call the underlying store's query implementation, but similar_search_with_scores exposes the raw similarity metric used by the vector database.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →