How to Use the Derisk Embedding API for Vector Storage and Similarity Search

The Derisk embedding API bridges language model encoders and vector store backends through the VectorStoreConnector, enabling seamless document indexing via DefaultEmbeddingFactory and semantic retrieval using similar_search with configurable scoring.

The derisk-ai/openderisk repository provides a modular embedding API designed to unify vector storage operations for retrieval-augmented generation (RAG) pipelines. This system abstracts database-specific implementations behind a consistent connector interface, allowing you to index text chunks and execute similarity searches across Chroma, PGVector, Milvus, and other backends without modifying core application logic.

Core Architecture of the Embedding API

The embedding API follows a factory-to-connector pattern that separates embedding generation from storage concerns.

Embedding Factories

The DefaultEmbeddingFactory class in derisk/rag/embedding/embedding_factory.py resolves concrete Embeddings implementations based on model names or API endpoints. It supports:

  • Local HuggingFace models via HuggingFaceEmbeddings (sentence-transformers)
  • Remote OpenAI-compatible endpoints via wrapped API clients
  • Custom providers through the abstract EmbeddingFactory base class

Vector Store Connector

The VectorStoreConnector class in derisk_serve/rag/connector.py acts as the central dispatcher. It registers concrete store implementations at import time using the @register_resource decorator, then maps VectorStoreConfig objects to specific backends like ChromaStore or PGVectorStore. This design enables the connector to forward load_document, similar_search, and similar_search_with_scores calls to the appropriate underlying database.

Concrete Vector Store Implementations

Actual storage and retrieval logic resides in backend-specific classes:

These implementations handle low-level CRUD operations, index management, and distance calculations.

Initializing Embedding Models

You must instantiate an embeddings provider before connecting to a vector store.

Local HuggingFace Sentence Transformers

Use DefaultEmbeddingFactory.default() to load a local sentence-transformers model. If you omit the model_name parameter, it falls back to the DEFAULT_MODEL_NAME defined in the factory configuration.

from derisk.rag.embedding import DefaultEmbeddingFactory

embeddings = DefaultEmbeddingFactory.default(
    model_name="sentence-transformers/all-mpnet-base-v2"
)

Remote OpenAI-Compatible Endpoints

For cloud-based or self-hosted embedding services, use DefaultEmbeddingFactory.openai() to configure API access:

embeddings = DefaultEmbeddingFactory.openai(
    api_url="https://my-embedding-service/v1/embeddings",
    api_key="MY_API_KEY",
    model_name="text-embedding-3-small",
)

Configuring Vector Storage Backends

The VectorStoreConnector requires a configuration object specific to your chosen backend.

ChromaDB Setup

Import ChromaVectorConfig to define collection names and persistence settings:

from derisk_ext.storage.vector_store.chroma_store import ChromaVectorConfig

vector_cfg = ChromaVectorConfig(name="my_collection")
connector = VectorStoreConnector.from_default(
    vector_store_type="Chroma",
    embedding_fn=embeddings,
    vector_store_config=vector_cfg,
)

PGVector Setup

For PostgreSQL with the pgvector extension, use PGVectorConfig with a connection string:

from derisk_ext.storage.vector_store.pgvector_store import PGVectorConfig

pg_cfg = PGVectorConfig(
    name="my_pg_collection",
    connection_string="postgresql://user:pwd@localhost:5432/db"
)
connector = VectorStoreConnector.from_default(
    vector_store_type="PGVector",
    embedding_fn=embeddings,
    vector_store_config=pg_cfg,
)

Once configured, the connector accepts Chunk objects from derisk.core and manages the indexing and query pipeline.

Synchronous Document Indexing

Convert raw text into Chunk instances, load them into the store, and execute vector similarity search:

from derisk.core import Chunk
from derisk.rag.embedding import DefaultEmbeddingFactory
from derisk_serve.rag.connector import VectorStoreConnector
from derisk_ext.storage.vector_store.chroma_store import ChromaVectorConfig

# ① Create embeddings

embeddings = DefaultEmbeddingFactory.default(
    model_name="sentence-transformers/all-mpnet-base-v2"
)

# ② Build connector

vector_cfg = ChromaVectorConfig(name="my_collection")
connector = VectorStoreConnector.from_default(
    vector_store_type="Chroma",
    embedding_fn=embeddings,
    vector_store_config=vector_cfg,
)

# ③ Prepare chunks

texts = [
    "Derisk is an AI-augmented risk-analysis platform.",
    "Vector stores enable fast similarity search."
]
chunks = [Chunk(content=t, metadata={"source": i}) for i, t in enumerate(texts)]

# ④ Load documents

doc_ids = connector.load_document(chunks)
print("Indexed docs:", doc_ids)

# ⑤ Similarity search

query = "What does Derisk do?"
hits = connector.similar_search(query, top_k=3)
for hit in hits:
    print(f"→ {hit.content} (score: {hit.score:.2f})")

Asynchronous Operations for Web Services

For high-concurrency applications, use the async API methods aload_document and asimilar_search_with_scores:

import asyncio
from derisk.core import Chunk
from derisk.rag.embedding import DefaultEmbeddingFactory
from derisk_serve.rag.connector import VectorStoreConnector

async def demo():
    embed = DefaultEmbeddingFactory.default(
        model_name="sentence-transformers/all-mpnet-base-v2"
    )
    connector = VectorStoreConnector.from_default(
        embedding_fn=embed,
        vector_store_config=ChromaVectorConfig(name="async_demo")
    )
    
    # Async load

    await connector.aload_document([
        Chunk(content="async doc 1"), 
        Chunk(content="async doc 2")
    ])
    
    # Async search with scores and threshold

    results = await connector.asimilar_search_with_scores(
        "async query", top_k=4, score_threshold=0.1
    )
    for r in results:
        print(r.content, r.score)

asyncio.run(demo())

Advanced Search Parameters

The similar_search_with_scores method accepts additional filtering parameters:

  • top_k: Maximum number of results to return (default varies by backend)
  • score_threshold: Minimum similarity score (0.0 to 1.0) to include in results
  • filter: Backend-specific metadata filters (e.g., {"source": 1})

These parameters are forwarded to the underlying store implementation in chroma_store.py or pgvector_store.py.

Summary

  • The Derisk embedding API unifies embedding generation and vector storage through the VectorStoreConnector class in derisk_serve/rag/connector.py.
  • Use DefaultEmbeddingFactory to instantiate local HuggingFace models or remote OpenAI-compatible endpoints.
  • Vector stores are configured via VectorStoreConfig subclasses like ChromaVectorConfig or PGVectorConfig.
  • Documents must be wrapped in Chunk objects before calling load_document or aload_document.
  • Perform retrieval using similar_search for basic results or similar_search_with_scores when you need relevance thresholds.

Frequently Asked Questions

What embedding models are supported by the Derisk embedding API?

The API supports any sentence-transformers model via HuggingFaceEmbeddings and any OpenAI-compatible REST endpoint via the openai factory method. You can extend support by subclassing EmbeddingFactory in derisk/rag/embedding/embedding_factory.py to add custom providers.

How does VectorStoreConnector select the vector store backend?

The connector uses the @register_resource decorator to build a registry of store classes at import time. When you call from_default() with a vector_store_type string (e.g., "Chroma" or "PGVector"), it looks up the registered class in derisk_serve/rag/connector.py and instantiates it with your VectorStoreConfig.

Can I use async operations for high-throughput applications?

Yes. The connector exposes aload_document and asimilar_search_with_scores methods that delegate to async implementations in the underlying store classes. Use these in FastAPI or other async web frameworks to avoid blocking the event loop during embedding generation or database queries.

What is the difference between similar_search and similar_search_with_scores?

similar_search returns a list of Chunk objects ranked by relevance, while similar_search_with_scores returns tuples of (chunk, score) allowing you to filter results by the score_threshold parameter. Both methods call the underlying store's query implementation, but similar_search_with_scores exposes the raw similarity metric used by the vector database.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →