How the OpenRisk RAG System Implements Knowledge Retrieval and Supports Multiple Embedding Models

The OpenRisk RAG system implements a two-stage retrieval pipeline that first attempts exact question matching via QARetriever, then falls back to semantic similarity search through EmbeddingRetriever, while supporting HuggingFace sentence-transformers, OpenAI embeddings, and any OpenAI-compatible API via a flexible factory pattern.

The derisk-ai/openderisk repository provides a modular RAG implementation that separates knowledge space management from vector similarity search. This architecture allows the system to combine exact-match retrieval with embedding-based semantic search while maintaining pluggable support for multiple embedding providers.

Two-Stage Knowledge Retrieval Architecture

The retrieval flow is orchestrated by KnowledgeSpaceRetriever in packages/derisk-serve/src/derisk_serve/rag/retriever/knowledge_space.py. This class builds a RetrieverChain that sequences two distinct retrieval strategies.

Knowledge Space Selection and Retriever Chain

When initialized with a space_id, the KnowledgeSpaceRetriever constructs a chain containing:

  1. QARetriever – Attempts exact matching against stored questions (implemented in packages/derisk-serve/src/derisk_serve/rag/retriever/qa_retriever.py).
  2. EmbeddingRetriever – Performs semantic similarity search if the QA retriever returns no results.

The retrieve() method executes this chain sequentially, returning Chunk objects that include content, metadata, and the source retriever name.

Vector Store Similarity via EmbeddingRetriever

The EmbeddingRetriever class in packages/derisk-serve/src/derisk_serve/rag/retriever/embedding.py handles the semantic search stage. It:

  • Obtains an embedding function from the embedding factory (see below).
  • Queries the configured vector store connector using similar_search or similar_search_with_scores.
  • Optionally rewrites queries and re-ranks results before returning the top-k chunks.

Supported Embedding Models and Factories

All embedding providers are abstracted behind EmbeddingFactory in packages/derisk-core/src/derisk/rag/embedding/embedding_factory.py. The system supports three primary factory methods plus a wrapper for custom implementations.

HuggingFace Sentence-Transformers

The DefaultEmbeddingFactory.default() method instantiates HuggingFaceEmbeddings from packages/derisk-core/src/derisk/rag/embedding/embeddings.py. This class:

  • Wraps sentence_transformers.SentenceTransformer.
  • Accepts any model name compatible with the HuggingFace Hub.
  • Supports multi-process encoding via encode_kwargs.

Example model: "sentence-transformers/all-mpnet-base-v2"

OpenAI and OpenAI-Compatible APIs

The factory provides two additional methods for cloud-based embeddings:

  • DefaultEmbeddingFactory.openai() – Connects to OpenAI or Azure OpenAI endpoints (e.g., "text-embedding-3-small"), requiring OPENAI_API_KEY.
  • DefaultEmbeddingFactory.remote() – Uses OpenAPIEmbeddings to communicate with any OpenAI-compatible HTTP endpoint, enabling self-hosted models.

Custom Embedding Wrappers

For pre-instantiated embedding objects, WrappedEmbeddingFactory allows users to inject custom Embeddings implementations without modifying the factory logic. All embedding classes register as resources (@register_resource) for UI discoverability.

Implementation Example

The following code demonstrates configuring a HuggingFace model, connecting it to a Chroma vector store, and executing retrieval through the full RAG pipeline:


# ---------------------------------------------------------

# 1️⃣ Create an embedding function (HF model)

# ---------------------------------------------------------

from derisk.rag.embedding import DefaultEmbeddingFactory

embedding_fn = DefaultEmbeddingFactory.default(
    model_name="sentence-transformers/all-mpnet-base-v2"
)

# ---------------------------------------------------------

# 2️⃣ Build a vector‑store connector (e.g. Chroma)

# ---------------------------------------------------------

from derisk.storage.vector_store.connector import VectorStoreConnector
from derisk.storage.vector_store.chroma_store import ChromaVectorConfig

vector_cfg = ChromaVectorConfig(name="risk_docs", embedding_fn=embedding_fn)
store = VectorStoreConnector(
    vector_store_type="Chroma", vector_store_config=vector_cfg
)

# ---------------------------------------------------------

# 3️⃣ Instantiate the embedding retriever

# ---------------------------------------------------------

from derisk.rag.retriever.embedding import EmbeddingRetriever

emb_ret = EmbeddingRetriever(index_store=store, top_k=5)

# ---------------------------------------------------------

# 4️⃣ Use KnowledgeSpaceRetriever for full RAG flow

# ---------------------------------------------------------

from derisk_serve.rag.retriever.knowledge_space import KnowledgeSpaceRetriever

rag = KnowledgeSpaceRetriever(
    space_id="risk_knowledge_space",
    top_k=5,
    embedding_model="sentence-transformers/all-mpnet-base-v2",
    system_app=system_app,
)

chunks = rag.retrieve("What are the main credit‑risk factors?")
for c in chunks:
    print(f"• {c.content[:120]}…")

Summary

Frequently Asked Questions

What is the two-stage retrieval process in OpenRisk RAG?

The OpenRisk RAG system implements a sequential retrieval chain. First, the QARetriever attempts exact matching against pre-stored questions. If no matches are found, the system automatically falls back to the EmbeddingRetriever, which performs semantic similarity search against the vector store using the configured embedding model.

Which embedding models does OpenRisk support out of the box?

OpenRisk supports three primary embedding families through the DefaultEmbeddingFactory: HuggingFace sentence-transformers (any model from the Hub), OpenAI embeddings (including Azure OpenAI), and generic OpenAI-compatible APIs for self-hosted models. Additionally, WrappedEmbeddingFactory allows injection of custom embedding implementations.

How do I configure a custom HuggingFace model in OpenRisk?

To use a specific HuggingFace model, invoke DefaultEmbeddingFactory.default() with the model_name parameter set to the desired HuggingFace Hub identifier (e.g., "sentence-transformers/all-mpnet-base-v2"). This instantiates HuggingFaceEmbeddings, which wraps sentence_transformers.SentenceTransformer and supports additional encoding arguments via encode_kwargs.

Can I use OpenAI-compatible endpoints with the OpenRisk RAG system?

Yes. The DefaultEmbeddingFactory.remote() method creates an OpenAPIEmbeddings instance that communicates with any OpenAI-compatible HTTP endpoint. This enables integration with self-hosted embedding services or alternative providers that implement the OpenAI /embeddings API schema, requiring only the endpoint URL and optional API key configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →