How the OpenRisk RAG System Implements Knowledge Retrieval and Supports Multiple Embedding Models
The OpenRisk RAG system implements a two-stage retrieval pipeline that first attempts exact question matching via QARetriever, then falls back to semantic similarity search through EmbeddingRetriever, while supporting HuggingFace sentence-transformers, OpenAI embeddings, and any OpenAI-compatible API via a flexible factory pattern.
The derisk-ai/openderisk repository provides a modular RAG implementation that separates knowledge space management from vector similarity search. This architecture allows the system to combine exact-match retrieval with embedding-based semantic search while maintaining pluggable support for multiple embedding providers.
Two-Stage Knowledge Retrieval Architecture
The retrieval flow is orchestrated by KnowledgeSpaceRetriever in packages/derisk-serve/src/derisk_serve/rag/retriever/knowledge_space.py. This class builds a RetrieverChain that sequences two distinct retrieval strategies.
Knowledge Space Selection and Retriever Chain
When initialized with a space_id, the KnowledgeSpaceRetriever constructs a chain containing:
QARetriever– Attempts exact matching against stored questions (implemented inpackages/derisk-serve/src/derisk_serve/rag/retriever/qa_retriever.py).EmbeddingRetriever– Performs semantic similarity search if the QA retriever returns no results.
The retrieve() method executes this chain sequentially, returning Chunk objects that include content, metadata, and the source retriever name.
Vector Store Similarity via EmbeddingRetriever
The EmbeddingRetriever class in packages/derisk-serve/src/derisk_serve/rag/retriever/embedding.py handles the semantic search stage. It:
- Obtains an embedding function from the embedding factory (see below).
- Queries the configured vector store connector using
similar_searchorsimilar_search_with_scores. - Optionally rewrites queries and re-ranks results before returning the top-k chunks.
Supported Embedding Models and Factories
All embedding providers are abstracted behind EmbeddingFactory in packages/derisk-core/src/derisk/rag/embedding/embedding_factory.py. The system supports three primary factory methods plus a wrapper for custom implementations.
HuggingFace Sentence-Transformers
The DefaultEmbeddingFactory.default() method instantiates HuggingFaceEmbeddings from packages/derisk-core/src/derisk/rag/embedding/embeddings.py. This class:
- Wraps
sentence_transformers.SentenceTransformer. - Accepts any model name compatible with the HuggingFace Hub.
- Supports multi-process encoding via
encode_kwargs.
Example model: "sentence-transformers/all-mpnet-base-v2"
OpenAI and OpenAI-Compatible APIs
The factory provides two additional methods for cloud-based embeddings:
DefaultEmbeddingFactory.openai()– Connects to OpenAI or Azure OpenAI endpoints (e.g.,"text-embedding-3-small"), requiringOPENAI_API_KEY.DefaultEmbeddingFactory.remote()– UsesOpenAPIEmbeddingsto communicate with any OpenAI-compatible HTTP endpoint, enabling self-hosted models.
Custom Embedding Wrappers
For pre-instantiated embedding objects, WrappedEmbeddingFactory allows users to inject custom Embeddings implementations without modifying the factory logic. All embedding classes register as resources (@register_resource) for UI discoverability.
Implementation Example
The following code demonstrates configuring a HuggingFace model, connecting it to a Chroma vector store, and executing retrieval through the full RAG pipeline:
# ---------------------------------------------------------
# 1️⃣ Create an embedding function (HF model)
# ---------------------------------------------------------
from derisk.rag.embedding import DefaultEmbeddingFactory
embedding_fn = DefaultEmbeddingFactory.default(
model_name="sentence-transformers/all-mpnet-base-v2"
)
# ---------------------------------------------------------
# 2️⃣ Build a vector‑store connector (e.g. Chroma)
# ---------------------------------------------------------
from derisk.storage.vector_store.connector import VectorStoreConnector
from derisk.storage.vector_store.chroma_store import ChromaVectorConfig
vector_cfg = ChromaVectorConfig(name="risk_docs", embedding_fn=embedding_fn)
store = VectorStoreConnector(
vector_store_type="Chroma", vector_store_config=vector_cfg
)
# ---------------------------------------------------------
# 3️⃣ Instantiate the embedding retriever
# ---------------------------------------------------------
from derisk.rag.retriever.embedding import EmbeddingRetriever
emb_ret = EmbeddingRetriever(index_store=store, top_k=5)
# ---------------------------------------------------------
# 4️⃣ Use KnowledgeSpaceRetriever for full RAG flow
# ---------------------------------------------------------
from derisk_serve.rag.retriever.knowledge_space import KnowledgeSpaceRetriever
rag = KnowledgeSpaceRetriever(
space_id="risk_knowledge_space",
top_k=5,
embedding_model="sentence-transformers/all-mpnet-base-v2",
system_app=system_app,
)
chunks = rag.retrieve("What are the main credit‑risk factors?")
for c in chunks:
print(f"• {c.content[:120]}…")
Summary
- Two-stage retrieval:
KnowledgeSpaceRetrieverorchestrates aRetrieverChainthat first attempts exact question matching viaQARetriever, then falls back to semantic search viaEmbeddingRetriever. - Flexible embedding support: The
EmbeddingFactoryabstraction supports HuggingFace sentence-transformers, OpenAI embeddings, and any OpenAI-compatible API endpoint. - Key implementation files:
packages/derisk-serve/src/derisk_serve/rag/retriever/knowledge_space.pyhandles orchestration, whilepackages/derisk-core/src/derisk/rag/embedding/embeddings.pyprovides concrete embedding implementations.
Frequently Asked Questions
What is the two-stage retrieval process in OpenRisk RAG?
The OpenRisk RAG system implements a sequential retrieval chain. First, the QARetriever attempts exact matching against pre-stored questions. If no matches are found, the system automatically falls back to the EmbeddingRetriever, which performs semantic similarity search against the vector store using the configured embedding model.
Which embedding models does OpenRisk support out of the box?
OpenRisk supports three primary embedding families through the DefaultEmbeddingFactory: HuggingFace sentence-transformers (any model from the Hub), OpenAI embeddings (including Azure OpenAI), and generic OpenAI-compatible APIs for self-hosted models. Additionally, WrappedEmbeddingFactory allows injection of custom embedding implementations.
How do I configure a custom HuggingFace model in OpenRisk?
To use a specific HuggingFace model, invoke DefaultEmbeddingFactory.default() with the model_name parameter set to the desired HuggingFace Hub identifier (e.g., "sentence-transformers/all-mpnet-base-v2"). This instantiates HuggingFaceEmbeddings, which wraps sentence_transformers.SentenceTransformer and supports additional encoding arguments via encode_kwargs.
Can I use OpenAI-compatible endpoints with the OpenRisk RAG system?
Yes. The DefaultEmbeddingFactory.remote() method creates an OpenAPIEmbeddings instance that communicates with any OpenAI-compatible HTTP endpoint. This enables integration with self-hosted embedding services or alternative providers that implement the OpenAI /embeddings API schema, requiring only the endpoint URL and optional API key configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →