BM25 vs Embedding-Based Semantic Retrieval in OpenRisk AI: A Technical Comparison
BM25 retriever excels at fast, deterministic keyword matching using lexical term frequencies, while embedding-based semantic retrieval captures conceptual meaning through dense vector similarity, and OpenRisk AI combines both via a configurable retriever chain for optimal RAG performance.
OpenRisk AI (derisk-ai/openderisk) implements a dual-track retrieval system that supports both classical BM25 lexical search and modern embedding-based semantic search. Understanding how these retrievers compare—and when to use each—is critical for building effective Retrieval-Augmented Generation (RAG) pipelines in financial risk analysis and knowledge management systems.
Core Architectural Differences
BM25: Probabilistic Term Matching
The BM25 retriever implements the Okapi BM25 algorithm, a classical probabilistic model that scores documents based on term frequency, inverse document frequency, and document length normalization. In packages/derisk-ext/src/derisk_ext/rag/retriever/bm25.py, the BM25Retriever class inherits from BaseRetriever and interfaces with an Elasticsearch-style full-text index defined in packages/derisk-ext/src/derisk_ext/storage/full_text/elasticsearch.py.
This approach requires exact token overlap (or near-exact after stemming and text analysis) between queries and documents. It performs exceptionally well for keyword-driven searches, structured identifiers, and short fields where deterministic, reproducible ranking is required.
Embedding: Dense Vector Similarity
The embedding-based retriever operates on dense vector representations of text. Implemented in packages/derisk-core/src/derisk_core/rag/retriever/embedding.py, the EmbeddingRetriever class wraps a vector store (such as FAISS or Qdrant) and an LLM embedding model (e.g., text-embedding-ada-002).
During retrieval, the system embeds the query on-the-fly and performs a nearest-neighbor search (typically cosine similarity or inner-product) against pre-computed document vectors. This captures semantic relationships, allowing the system to match documents that mean the same thing even when they use different wording, synonyms, or paraphrases.
Implementation in OpenRisk AI
BM25 Retriever Class
The BM25Retriever in packages/derisk-ext/src/derisk_ext/rag/retriever/bm25.py initializes with an index name and Elasticsearch-compatible schema where the field mapping specifies "type": "BM25". The retrieve() method constructs a BM25 query from user text, executes it against the inverted index, and returns top-k document IDs with raw relevance scores.
Embedding Retriever Class
The EmbeddingRetriever in packages/derisk-core/src/derisk_core/rag/retriever/embedding.py loads the configured embedding model during initialization. Documents are pre-vectorized and stored in the vector database; queries are embedded at runtime. The retrieval latency depends on embedding generation time and vector search complexity, trading speed for semantic depth.
Performance and Use Case Comparison
When deciding between BM25 and embedding-based retrieval in OpenRisk AI, consider the following trade-offs:
-
Speed and Resource Efficiency: BM25 requires only a lightweight inverted index and performs very fast lookups without GPU resources. Embedding retrieval incurs latency from neural model inference and high-dimensional vector search.
-
Scalability: BM25 scales linearly with indexed tokens and handles billions of documents via Elasticsearch clusters. Embedding scalability depends on vector store limitations; high-dimensional indexes require significant RAM and often use approximate nearest neighbor (ANN) structures that trade recall for speed.
-
Robustness: BM25 is brittle to misspellings and synonym usage unless augmented with analysis pipelines. Embeddings naturally tolerate paraphrases, typos, and semantic variations due to their dense representation learning.
-
Determinism: BM25 produces reproducible, explainable scores based on term statistics. Embedding similarity is opaque and can vary with model updates or embedding drift.
Hybrid Retrieval Strategy
OpenRisk AI does not force an either/or choice. The KnowledgeSpaceRetriever in packages/derisk-serve/src/derisk_serve/rag/retriever/knowledge_space.py constructs a retriever chain using the RetrieverChain class from packages/derisk-serve/src/derisk_serve/rag/retriever/retriever_chain.py.
This architecture runs BM25 first for lexical precision on structured fields, then invokes the EmbeddingRetriever for semantic depth on unstructured content. The chain merges candidates from both sources, optionally applies a reranker, and returns a unified result list. This hybrid approach leverages BM25’s speed for exact matches while capturing semantic nuance through embeddings, delivering superior RAG performance for financial risk documents.
Code Examples
The following snippets demonstrate how to instantiate and use each retriever in the OpenRisk AI framework.
# Example 1 – BM25 retriever (lexical)
from derisk_ext.rag.retriever.bm25 import BM25Retriever
# Create a BM25 retriever that points to the “knowledge” collection
bm25 = BM25Retriever(
index_name="knowledge", # Elasticsearch index name
top_k=5, # number of hits to return
)
# Run a lexical query
hits = bm25.retrieve("What is the definition of VaR?")
print("BM25 results:", hits)
# Example 2 – Embedding‑based retriever (semantic)
from derisk.rag.retriever.embedding import EmbeddingRetriever
# Create an embedding retriever that uses the default LLM embedding model
embed = EmbeddingRetriever(
vector_store="faiss", # could be faiss, qdrant, etc.
top_k=5,
embed_model="text-embedding-ada-002",
)
# Run a semantically similar query
hits = embed.retrieve("Explain value‑at‑risk in simple terms")
print("Embedding results:", hits)
# Example 3 – Hybrid chain (BM25 + Embedding)
from derisk.rag.retriever.retriever_chain import RetrieverChain
from derisk_ext.rag.retriever.bm25 import BM25Retriever
from derisk.rag.retriever.embedding import EmbeddingRetriever
bm25 = BM25Retriever(index_name="knowledge", top_k=5)
embed = EmbeddingRetriever(vector_store="faiss", top_k=5)
# Combine them: BM25 first, then embed‑based reranking
chain = RetrieverChain(retrievers=[bm25, embed])
results = chain.retrieve("What does \"risk‑adjusted return\" mean?")
print("Hybrid results:", results)
Summary
-
BM25 Retriever (
BM25Retrieverinpackages/derisk-ext/src/derisk_ext/rag/retriever/bm25.py) provides fast, deterministic lexical matching using probabilistic term weighting, ideal for keyword searches and structured fields. -
Embedding Retriever (
EmbeddingRetrieverinpackages/derisk-core/src/derisk_core/rag/retriever/embedding.py) enables semantic search by matching dense vector representations, capturing meaning beyond exact word matches but requiring more compute resources. -
Hybrid Architecture via
RetrieverChainandKnowledgeSpaceRetrieverallows OpenRisk AI to combine both approaches, leveraging BM25 for speed and precision on exact terms while using embeddings for semantic depth and paraphrase handling.
Frequently Asked Questions
What is the main difference between BM25 and embedding-based retrieval in OpenRisk AI?
BM25 is a lexical retrieval method that scores documents based on term frequency and inverse document frequency, requiring exact or near-exact word matches. Embedding-based retrieval converts text into dense vectors and measures semantic similarity, allowing it to find conceptually related documents even when they share no common keywords.
When should I use BM25 over embedding-based retrieval?
Use BM25 when you need fast, reproducible results for keyword-driven searches, structured identifiers (like document IDs or regulatory codes), or short text fields where exact term presence matters. It requires minimal computational resources and scales efficiently to billions of documents via Elasticsearch-style indexes.
How does OpenRisk AI combine BM25 and embedding retrievers?
OpenRisk AI uses the RetrieverChain class (located in packages/derisk-serve/src/derisk_serve/rag/retriever/retriever_chain.py) to orchestrate multiple retrievers. The KnowledgeSpaceRetriever typically configures this chain to run BM25 first for lexical candidate generation, then invokes the EmbeddingRetriever to rerank or augment results based on semantic similarity, merging both signal types into a unified output.
Is BM25 sufficient for modern RAG applications, or is embedding retrieval necessary?
While BM25 remains effective for structured and keyword-heavy domains, modern RAG applications typically benefit from embedding retrieval to handle natural language paraphrasing, semantic nuances, and long-form unstructured documents. OpenRisk AI’s architecture treats these as complementary: BM25 provides precision and speed for exact matches, while embeddings provide recall and semantic coverage, with the hybrid approach delivering superior overall retrieval quality.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →