How to Apply Metadata Filtering in RAG Queries for Precise Results in OpenDeRisk
OpenDeRisk enables precise metadata filtering in RAG queries by accepting a MetadataFilters object in KnowledgeSearchRequest, which the service validates and translates into vector-store filter expressions before executing filtered similarity search.
OpenDeRisk is an open-source risk management platform that implements a sophisticated Retrieval-Augmented Generation (RAG) pipeline. When you need to constrain search results to specific document sources, categories, or custom tags, you can implement metadata filtering in RAG queries to retrieve only the most relevant chunks. This guide explains the architecture and implementation based on the actual source code in the derisk-ai/openderisk repository.
Understanding the Metadata Filtering Architecture
The metadata filtering system in OpenDeRisk spans multiple layers, from API schema definitions to low-level vector store operations.
Core Filter Models (filters.py)
The foundational data structures reside in packages/derisk-core/src/derisk/storage/vector_store/filters.py. This module defines:
FilterOperator: Enum for comparison operators (EQ,IN,EXISTS, etc.)FilterCondition: Logical condition (AND/OR) binding multiple filtersMetadataFilter: Individual filter withkey,operator, andvalueMetadataFilters: Container combining condition and filter list
API Schema Integration (schemas.py)
The KnowledgeSearchRequest class in packages/derisk-serve/src/derisk_serve/rag/api/schemas.py accepts an optional metadata_filters field (lines 321-324). This allows clients to pass complex filtering criteria alongside the natural language query.
Implementing Metadata Filters in Practice
To apply metadata filtering in RAG queries, you construct filter objects and attach them to search requests.
Building MetadataFilters Objects
Create filter specifications using the classes from filters.py:
from derisk.storage.vector_store.filters import (
MetadataFilters, MetadataFilter, FilterOperator, FilterCondition
)
# Example: retrieve only chunks where `source` == "yuque" AND `category` is in ["finance", "risk"]
metadata_filters = MetadataFilters(
condition=FilterCondition.AND,
filters=[
MetadataFilter(
key="source",
operator=FilterOperator.EQ,
value="yuque"
),
MetadataFilter(
key="category",
operator=FilterOperator.IN,
value=["finance", "risk"]
),
],
)
Attaching Filters to Search Requests
Pass the constructed filters to the RAG service via KnowledgeSearchRequest:
from derisk_serve.rag.api.schemas import KnowledgeSearchRequest
from derisk_serve.rag.service.service import Service
# Build the request
request = KnowledgeSearchRequest(
query="What are the latest credit risk regulations?",
knowledge_ids=["space_123"],
metadata_filters=metadata_filters, # Apply filtering here
top_k=20,
similarity_score_threshold=0.2,
)
# Execute retrieval
chunks = await service.aretrieve_with_scores(
knowledge_id="space_123",
request=request,
knowledge_space_retriever=await service.acreate_knowledge_space_retriever(
knowledge_id="space_123",
top_k=20,
),
)
Using the REST API
For HTTP clients, embed filters in the JSON payload:
{
"query": "Explain the impact of Basel III on loan pricing",
"knowledge_ids": ["space_987"],
"metadata_filters": {
"condition": "and",
"filters": [
{ "key": "source", "operator": "==", "value": "yuque" },
{ "key": "tags", "operator": "in", "value": ["basel3", "pricing"] }
]
},
"top_k": 10,
"similarity_score_threshold": 0.15
}
POST this to the /rag/search endpoint. The backend validates the structure via Service.check_metadata_filters before execution.
Advanced Features: LLM-Driven Filter Expansion
OpenDeRisk can automatically populate missing filter values using the TagsExtractor LLM helper. When you provide a filter with an empty value list, the Service.aget_metadata_filter method (defined around line 5049 in service.py) performs the following:
- Detects the empty
filter_value - Retrieves all possible values for that metadata key from the knowledge space
- Prompts the LLM to select relevant tags based on the query context
- Substitutes inferred values before passing filters to
KnowledgeSpaceRetriever
Example of triggering auto-generation:
metadata_filters = MetadataFilters(
condition=FilterCondition.AND,
filters=[
MetadataFilter(key="category", operator=FilterOperator.IN, value=[]), # Empty triggers LLM
],
)
Summary
MetadataFiltersobjects infilters.pyprovide the type-safe structure for constraints.- Pass filters via
metadata_filtersinKnowledgeSearchRequest(schemas.py) to constrain RAG retrieval. - The service layer validates filters via
Service.check_metadata_filtersand can auto-generate values viaTagsExtractor. KnowledgeSpaceRetrievertranslatesMetadataFiltersinto vector-store specific expressions for execution.- Use
FilterCondition.ANDorORto combine multiple constraints, and operators likeEQ,IN, orEXISTSfor precise matching.
Frequently Asked Questions
What operators are supported in OpenDeRisk metadata filters?
OpenDeRisk supports standard comparison operators defined in FilterOperator within packages/derisk-core/src/derisk/storage/vector_store/filters.py. These include equality (EQ), membership (IN), existence checks (EXISTS), and comparison operators for numeric or string fields. The specific availability depends on the underlying vector store implementation in IndexStoreBase.
Can I combine multiple metadata filters with OR logic instead of AND?
Yes. Set the condition field in your MetadataFilters object to FilterCondition.OR instead of FilterCondition.AND. When using OR logic, the retriever returns chunks that satisfy any one of the specified filters rather than requiring all conditions to match. This is useful for searching across multiple categories or sources simultaneously.
How does OpenDeRisk handle missing filter values?
When a MetadataFilter has an empty value list, the Service.aget_metadata_filter method automatically invokes the TagsExtractor LLM helper. This process extracts candidate values from the knowledge space metadata, prompts the LLM to select relevant tags based on the query context, and populates the filter values before retrieval execution. If the LLM cannot determine appropriate values, the filter may be skipped or return no results depending on the operator.
Where is the metadata filter validation performed?
Validation occurs in Service.check_metadata_filters within packages/derisk-serve/src/derisk_serve/rag/service/service.py (lines 5049-5082). This method ensures that filter structures are valid, operators are supported, and values are present or can be generated before the retrieval pipeline executes. Invalid filters raise errors before reaching the vector store layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →