How Hybrid Search with RRF Works in Neko Image Gallery: Combining Vision and OCR Embeddings

Neko Image Gallery implements hybrid search by fusing vision and OCR text embeddings using the Reciprocal Rank Fusion (RRF) algorithm via Qdrant's FusionQuery API, merging ranked results from multiple vector spaces into a unified relevance score.

The hv0905/nekoimagegallery repository is an open-source image management system that supports multi-modal search capabilities. When users query using both visual similarity and text extracted via OCR, the system employs hybrid search with RRF to deliver results that consider both modalities simultaneously. This approach avoids the limitations of single-vector searches by combining ranked lists from independent embedding spaces.

Architecture of the Hybrid Search Pipeline

The hybrid search orchestrates multiple embedding types through Qdrant's native fusion capabilities. When a query targets both vision and OCR bases, the VectorDbContext.query_search method constructs separate prefetch queries for each modality before delegating the actual fusion to the vector database.

Building Per-Basis Prefetch Queries

For each search basis in the request, the service creates a Prefetch object targeting the appropriate vector field. In app/Services/vector_db_context.py (lines 141-149), the code maps SearchBasisEnum.vision to the "image_vector" field and SearchBasisEnum.ocr to the "text_contain_vector" field:

prefetches = [
    Prefetch(
        query=self._convert_basis_to_qdrant_query(v),
        using=self.vector_name_for_basis(k),
        filter=filters,
        limit=(top_k + skip) * 2
    )
    for (k, v) in query.criteria.items()
]

The vector_name_for_basis method handles the enum-to-string mapping, while _convert_basis_to_qdrant_query transforms internal DbQueryBasis objects into Qdrant-compatible NearestQuery or RecommendQuery instances. The limit calculation fetches extra results per basis to improve the quality of the final RRF ranking.

Query Transformation Logic

The _convert_basis_to_qdrant_query method processes both single-vector nearest neighbor searches and complex recommendation queries that include positive and negative example vectors. This allows the hybrid system to support advanced search patterns like "find images similar to this photo but without this text content" across multiple embedding spaces simultaneously.

How RRF Calculates the Fusion Score

Reciprocal Rank Fusion combines ranked lists from independent sources using a rank-based formula that prioritizes documents appearing high in multiple lists. Qdrant implements the standard RRF scoring function:

[ \text{score}\text{RRF}(d) = \sum{i=1}^{N} \frac{1}{k + \text{rank}_i(d)} ]

Where N represents the number of search bases (e.g., vision plus OCR), rank_i(d) is the position of document d in the ith result list, and k is a smoothing constant. Qdrant uses a default k = 60, which dampens the impact of absolute rank positions while ensuring that overlap between lists significantly boosts relevance scores.

Implementation in vector_db_context.py

The actual fusion trigger resides in the query_search method of app/Services/vector_db_context.py. Lines 150-156 demonstrate how the service passes the constructed prefetches to Qdrant with a FusionQuery:

result = await self._client.query_points(
    collection_name=self.collection_name,
    prefetch=prefetches,
    query=FusionQuery(fusion=models.Fusion.RRF),
    limit=top_k,
    offset=skip,
    with_payload=True
)

The FusionQuery(fusion=models.Fusion.RRF) parameter instructs Qdrant to execute the RRF algorithm server-side. The database independently runs each prefetch query, calculates reciprocal rank scores for all returned documents, and returns a unified list of ScoredPoint objects ordered by their fused relevance.

Post-Processing the Fused Results

After receiving the RRF-ranked results, lines 172-173 in vector_db_context.py convert the raw Qdrant points into application-specific SearchResult objects:

return [self._get_search_result_from_scored_point(t) for t in result.points]

Each resulting object contains the merged RRF score representing the document's relevance across all specified search bases.

API Entry Point in search.py

The HTTP interface exposing this functionality resides in app/Controllers/search.py. The /hybrid endpoint (lines 173-190) constructs a DbQuery containing both vision and OCR bases, which automatically triggers the RRF code path:

@search_router.post(
    "/hybrid",
    description="Search with hybrid criteria (vision and ocr). Will use RRF algorithm to combine the hybrid results ..."
)
async def hybrid_search(
    model: HybridSearchModel,
    filter_param: Annotated[FilterParams, Depends(FilterParams)],
    paging: Annotated[SearchPagingParams, Depends(SearchPagingParams)]
):
    return await query_and_postprocess(
        query=DbQuery(criteria={
            SearchBasisEnum.vision: process_advanced_search_vectors(...),
            SearchBasisEnum.ocr: process_advanced_search_vectors(...)
        }),
        paging=paging,
        filter_param=filter_param,
    )

When the criteria dictionary contains multiple entries, the downstream query_search method detects the multi-basis condition and activates the prefetch construction and RRF fusion logic.

Practical Usage Examples

Querying via the HTTP API

Send a POST request to the hybrid endpoint with criteria for both vision and OCR bases:

curl -X POST https://your-host/api/search/hybrid \
  -H "Content-Type: application/json" \
  -d '{
        "vision": {
          "criteria": ["cat playing piano"],
          "negative_criteria": []
        },
        "ocr": {
          "criteria": ["musical notes"],
          "negative_criteria": []
        },
        "paging": { "page": 1, "pageSize": 20 }
      }'

The response contains images ranked by their combined relevance to both the visual content and the extracted text.

Direct Python SDK Invocation

For programmatic access within the application stack:

from app.Models.db_queries import DbQuery, DbQueryBasis, DbQueryCriteriaVector
from app.Models.api_models.search_api_model import SearchBasisEnum
from app.Services.vector_db_context import VectorDbContext
from app.Models.query_params import FilterParams

# Construct bases for each modality

vision_basis = DbQueryBasis(
    positive=[DbQueryCriteriaVector(vector=vision_embedding)]
)
ocr_basis = DbQueryBasis(
    positive=[DbQueryCriteriaVector(vector=ocr_embedding)]
)

# Create multi-basis query

query = DbQuery(criteria={
    SearchBasisEnum.vision: vision_basis,
    SearchBasisEnum.ocr: ocr_basis,
})

# Execute hybrid search

ctx = VectorDbContext()
results = await ctx.query_search(
    query=query, 
    top_k=10, 
    skip=0, 
    filter_param=FilterParams()
)

# Results contain RRF-combined scores

for res in results:
    print(f"Image {res.img.id}: RRF score {res.score}")

Summary

  • Multi-basis queries trigger automatic RRF fusion when DbQuery.criteria contains more than one entry.
  • The VectorDbContext.query_search method in app/Services/vector_db_context.py builds Prefetch objects for each basis, targeting specific vector fields ("image_vector" for vision, "text_contain_vector" for OCR).
  • Qdrant's FusionQuery executes the RRF calculation server-side using the formula with default k = 60.
  • The system fetches (top_k + skip) * 2 results per basis to ensure high-quality fusion before applying the final limit.
  • The /api/search/hybrid endpoint in app/Controllers/search.py provides the HTTP interface for combined vision and OCR searches.

Frequently Asked Questions

What is the default k value used in the RRF calculation?

According to the Qdrant implementation used by Neko Image Gallery, the default smoothing constant k = 60. This value balances the influence of high rankings across different result lists, preventing any single modality from dominating the final score purely based on rank position.

The _convert_basis_to_qdrant_query method processes negative criteria by constructing RecommendQuery instances rather than simple NearestQuery vectors. When building the hybrid request, these recommendation queries—including both positive and negative example vectors—are wrapped in Prefetch objects and fused via RRF alongside other bases, allowing the system to exclude undesirable features from the final ranking.

Can hybrid search with RRF be combined with metadata filters?

Yes. The filters parameter passed to each Prefetch object applies user-specified constraints before the RRF fusion occurs. This means the system filters each basis independently using the same FilterParams, then fuses only the filtered result sets. This implementation ensures that metadata constraints like date ranges or tags are respected while still benefiting from cross-modal relevance scoring.

Why does the system fetch (top_k + skip) * 2 results per basis?

Fetching twice the required number of results provides a quality buffer for the RRF algorithm. Since RRF recalculates relevance based on reciprocal ranks rather than raw vector similarity, an item ranked 15th in the vision list but 1st in the OCR list might deserve a top-10 position in the final results. By retrieving extra candidates from each basis, the system ensures that such high-potential documents are available for consideration during the fusion stage.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →