How the embed_search Tool Performs Multi-Embedding Fusion for Video Archives
The embed_search tool performs multi-embedding fusion by generating query vectors from text, image, or video inputs, retrieving candidate matches via Elasticsearch KNN search, and optionally fusing embedding similarity scores with attribute-based scores using weighted linear or reciprocal rank fusion strategies.
The NVIDIA AI Blueprints video search and summarization repository implements a sophisticated retrieval architecture for video archives through the embed_search tool. This component handles multiple embedding modalities and performs advanced score fusion to deliver accurate, context-aware video segment ranking. The implementation spans three logical stages: query vector generation, Elasticsearch nearest-neighbor lookup, and multi-factor relevance fusion.
Query Embedding Generation
The fusion process begins with embed_search.py, where the _generate_query_embedding function transforms heterogeneous inputs into a unified query vector. The tool accepts text queries, image URLs, video URLs, or pre-computed embedding arrays via the _str_input_converter and _chat_request_input_converter parsers.
Depending on the input type, the CosmosEmbedClient (defined in cosmos_embed.py) dispatches to specific endpoints:
- Text:
get_text_embeddingcalls/v1/generate_text_embeddings - Image:
get_image_embeddingencodes and sends to/v1/generate_image_embeddings - Video:
get_video_embeddingbatches URLs viaget_video_embeddings_from_urlsto/v1/generate_video_embeddings
When an embeddings array is provided directly, the pipeline extracts the first vector, ensuring only one modality drives the initial retrieval phase.
Elasticsearch KNN Retrieval
With the query vector prepared, the _build_es_query function (lines 224+) constructs a nested KNN query targeting the llm.visionEmbeddings.vector field in Elasticsearch. The implementation dynamically scales the k parameter based on applied filters and similarity thresholds (lines 441-447 in embed_search.py).
The query structure supports optional boolean filters for video_sources, description text, and timestamp ranges, allowing users to constrain the vector search space before fusion occurs. The async Elasticsearch client executes this query against the video archive index, returning candidate documents containing stored vision embeddings and metadata.
Multi-Embedding Fusion Pipeline
When attribute search is enabled via the use_attribute_search flag in the higher-level search tool, the pipeline transitions from pure vector retrieval to multi-embedding fusion. This stage involves three critical transformations:
Similarity Score Normalization
The _process_search_hit function (line 398 in embed_search.py) converts raw Elasticsearch scores into interpretable cosine similarity values using the transformation round(2 * _score - 1, 2). Each hit becomes an EmbedSearchResultItem containing the normalized embedding similarity, video metadata, and temporal boundaries.
Attribute Search Integration
The fusion_search_rerank function in search.py (line 487) orchestrates the fusion workflow. For each embedding-based result, it executes an attribute search to retrieve object IDs and attribute-specific similarity scores. These attribute scores undergo normalization through averaging across all detected attributes, producing a normalised_attribute_score between 0 and 1.
Fusion Strategy Implementation
The system implements three distinct fusion algorithms selected via the fusion_method parameter in SearchAgentConfig:
Weighted Linear Fusion (_apply_weighted_linear_fusion, line 355):
fused_score = w_embed * embed_score + w_attribute * attribute_score
Reciprocal Rank Fusion (RRF) (_apply_rrf_fusion, line 401):
fused_score = 1/(rank_embed + k) + w * normalised_attribute_score
RRF with Attribute Rank:
This variant combines both embedding and attribute ranks in the reciprocal rank formula, accommodating scenarios where attribute-based ordering provides complementary signals to vector similarity.
The fused score replaces the original embedding similarity as the final ranking metric, with results re-sorted before returning the EmbedSearchOutput to the caller.
Implementation Architecture
The fusion workflow follows this execution path through the codebase:
- Input Parsing:
embed_searchreceives requests through standardized converters, extractingquery,image_url,video_url, orembeddingsparameters - Vector Generation:
CosmosEmbedClientgenerates the query embedding via NVIDIA Cosmos endpoints - Initial Retrieval:
_build_es_queryconstructs the nested KNN query with optional filters - Hit Processing:
_process_search_hitextracts metadata and converts ES scores to cosine similarity - Fusion Execution:
fusion_search_rerankinsearch.pycombines embedding and attribute scores using the configured strategy - Result Ranking: Final
SearchResultobjects are sorted by the fused similarity score
Key configuration resides in search_agent.py, which wires embed_search and attribute_search together and exposes the fusion_method selection to higher-level agents.
Summary
- The
embed_searchtool inagent/src/vss_agents/tools/embed_search.pyserves as the primary entry point for semantic video search, supporting text, image, video, and pre-computed vector inputs - Query embeddings are generated through
CosmosEmbedClientand matched against video archives using Elasticsearch KNN queries targetingllm.visionEmbeddings.vector - Raw Elasticsearch scores undergo normalization via the formula
2 * _score - 1to produce cosine similarity metrics - Multi-embedding fusion occurs through
fusion_search_rerankinsearch.py, which integrates attribute-based retrieval scores with embedding similarity - Three fusion strategies are available: weighted linear fusion, reciprocal rank fusion (RRF), and RRF with attribute rank, configured via
SearchAgentConfig - The final output provides a unified relevance ranking that combines semantic vector similarity with structured attribute metadata
Frequently Asked Questions
What embedding modalities does the embed_search tool support?
The tool supports four input modalities processed through CosmosEmbedClient: text queries (via get_text_embedding), image URLs (via get_image_embedding), video URLs (via get_video_embedding), and pre-computed embedding vectors supplied directly in the request payload. Only one modality is used per query to generate the initial search vector.
How does the fusion_search_rerank function combine scores?
The fusion_search_rerank function (line 487 in search.py) retrieves attribute scores for each embedding result, normalizes them by averaging across detected attributes, and applies one of three fusion algorithms. The function takes the embedding similarity score and the normalized attribute score as inputs, applying either weighted linear combination or reciprocal rank fusion to produce a final relevance score.
What is the difference between weighted linear fusion and RRF?
Weighted linear fusion (_apply_weighted_linear_fusion) calculates the final score as a simple weighted sum: w_embed * embed_score + w_attribute * attribute_score. Reciprocal Rank Fusion (RRF) (_apply_rrf_fusion) instead uses rank-based reciprocal values: 1/(rank_embed + k) + w * normalised_attribute_score, which tends to be more robust when comparing scores from different distributions or scales. RRF with attribute rank extends this by incorporating both embedding and attribute ranks into the reciprocal calculation.
Where is the fusion method configured in the codebase?
The fusion method is configured in search_agent.py through the SearchAgentConfig class, which exposes the fusion_method parameter. Valid options include "weighted_linear", "rrf", and "rrf_with_attribute_rank". This configuration determines which fusion strategy fusion_search_rerank applies when combining embedding and attribute search results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →