# How the embed_search Tool Performs Multi-Embedding Fusion for Video Archives

> Discover how embed_search fuses multiple embeddings for video archives. It uses text, image, or video inputs, Elasticsearch KNN, and fusion strategies for optimal retrieval.

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: deep-dive
- Published: 2026-05-15

---

**The embed_search tool performs multi-embedding fusion by generating query vectors from text, image, or video inputs, retrieving candidate matches via Elasticsearch KNN search, and optionally fusing embedding similarity scores with attribute-based scores using weighted linear or reciprocal rank fusion strategies.**

The NVIDIA AI Blueprints video search and summarization repository implements a sophisticated retrieval architecture for video archives through the `embed_search` tool. This component handles multiple embedding modalities and performs advanced score fusion to deliver accurate, context-aware video segment ranking. The implementation spans three logical stages: query vector generation, Elasticsearch nearest-neighbor lookup, and multi-factor relevance fusion.

## Query Embedding Generation

The fusion process begins with [`embed_search.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/embed_search.py), where the `_generate_query_embedding` function transforms heterogeneous inputs into a unified query vector. The tool accepts text queries, image URLs, video URLs, or pre-computed embedding arrays via the `_str_input_converter` and `_chat_request_input_converter` parsers.

Depending on the input type, the `CosmosEmbedClient` (defined in [`cosmos_embed.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/cosmos_embed.py)) dispatches to specific endpoints:

- **Text**: `get_text_embedding` calls `/v1/generate_text_embeddings`
- **Image**: `get_image_embedding` encodes and sends to `/v1/generate_image_embeddings`  
- **Video**: `get_video_embedding` batches URLs via `get_video_embeddings_from_urls` to `/v1/generate_video_embeddings`

When an `embeddings` array is provided directly, the pipeline extracts the first vector, ensuring only one modality drives the initial retrieval phase.

## Elasticsearch KNN Retrieval

With the query vector prepared, the `_build_es_query` function (lines 224+) constructs a nested KNN query targeting the `llm.visionEmbeddings.vector` field in Elasticsearch. The implementation dynamically scales the `k` parameter based on applied filters and similarity thresholds (lines 441-447 in [`embed_search.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/embed_search.py)).

The query structure supports optional boolean filters for `video_sources`, description text, and timestamp ranges, allowing users to constrain the vector search space before fusion occurs. The async Elasticsearch client executes this query against the video archive index, returning candidate documents containing stored vision embeddings and metadata.

## Multi-Embedding Fusion Pipeline

When **attribute search** is enabled via the `use_attribute_search` flag in the higher-level `search` tool, the pipeline transitions from pure vector retrieval to multi-embedding fusion. This stage involves three critical transformations:

### Similarity Score Normalization

The `_process_search_hit` function (line 398 in [`embed_search.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/embed_search.py)) converts raw Elasticsearch scores into interpretable cosine similarity values using the transformation `round(2 * _score - 1, 2)`. Each hit becomes an `EmbedSearchResultItem` containing the normalized embedding similarity, video metadata, and temporal boundaries.

### Attribute Search Integration

The `fusion_search_rerank` function in [`search.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/search.py) (line 487) orchestrates the fusion workflow. For each embedding-based result, it executes an attribute search to retrieve object IDs and attribute-specific similarity scores. These attribute scores undergo normalization through averaging across all detected attributes, producing a `normalised_attribute_score` between 0 and 1.

### Fusion Strategy Implementation

The system implements three distinct fusion algorithms selected via the `fusion_method` parameter in `SearchAgentConfig`:

**Weighted Linear Fusion** (`_apply_weighted_linear_fusion`, line 355):

```python
fused_score = w_embed * embed_score + w_attribute * attribute_score

```

**Reciprocal Rank Fusion (RRF)** (`_apply_rrf_fusion`, line 401):

```python
fused_score = 1/(rank_embed + k) + w * normalised_attribute_score

```

**RRF with Attribute Rank**:

This variant combines both embedding and attribute ranks in the reciprocal rank formula, accommodating scenarios where attribute-based ordering provides complementary signals to vector similarity.

The fused score replaces the original embedding similarity as the final ranking metric, with results re-sorted before returning the `EmbedSearchOutput` to the caller.

## Implementation Architecture

The fusion workflow follows this execution path through the codebase:

1. **Input Parsing**: `embed_search` receives requests through standardized converters, extracting `query`, `image_url`, `video_url`, or `embeddings` parameters
2. **Vector Generation**: `CosmosEmbedClient` generates the query embedding via NVIDIA Cosmos endpoints
3. **Initial Retrieval**: `_build_es_query` constructs the nested KNN query with optional filters
4. **Hit Processing**: `_process_search_hit` extracts metadata and converts ES scores to cosine similarity
5. **Fusion Execution**: `fusion_search_rerank` in [`search.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/search.py) combines embedding and attribute scores using the configured strategy
6. **Result Ranking**: Final `SearchResult` objects are sorted by the fused similarity score

Key configuration resides in [`search_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/search_agent.py), which wires `embed_search` and `attribute_search` together and exposes the `fusion_method` selection to higher-level agents.

## Summary

- The `embed_search` tool in [`agent/src/vss_agents/tools/embed_search.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/tools/embed_search.py) serves as the primary entry point for semantic video search, supporting text, image, video, and pre-computed vector inputs
- Query embeddings are generated through `CosmosEmbedClient` and matched against video archives using Elasticsearch KNN queries targeting `llm.visionEmbeddings.vector`
- Raw Elasticsearch scores undergo normalization via the formula `2 * _score - 1` to produce cosine similarity metrics
- Multi-embedding fusion occurs through `fusion_search_rerank` in [`search.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/search.py), which integrates attribute-based retrieval scores with embedding similarity
- Three fusion strategies are available: **weighted linear fusion**, **reciprocal rank fusion (RRF)**, and **RRF with attribute rank**, configured via `SearchAgentConfig`
- The final output provides a unified relevance ranking that combines semantic vector similarity with structured attribute metadata

## Frequently Asked Questions

### What embedding modalities does the embed_search tool support?

The tool supports four input modalities processed through `CosmosEmbedClient`: text queries (via `get_text_embedding`), image URLs (via `get_image_embedding`), video URLs (via `get_video_embedding`), and pre-computed embedding vectors supplied directly in the request payload. Only one modality is used per query to generate the initial search vector.

### How does the fusion_search_rerank function combine scores?

The `fusion_search_rerank` function (line 487 in [`search.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/search.py)) retrieves attribute scores for each embedding result, normalizes them by averaging across detected attributes, and applies one of three fusion algorithms. The function takes the embedding similarity score and the normalized attribute score as inputs, applying either weighted linear combination or reciprocal rank fusion to produce a final relevance score.

### What is the difference between weighted linear fusion and RRF?

**Weighted linear fusion** (`_apply_weighted_linear_fusion`) calculates the final score as a simple weighted sum: `w_embed * embed_score + w_attribute * attribute_score`. **Reciprocal Rank Fusion (RRF)** (`_apply_rrf_fusion`) instead uses rank-based reciprocal values: `1/(rank_embed + k) + w * normalised_attribute_score`, which tends to be more robust when comparing scores from different distributions or scales. RRF with attribute rank extends this by incorporating both embedding and attribute ranks into the reciprocal calculation.

### Where is the fusion method configured in the codebase?

The fusion method is configured in [`search_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/search_agent.py) through the `SearchAgentConfig` class, which exposes the `fusion_method` parameter. Valid options include `"weighted_linear"`, `"rrf"`, and `"rrf_with_attribute_rank"`. This configuration determines which fusion strategy `fusion_search_rerank` applies when combining embedding and attribute search results.