# How the attribute_search Tool Works with Elasticsearch Query Builders in NVIDIA VSS

> Explore how the attribute_search tool uses Elasticsearch query builders for efficient video analysis. Discover its two-stage query pipeline for KNN vector search and result refinement.

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: how-to-guide
- Published: 2026-05-15

---

**The attribute_search tool executes a two-stage Elasticsearch query pipeline that first performs a K-nearest-neighbor (KNN) vector search on behavior embeddings to find candidate objects, then optionally refines results with frame-level Painless script scoring to extract precise timestamps and bounding boxes.**

The attribute_search tool in the NVIDIA-AI-Blueprints/video-search-and-summarization repository enables AI agents to locate objects in video streams by their visual attributes through sophisticated Elasticsearch query construction. By combining vector similarity search with structured filtering, the tool bridges the gap between semantic embeddings and time-series video metadata. This implementation follows the same query-building patterns found throughout the VSS codebase, demonstrating how specialized search tools integrate with the core architecture in [`agent/src/vss_agents/video_analytics/query_builders.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/video_analytics/query_builders.py).

## Two-Stage Query Architecture

The attribute_search tool processes visual queries through two distinct Elasticsearch operations orchestrated by the `search_by_attributes` function in [`agent/src/vss_agents/tools/attribute_search.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/tools/attribute_search.py). 

First, `_search_behavior` retrieves candidate objects using vector similarity on behavior embeddings. Second, `_get_frame_from_behavior` optionally performs a scripted score query to pinpoint the exact frame, timestamp, and bounding box for each candidate. Both stages reuse the query-building philosophy from [`query_builders.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/query_builders.py): deep-copying a base query template and incrementally adding filters.

## Stage 1: Behavior-Level KNN Search

### Input Configuration Models

The tool accepts structured parameters through `AttributeSearchInput`, which defines user-facing parameters such as query text, time windows, video sources, `top_k`, and similarity thresholds. Service endpoints and index names are managed by `AttributeSearchConfig`. These models validate inputs before query construction begins, ensuring vector dimensions and index names align with the Elasticsearch cluster configuration.

### Constructing the Vector Search Payload

The `_search_behavior` method builds a KNN query that executes directly within Elasticsearch's vector search capability. The query payload specifies:

- The target embedding field in the behavior index
- The query vector generated by the RTVI CV embedding client
- The `k` parameter for top results and `num_candidates` for the approximate nearest neighbor algorithm

This approach allows vector similarity computation to occur alongside metadata filtering in a single request, minimizing network round-trips.

### Integrating Temporal and Source Filters

Following the pattern established in [`query_builders.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/query_builders.py), the tool constructs `filter_clauses` for time ranges and video source wildcards, then injects them into the KNN payload's `filter` parameter. The final search body combines the KNN payload with `_source` field specifications to retrieve only the necessary metadata for subsequent processing.

```python

# Conceptual query structure built by _search_behavior

search_query = {
    "knn": {
        "field": "embedding",
        "query_vector": query_embedding,
        "k": top_k,
        "num_candidates": 100,
        "filter": {
            "bool": {
                "must": [
                    {"range": {"timestamp": {"gte": start_time, "lte": end_time}}},
                    {"wildcard": {"source_id": source_pattern}}
                ]
            }
        }
    },
    "_source": ["sensor_id", "object_id", "object_type", "timestamp"]
}

```

## Stage 2: Frame-Level Refinement

### Painless Script for Cosine Similarity

When `enable_frame_lookup` is true (the default behavior), the tool executes `_get_frame_from_behavior` for each candidate object. This function constructs a query containing a **Painless script** that iterates over the `objects` array in frame documents. The script extracts the specific object's embedding vector using its ID, computes the dot-product and L2 norms, and returns a normalized cosine similarity score.

### Query Body and Result Processing

The frame lookup query combines a `bool` filter (matching sensor ID, timestamp range, and nested `objects.id`) with a `script_score` function containing the Painless implementation. After execution, the tool extracts the best-matching frame ID, timestamp, and bounding box coordinates for each candidate object.

### Parallel Execution Strategy

The tool leverages `asyncio.gather` to execute frame lookup queries concurrently for all behavior candidates. This parallel approach significantly reduces latency when processing multiple objects, returning a list of `(frame_id, bbox, frame_score, timestamp)` tuples that feed into the final result assembly.

## Result Assembly and Deduplication

The `_build_result` function merges behavior-level metadata with frame-level data, creating `AttributeSearchResult` objects containing normalized timestamps, bounding boxes, and dual scores (`behavior_score` and `frame_score`). 

The process then applies two final transformations:

- **Deduplication**: The `_deduplicate_by_object` method removes redundant entries where the same object appears multiple times across different segments, filtering by unique sensor and object ID combinations.
- **Exclusion filtering**: Any videos listed in `exclude_videos` are removed from the result set.
- **Top-k limiting**: The final ranked list is truncated to the requested `top_k` results.

## Alignment with Query Builder Patterns

The attribute_search tool adheres to the architectural standards defined in [`agent/src/vss_agents/video_analytics/query_builders.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/video_analytics/query_builders.py). Like `IncidentQueryBuilder` and `BehaviorQueryBuilder`, it conceptually starts from `BASE_QUERY_TEMPLATE` imported from [`es_client.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/es_client.py) and layers additional constraints into the `bool.must` clause.

While generic builders focus on term and range queries for incidents and frames, attribute_search extends this pattern to vector search by embedding the KNN payload within the same filtering framework. This consistency ensures that modifications to base templates or common filter handling affect all tools uniformly across the repository.

## Practical Implementation Example

```python
import asyncio
from datetime import datetime, timedelta
from vss_agents.tools.attribute_search import AttributeSearchInput, search_attributes
from vss_agents.embed.rtvi_cv_embed import RTVICVEmbedClient
from elasticsearch import AsyncElasticsearch

async def find_red_hat_person():
    """Search for a person wearing a red hat in recent video footage."""
    embed_client = RTVICVEmbedClient(endpoint="http://localhost:9000")
    es_client = AsyncElasticsearch(["http://localhost:9200"])
    
    search_input = AttributeSearchInput(
        query="person with red hat",
        source_type="rtsp",
        timestamp_start=datetime.utcnow() - timedelta(seconds=30),
        timestamp_end=datetime.utcnow(),
        top_k=5,
        min_similarity=0.4,
    )
    
    results = await search_attributes(
        search_input=search_input,
        embed_client=embed_client,
        es_client=es_client,
        index="mdx-behavior-2026-01-06",
        vst_external_url="http://vst.example.com",
        frames_index="mdx-raw-2026-01-09",
    )
    
    for result in results:
        print(f"Object: {result.metadata.object_type}")
        print(f"Timestamp: {result.metadata.frame_timestamp}")
        print(f"Bounding Box: {result.metadata.bbox}")

asyncio.run(find_red_hat_person())

```

This example triggers both the initial KNN search (`_search_behavior`) and the subsequent frame-level lookups (`_get_frame_from_behavior`), returning structured metadata with precise spatial and temporal coordinates.

## Summary

- The attribute_search tool implements a two-stage query pipeline: KNN vector search on behavior embeddings followed by scripted frame scoring for exact matches.
- Behavior-level queries combine vector similarity with temporal and source filters using the same `bool.must` patterns as generic query builders.
- Frame-level refinement uses Painless scripts to calculate exact cosine similarity against stored embeddings in the frames index.
- The tool follows VSS repository conventions by leveraging `BASE_QUERY_TEMPLATE` patterns and async execution via `asyncio.gather`.
- Results undergo deduplication by object ID and exclusion filtering before returning the final ranked list with complete metadata.

## Frequently Asked Questions

### What is the difference between behavior-level and frame-level search in attribute_search?

The behavior-level search performs approximate KNN on behavior index embeddings to rapidly identify candidate objects matching the semantic query across large time windows. The frame-level search then executes an exact cosine similarity calculation using Painless scripts against specific object embeddings within the frames index to identify the precise timestamp and bounding box coordinates for each candidate.

### How does the tool integrate filters with vector similarity queries?

The tool constructs standard Elasticsearch `bool` filter clauses for time ranges and video sources, then embeds them directly into the KNN payload's `filter` parameter. This ensures vector similarity calculations only occur on documents matching the metadata constraints, maintaining consistency with the filtering approach used in `IncidentQueryBuilder` and `BehaviorQueryBuilder`.

### Why does attribute_search use a Painless script for frame lookup instead of another KNN query?

Frame lookup requires exact cosine similarity calculation against a specific object ID within a nested `objects` array. The Painless script iterates through this array, extracts the matching vector by ID, and computes similarity precisely using dot-products and norms. This allows per-frame scoring of specific objects, which wouldn't be possible with standard KNN queries that operate on entire document embeddings without nested field isolation.

### How does the query construction in attribute_search relate to query_builders.py?

While [`query_builders.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/query_builders.py) provides generic builders for incidents and frames using term and range queries, `attribute_search` applies the same architectural philosophy—starting from `BASE_QUERY_TEMPLATE` and building `bool.must` clauses—but specializes it for vector search. Instead of simple match queries, it constructs KNN payloads with embedded filters, extending the repository's uniform query-building approach to semantic video search.