# How to Configure Video Analytics Elasticsearch Integration in VSS

> Easily configure video analytics Elasticsearch integration in VSS. Set environment variables and use ESClient for asynchronous queries to incidents and frame data. Get started today.

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: how-to-guide
- Published: 2026-05-15

---

**Configure video analytics Elasticsearch integration in VSS by setting the `ELASTICSEARCH_URL`, connection retry parameters, and vector dimension environment variables, then use the `ESClient` class to query incidents and frame data asynchronously.**

The NVIDIA AI Blueprints Video Search & Summarization (VSS) repository provides a robust asynchronous interface for connecting to Elasticsearch. To configure video analytics Elasticsearch integration in VSS correctly, you must align environment variables with the embedding dimensions used by your RTVI and Vision-LLM models while initializing the `ESClient` wrapper for safe index access.

## Core Components

### ESClient Async Wrapper

The integration centers on the `ESClient` class defined in [`agent/src/vss_agents/video_analytics/es_client.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/video_analytics/es_client.py). This wrapper extends `AsyncElasticsearch` to provide production-safe defaults and a whitelist-based security model. The class exposes four primary helper methods—`search()`, `aggregate()`, `get_by_id()`, and `count()`—that validate all requests against a hardcoded index whitelist (lines 34‑41) before execution.

### Index Whitelist Security

VSS maps logical index names (e.g., `incidents`, `frames`, `behavior`) to physical Elasticsearch index patterns through an internal whitelist. Attempting to query an unlisted index raises a clear `ValueError`, preventing accidental cross-tenant data access or typo-induced queries against non-existent shards.

### Vector Field Configuration

Video analytics data relies on dense vectors for semantic search. The integration requires two specific dimension variables:
- **`ELASTICSEARCH_RTVI_CV_EMBEDDINGS_DIM`**: Dimension for RTVI (real-time video intelligence) frame embeddings (default: 1536)
- **`ELASTICSEARCH_VISION_LLM_EMBEDDINGS_DIM`**: Dimension for Vision-LLM embeddings (default: 768)

These values must match the output dimensions of the models feeding data into Elasticsearch; otherwise, indexing operations will fail with mapping conflicts.

## Environment Variable Configuration

Set these variables in your deployment `.env` files (typically located in `deployments/**/.env`):

1. **`ELASTICSEARCH_URL`** – The full endpoint address (default: `http://localhost:9200`).
2. **`ELASTICSEARCH_CONNECTION_MAX_ATTEMPTS`** – Retry count for initial connection bootstrap before aborting deployment.
3. **`ELASTICSEARCH_RTVI_CV_EMBEDDINGS_DIM`** – Must equal the RTVI model embedding size.
4. **`ELASTICSEARCH_VISION_LLM_EMBEDDINGS_DIM`** – Must equal the Vision-LLM embedding size.
5. **`ELASTICSEARCH_ILM_MIN_AGE`** – (Optional) Index Lifecycle Management threshold for automatic shard deletion (e.g., `4h`).

The `ESClient` constructor accepts an optional `index_prefix` parameter, allowing multi-tenant deployments to isolate indices by prepending tenant identifiers (e.g., `tenant-a-incidents`).

## Implementation Example

Typical usage within a VSS video-analytics tool follows this pattern:

```python
import os
from copy import deepcopy
from vss_agents.video_analytics.es_client import ESClient
from vss_agents.video_analytics.utils import BASE_QUERY_TEMPLATE

# 1️⃣ Initialise the client with tenant isolation

es = ESClient(es_url=os.getenv("ELASTICSEARCH_URL"), index_prefix="production-")

# 2️⃣ Build a filtered query using the base template

query = deepcopy(BASE_QUERY_TEMPLATE)
query["query"]["bool"]["must"].append({"match": {"camera_id": "cam_001"}})

# 3️⃣ Execute search against whitelisted 'incidents' index

results = await es.search(
    index_key="incidents",
    query_body=query,
    size=50,
    sort="timestamp:desc",
    source_includes=["timestamp", "severity", "description"],
)

# 4️⃣ Run aggregations (e.g., count by severity)

agg_body = {"terms": {"field": "severity"}}
agg_result = await es.aggregate(
    "incidents", 
    query_body=query, 
    aggs={"by_severity": agg_body}
)

# 5️⃣ Clean up connections

await es.close()

```

All VSS analytics tools import `ESClient` from the same module, ensuring consistent connection pooling and error handling across the codebase.

## Deployment Initialization Scripts

Before the application starts, bootstrap scripts in `deployments/foundational/elk/init-scripts/` configure the Elasticsearch cluster:

- **[`elasticsearch-template-creation.sh`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/elasticsearch-template-creation.sh)** – Creates index templates referencing `${ELASTICSEARCH_RTVI_CV_EMBEDDINGS_DIM}` and `${ELASTICSEARCH_VISION_LLM_EMBEDDINGS_DIM}` to define dense_vector field mappings.
- **[`elasticsearch-ilm-policy-creation.sh`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/elasticsearch-ilm-policy-creation.sh)** – Applies Index Lifecycle Management policies using `${ELASTICSEARCH_ILM_MIN_AGE}` to automate rollover and deletion of aged analytics data.

These scripts run during Docker Compose startup, ensuring the cluster schema matches the application's expected vector dimensions before `ESClient` attempts to write data.

## Summary

- **Environment-driven configuration**: Set `ELASTICSEARCH_URL`, retry attempts, and vector dimensions in `.env` files before deployment.
- **Dimension alignment**: Ensure `ELASTICSEARCH_RTVI_CV_EMBEDDINGS_DIM` and `ELASTICSEARCH_VISION_LLM_EMBEDDINGS_DIM` match your model output sizes (1536 and 768 by default).
- **Secure access**: The `ESClient` whitelist in [`agent/src/vss_agents/video_analytics/es_client.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/video_analytics/es_client.py) restricts queries to approved indices only.
- **Multi-tenant support**: Pass `index_prefix` during `ESClient` initialization to isolate tenant data.
- **Automated lifecycle**: Use `ELASTICSEARCH_ILM_MIN_AGE` to control automatic deletion of old analytics shards.

## Frequently Asked Questions

### What environment variables are required to configure Elasticsearch in VSS?

At minimum, you must define `ELASTICSEARCH_URL` to point to your cluster endpoint. For production deployments, also set `ELASTICSEARCH_CONNECTION_MAX_ATTEMPTS` for bootstrap resilience and the two embedding dimension variables (`ELASTICSEARCH_RTVI_CV_EMBEDDINGS_DIM` and `ELASTICSEARCH_VISION_LLM_EMBEDDINGS_DIM`) to match your model outputs. Optional variables include `ELASTICSEARCH_ILM_MIN_AGE` for data retention policies.

### How does VSS handle vector dimension mismatches in Elasticsearch?

If the environment variables defining vector dimensions do not match the actual embedding size produced by the RTVI or Vision-LLM models, Elasticsearch will reject indexing requests with mapping exceptions. The initialization scripts in [`deployments/foundational/elk/init-scripts/elasticsearch-template-creation.sh`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/deployments/foundational/elk/init-scripts/elasticsearch-template-creation.sh) use these variables to create index templates; a mismatch between template definitions and incoming data causes pipeline failures.

### What is the purpose of the index whitelist in the ESClient class?

The whitelist in [`agent/src/vss_agents/video_analytics/es_client.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/video_analytics/es_client.py) (lines 34‑41) maps logical keys like `incidents` or `frames` to physical index patterns. This security layer prevents the application from querying arbitrary indices, ensuring that video analytics tools can only access data within their authorized scope. Attempting to use an unlisted key raises a `ValueError` immediately.

### Where are the Elasticsearch index templates defined in the VSS repository?

Index templates are created by shell scripts located in `deployments/foundational/elk/init-scripts/`. Specifically, [`elasticsearch-template-creation.sh`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/elasticsearch-template-creation.sh) generates templates that include dense_vector field definitions using the dimension variables, while [`elasticsearch-ilm-policy-creation.sh`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/elasticsearch-ilm-policy-creation.sh) configures retention policies. These scripts execute during container initialization to prepare the cluster before the application begins ingesting video analytics data.