# Best Practices for Real-Time Video Alerts with VLM Verification in NVIDIA VSS

> Implement real-time video alerts with VLM verification using NVIDIA VSS. Discover best practices for automatic incident detection and efficient stream processing. Learn to leverage the POST /v1/generate_captions_alerts endpoint...

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: best-practices
- Published: 2026-05-15

---

**Use the `POST /v1/generate_captions_alerts` endpoint with `stream=true`, enable Kafka via the `RTVI_VLM_KAFKA_ENABLED` environment variable, and design prompts that force the VLM to output the tokens "yes" or "true" to trigger automatic incident detection.**

NVIDIA Video Search & Summarization (VSS) 3.1 ships with the **Real-Time Vision-Language Microservice (RT-VLM)**, which generates dense captions and incident alerts from stored video files or live RTSP streams. Implementing **real-time video alerts with VLM verification** requires understanding the automatic token detection mechanism, proper endpoint selection, and reliable Kafka consumption patterns. This guide walks through the architecture, configuration, and production-ready workflows for deterministic alert generation.

## Architecture and Alert Detection Logic

The RT-VLM service exposes REST endpoints under `/v1/*` that decode video, chunk streams, and run Vision-Language Models (VLMs) such as Cosmos-Reason 1/2 or any OpenAI-compatible model. According to the source code in [`skills/rt-vlm/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/rt-vlm/SKILL.md), alerts are **not** toggled per-request. Instead, the service automatically emits an incident when the VLM response contains the tokens **"yes"** or **"true"** (case-insensitive).

Key components include:

- **RT-VLM Service**: Processes video chunks and streams caption deltas via HTTP Server-Sent Events (SSE) when `stream=true` is set.
- **Kafka Bus**: Publishes every caption message to `KAFKA_TOPIC` and incident protobufs to `KAFKA_INCIDENT_TOPIC` when alert tokens are detected.
- **Environment Configuration**: Variables prefixed with `RTVI_VLM_KAFKA_*` enable the message bus and define topic names at container startup.

> **Critical:** You must use the endpoint `POST /v1/generate_captions_alerts` (note the `_alerts` suffix). Using the legacy `/v1/generate_captions` endpoint will silently fail to emit alerts in VSS 3.1 GA releases.

## Configuring the Environment for Alert Generation

Alert emission depends on environment variables set at container launch. These cannot be overridden per request.

```bash
export RTVI_VLM_KAFKA_ENABLED=true
export RTVI_VLM_KAFKA_TOPIC=vision-llm-messages
export RTVI_VLM_KAFKA_INCIDENT_TOPIC=vision-llm-events-incidents
export RTVI_VLM_ERROR_MESSAGE_TOPIC=vision-llm-errors
export HOST_IP=<kafka-host>

```

As documented in [`skills/rt-vlm/references/deploy-rt-vlm-service.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/rt-vlm/references/deploy-rt-vlm-service.md), these settings wire the RT-VLM container to the Kafka broker. If `RTVI_VLM_KAFKA_ENABLED` is false or unset, the incident topic will remain empty regardless of VLM output.

## Prompt Engineering for Reliable Verification

To ensure consistent **VLM verification** for safety incidents or anomalies, constrain the model to a Yes/No output format. The detection logic specifically looks for "yes" or "true" in the response text.

**Recommended prompt structure:**

```text
Anomaly Detected: Yes/No
Reason: <brief explanation>

```

Pair this with a strict system prompt:

```json
{
  "system_prompt": "Answer the user's question with a single Yes or No on the first line."
}

```

This pattern, referenced in [`skills/rt-vlm/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/rt-vlm/SKILL.md) and [`skills/rt-vlm/references/kafka-workflows.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/rt-vlm/references/kafka-workflows.md), maximizes the probability that the VLM emits the trigger tokens when an anomaly is present.

## End-to-End Implementation Workflows

### Processing Stored Video Files

First, upload the video file to obtain a file ID, then initiate caption generation with alerts enabled:

```bash

# Upload video (multipart/form-data)

FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
  -H "Authorization: Bearer $API_KEY" \
  -F "file=@/path/to/warehouse.mp4" \
  -F "purpose=vision" \
  -F "media_type=video" | jq -r '.id')

# Generate captions and alerts via SSE

curl -N -X POST "$BASE_URL/v1/generate_captions_alerts" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "id": "'"$FILE_ID"'",
    "prompt": "Anomaly Detected: Yes/No\nReason: Explain any safety breach.",
    "system_prompt": "Answer with Yes or No on the first line.",
    "model": "cosmos-reason1",
    "chunk_duration": 10,
    "stream": true
  }'

```

Each SSE event contains `start_ts`, `end_ts`, and `content`. When the VLM outputs "yes" or "true", an incident protobuf is automatically sent to the `KAFKA_INCIDENT_TOPIC`.

### Processing Live RTSP Streams

Register the live stream first, then start the inference loop:

```bash

# Register the stream

STREAM_ID=$(curl -fsS -X POST "$BASE_URL/v1/streams/add" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"streams":[{"liveStreamUrl":"rtsp://camera:8554/live","description":"Warehouse cam"}]}' \
  | jq -r '.results[0].id')

# Start real-time captioning with VLM verification

curl -N -X POST "$BASE_URL/v1/generate_captions_alerts" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "id": "'"$STREAM_ID"'",
    "prompt": "Anomaly Detected: Yes/No\nReason: Flag missing hard-hat or vest.",
    "system_prompt": "Answer with Yes or No on the first line.",
    "model": "cosmos-reason2",
    "chunk_duration": 10,
    "stream": true
  }' &

```

**Resource cleanup requires two calls** to prevent duplicate connections:

```bash
curl -X DELETE "$BASE_URL/v1/generate_captions_alerts/$STREAM_ID" \
  -H "Authorization: Bearer $API_KEY"
curl -X DELETE "$BASE_URL/v1/streams/delete/$STREAM_ID" \
  -H "Authorization: Bearer $API_KEY"

```

## Consuming Alerts from Kafka

The Kafka workflow provides durable, asynchronous access to alerts. Verify topic offsets and consume messages as follows:

```bash

# Check topic offsets

docker exec mdx-kafka kafka-get-offsets \
  --bootstrap-server 127.0.0.1:9092 \
  --topic vision-llm-events-incidents

# Consume incident metadata (keys and headers only)

docker exec mdx-kafka kafka-console-consumer \
  --bootstrap-server 127.0.0.1:9092 \
  --topic vision-llm-events-incidents \
  --from-beginning \
  --max-messages 20 \
  --property print.timestamp=true \
  --property print.key=true \
  --property print.headers=true \
  --property print.value=false

```

According to [`skills/rt-vlm/references/kafka-workflows.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/rt-vlm/references/kafka-workflows.md), incident records contain `isAnomaly=true` and `info["triggerPhrase"]` set to the matched token. The message key uses the format `<request_id>:<chunk_idx>`, allowing direct correlation between an incident and its parent caption message.

## Common Pitfalls and Resolution Strategies

| Symptom | Root Cause | Solution |
|---------|------------|----------|
| No alerts despite Yes/No questions | Prompt missing required pattern or `system_prompt` not constraining output | Use the two-line "Anomaly Detected: Yes/No" pattern with a strict system prompt enforcing single-word answers. |
| Alerts published but HTTP client sees nothing | `stream=false` (default) buffers until video completion | Set `stream=true` and use an SSE-compatible client like `curl -N`. |
| Empty `vision-llm-events-incidents` topic | `RTVI_VLM_KAFKA_ENABLED` was false at startup | Restart the RT-VLM container with `RTVI_VLM_KAFKA_ENABLED=true`. |
| Duplicate RTSP connections after shutdown | Only deleted `/v1/generate_captions_alerts/{id}`, not the stream registration | Call both DELETE endpoints: `/v1/generate_captions_alerts/{id}` and `/v1/streams/delete/{id}`. |
| Model OOM on long videos | `chunk_duration=0` sends the entire file as one request | Set `chunk_duration` to 10 seconds (or similar) to limit context window size. |

## Summary

- **Enable Kafka at deployment** by setting `RTVI_VLM_KAFKA_ENABLED=true` and configuring the three topic environment variables.
- **Use the correct endpoint**: `POST /v1/generate_captions_alerts` is required for alert emission in VSS 3.1.
- **Design deterministic prompts** that output "yes" or "true" on a dedicated line, paired with a restrictive `system_prompt`.
- **Stream continuously** by setting `stream=true` for real-time SSE updates and immediate Kafka publication.
- **Clean up resources properly** by deleting both the caption generation session and the stream registration for RTSP sources.
- **Consume from `vision-llm-events-incidents`** using the message key format `<request_id>:<chunk_idx>` to correlate alerts with specific video segments.

## Frequently Asked Questions

### Why am I not receiving alerts even though the VLM answers "Yes"?

The RT-VLM service performs case-insensitive token detection on the words "yes" or "true". If your prompt allows explanatory text before the answer (e.g., "There is a problem: Yes"), the token may not appear in the expected format. Force the VLM to output "Yes" or "No" on the first line using a strict `system_prompt` and verify that `RTVI_VLM_KAFKA_ENABLED` was set to `true` before the container started, as this cannot be changed at request time.

### What is the difference between `/v1/generate_captions` and `/v1/generate_captions_alerts`?

The `/v1/generate_captions` endpoint generates dense captions but does not inspect the VLM output for alert tokens. Only the `/v1/generate_captions_alerts` endpoint (available in VSS 3.1 and later) implements the token detection logic that publishes to `vision-llm-events-incidents`. Using the former will result in captions without the accompanying alert infrastructure.

### How do I correlate a Kafka incident with the original video chunk?

Each Kafka message in the incident topic uses a key formatted as `<request_id>:<chunk_idx>`, where `request_id` matches the caption generation session and `chunk_idx` indicates the temporal segment. The incident protobuf also includes `isAnomaly=true` and metadata about the trigger phrase, allowing you to join the alert with the specific SSE event or caption message from the `vision-llm-messages` topic.

### Can I change the alert trigger words from "yes"/"true" to custom tokens?

No. As implemented in the NVIDIA VSS 3.1 RT-VLM service, the token detection logic is hardcoded to recognize only "yes" and "true" (case-insensitive). To use custom alerting criteria, you must consume the raw caption stream from Kafka or HTTP SSE and implement custom filtering logic in your downstream application.