Best Practices for Real-Time Video Alerts with VLM Verification in NVIDIA VSS

Use the POST /v1/generate_captions_alerts endpoint with stream=true, enable Kafka via the RTVI_VLM_KAFKA_ENABLED environment variable, and design prompts that force the VLM to output the tokens "yes" or "true" to trigger automatic incident detection.

NVIDIA Video Search & Summarization (VSS) 3.1 ships with the Real-Time Vision-Language Microservice (RT-VLM), which generates dense captions and incident alerts from stored video files or live RTSP streams. Implementing real-time video alerts with VLM verification requires understanding the automatic token detection mechanism, proper endpoint selection, and reliable Kafka consumption patterns. This guide walks through the architecture, configuration, and production-ready workflows for deterministic alert generation.

Architecture and Alert Detection Logic

The RT-VLM service exposes REST endpoints under /v1/* that decode video, chunk streams, and run Vision-Language Models (VLMs) such as Cosmos-Reason 1/2 or any OpenAI-compatible model. According to the source code in skills/rt-vlm/SKILL.md, alerts are not toggled per-request. Instead, the service automatically emits an incident when the VLM response contains the tokens "yes" or "true" (case-insensitive).

Key components include:

  • RT-VLM Service: Processes video chunks and streams caption deltas via HTTP Server-Sent Events (SSE) when stream=true is set.
  • Kafka Bus: Publishes every caption message to KAFKA_TOPIC and incident protobufs to KAFKA_INCIDENT_TOPIC when alert tokens are detected.
  • Environment Configuration: Variables prefixed with RTVI_VLM_KAFKA_* enable the message bus and define topic names at container startup.

Critical: You must use the endpoint POST /v1/generate_captions_alerts (note the _alerts suffix). Using the legacy /v1/generate_captions endpoint will silently fail to emit alerts in VSS 3.1 GA releases.

Configuring the Environment for Alert Generation

Alert emission depends on environment variables set at container launch. These cannot be overridden per request.

export RTVI_VLM_KAFKA_ENABLED=true
export RTVI_VLM_KAFKA_TOPIC=vision-llm-messages
export RTVI_VLM_KAFKA_INCIDENT_TOPIC=vision-llm-events-incidents
export RTVI_VLM_ERROR_MESSAGE_TOPIC=vision-llm-errors
export HOST_IP=<kafka-host>

As documented in skills/rt-vlm/references/deploy-rt-vlm-service.md, these settings wire the RT-VLM container to the Kafka broker. If RTVI_VLM_KAFKA_ENABLED is false or unset, the incident topic will remain empty regardless of VLM output.

Prompt Engineering for Reliable Verification

To ensure consistent VLM verification for safety incidents or anomalies, constrain the model to a Yes/No output format. The detection logic specifically looks for "yes" or "true" in the response text.

Recommended prompt structure:

Anomaly Detected: Yes/No
Reason: <brief explanation>

Pair this with a strict system prompt:

{
  "system_prompt": "Answer the user's question with a single Yes or No on the first line."
}

This pattern, referenced in skills/rt-vlm/SKILL.md and skills/rt-vlm/references/kafka-workflows.md, maximizes the probability that the VLM emits the trigger tokens when an anomaly is present.

End-to-End Implementation Workflows

Processing Stored Video Files

First, upload the video file to obtain a file ID, then initiate caption generation with alerts enabled:


# Upload video (multipart/form-data)

FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
  -H "Authorization: Bearer $API_KEY" \
  -F "file=@/path/to/warehouse.mp4" \
  -F "purpose=vision" \
  -F "media_type=video" | jq -r '.id')

# Generate captions and alerts via SSE

curl -N -X POST "$BASE_URL/v1/generate_captions_alerts" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "id": "'"$FILE_ID"'",
    "prompt": "Anomaly Detected: Yes/No\nReason: Explain any safety breach.",
    "system_prompt": "Answer with Yes or No on the first line.",
    "model": "cosmos-reason1",
    "chunk_duration": 10,
    "stream": true
  }'

Each SSE event contains start_ts, end_ts, and content. When the VLM outputs "yes" or "true", an incident protobuf is automatically sent to the KAFKA_INCIDENT_TOPIC.

Processing Live RTSP Streams

Register the live stream first, then start the inference loop:


# Register the stream

STREAM_ID=$(curl -fsS -X POST "$BASE_URL/v1/streams/add" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"streams":[{"liveStreamUrl":"rtsp://camera:8554/live","description":"Warehouse cam"}]}' \
  | jq -r '.results[0].id')

# Start real-time captioning with VLM verification

curl -N -X POST "$BASE_URL/v1/generate_captions_alerts" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "id": "'"$STREAM_ID"'",
    "prompt": "Anomaly Detected: Yes/No\nReason: Flag missing hard-hat or vest.",
    "system_prompt": "Answer with Yes or No on the first line.",
    "model": "cosmos-reason2",
    "chunk_duration": 10,
    "stream": true
  }' &

Resource cleanup requires two calls to prevent duplicate connections:

curl -X DELETE "$BASE_URL/v1/generate_captions_alerts/$STREAM_ID" \
  -H "Authorization: Bearer $API_KEY"
curl -X DELETE "$BASE_URL/v1/streams/delete/$STREAM_ID" \
  -H "Authorization: Bearer $API_KEY"

Consuming Alerts from Kafka

The Kafka workflow provides durable, asynchronous access to alerts. Verify topic offsets and consume messages as follows:


# Check topic offsets

docker exec mdx-kafka kafka-get-offsets \
  --bootstrap-server 127.0.0.1:9092 \
  --topic vision-llm-events-incidents

# Consume incident metadata (keys and headers only)

docker exec mdx-kafka kafka-console-consumer \
  --bootstrap-server 127.0.0.1:9092 \
  --topic vision-llm-events-incidents \
  --from-beginning \
  --max-messages 20 \
  --property print.timestamp=true \
  --property print.key=true \
  --property print.headers=true \
  --property print.value=false

According to skills/rt-vlm/references/kafka-workflows.md, incident records contain isAnomaly=true and info["triggerPhrase"] set to the matched token. The message key uses the format <request_id>:<chunk_idx>, allowing direct correlation between an incident and its parent caption message.

Common Pitfalls and Resolution Strategies

Symptom Root Cause Solution
No alerts despite Yes/No questions Prompt missing required pattern or system_prompt not constraining output Use the two-line "Anomaly Detected: Yes/No" pattern with a strict system prompt enforcing single-word answers.
Alerts published but HTTP client sees nothing stream=false (default) buffers until video completion Set stream=true and use an SSE-compatible client like curl -N.
Empty vision-llm-events-incidents topic RTVI_VLM_KAFKA_ENABLED was false at startup Restart the RT-VLM container with RTVI_VLM_KAFKA_ENABLED=true.
Duplicate RTSP connections after shutdown Only deleted /v1/generate_captions_alerts/{id}, not the stream registration Call both DELETE endpoints: /v1/generate_captions_alerts/{id} and /v1/streams/delete/{id}.
Model OOM on long videos chunk_duration=0 sends the entire file as one request Set chunk_duration to 10 seconds (or similar) to limit context window size.

Summary

  • Enable Kafka at deployment by setting RTVI_VLM_KAFKA_ENABLED=true and configuring the three topic environment variables.
  • Use the correct endpoint: POST /v1/generate_captions_alerts is required for alert emission in VSS 3.1.
  • Design deterministic prompts that output "yes" or "true" on a dedicated line, paired with a restrictive system_prompt.
  • Stream continuously by setting stream=true for real-time SSE updates and immediate Kafka publication.
  • Clean up resources properly by deleting both the caption generation session and the stream registration for RTSP sources.
  • Consume from vision-llm-events-incidents using the message key format <request_id>:<chunk_idx> to correlate alerts with specific video segments.

Frequently Asked Questions

Why am I not receiving alerts even though the VLM answers "Yes"?

The RT-VLM service performs case-insensitive token detection on the words "yes" or "true". If your prompt allows explanatory text before the answer (e.g., "There is a problem: Yes"), the token may not appear in the expected format. Force the VLM to output "Yes" or "No" on the first line using a strict system_prompt and verify that RTVI_VLM_KAFKA_ENABLED was set to true before the container started, as this cannot be changed at request time.

What is the difference between /v1/generate_captions and /v1/generate_captions_alerts?

The /v1/generate_captions endpoint generates dense captions but does not inspect the VLM output for alert tokens. Only the /v1/generate_captions_alerts endpoint (available in VSS 3.1 and later) implements the token detection logic that publishes to vision-llm-events-incidents. Using the former will result in captions without the accompanying alert infrastructure.

How do I correlate a Kafka incident with the original video chunk?

Each Kafka message in the incident topic uses a key formatted as <request_id>:<chunk_idx>, where request_id matches the caption generation session and chunk_idx indicates the temporal segment. The incident protobuf also includes isAnomaly=true and metadata about the trigger phrase, allowing you to join the alert with the specific SSE event or caption message from the vision-llm-messages topic.

Can I change the alert trigger words from "yes"/"true" to custom tokens?

No. As implemented in the NVIDIA VSS 3.1 RT-VLM service, the token detection logic is hardcoded to recognize only "yes" and "true" (case-insensitive). To use custom alerting criteria, you must consume the raw caption stream from Kafka or HTTP SSE and implement custom filtering logic in your downstream application.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →