# How the RTVI VLM Alert Tool Processes Video Streams: A Technical Deep Dive

> Explore how the RTVI VLM alert tool processes video streams discover its real-time captioning and anomaly alert mechanisms via SSE and Kafka.

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: deep-dive
- Published: 2026-05-15

---

**The RTVI VLM alert tool processes video streams by first registering an RTSP source to obtain a unique stream identifier, then establishing a Server-Sent Events (SSE) connection to deliver real-time captions, while simultaneously publishing anomaly alerts to a dedicated Kafka topic for downstream consumers.**

The NVIDIA AI Blueprints video search and summarization repository provides a Real-Time Vision-Language Microservice (RTVI VLM) that enables dense captioning and incident detection on live video feeds. Understanding how this microservice ingests, analyzes, and distributes stream data is essential for building production video analytics pipelines that require sub-second latency and reliable event notification. The implementation details are specified in [`skills/rt-vlm/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/rt-vlm/SKILL.md), which defines the REST API contract and Kafka integration patterns.

## Stream Registration and RTSP Source Initialization

Every streaming session begins with registering the live video source. According to [`skills/rt-vlm/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/rt-vlm/SKILL.md) at line 441, clients must POST to `/v1/streams/add` with a JSON payload containing the RTSP URL and optional metadata.

### Registering the Live Stream

The registration endpoint requires the `liveStreamUrl` field (an `rtsp://` URI) and accepts additional parameters such as description, credentials, and location info. The service returns a **UUID** that uniquely identifies the stream for all subsequent operations.

```bash

# Register a live RTSP stream and capture the returned UUID

STREAM_ID=$(curl -fsS -X POST "$BASE_URL/v1/streams/add" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"streams":[{"liveStreamUrl":"rtsp://10.0.0.5:8554/warehouse","description":"warehouse cam"}]}' \
  | jq -r '.results[0].id')

```

## Real-Time Caption Generation and SSE Streaming

Once the stream is registered, clients invoke the `/v1/generate_captions_alerts` endpoint to begin analysis. As documented at line 71 in [`skills/rt-vlm/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/rt-vlm/SKILL.md), this endpoint accepts parameters that configure the VLM's output style and reasoning capabilities.

### Configuring the Caption Request

The request body must include the stream UUID (`id`), a text `prompt` describing the desired caption format, and the `stream=true` flag to enable Server-Sent Events. Optional flags such as `enable_reasoning`, `chunk_duration`, and `system_prompt` allow fine-tuning of the generation behavior.

```bash

# Start dense-caption + alert generation with SSE streaming

curl -N -X POST "$BASE_URL/v1/generate_captions_alerts" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
        \"id\": \"$STREAM_ID\",
        \"prompt\": \"Describe each event; start each sentence with a timestamp.\",
        \"model\": \"cosmos-reason1\",
        \"chunk_duration\": 10,
        \"stream\": true
      }"

```

### Processing the SSE Response

With `stream=true`, the service returns an **SSE** stream where each event contains a JSON payload. These payloads include the caption content (`content`) along with precise timestamps (`start_ts`, `end_ts`) indicating the temporal boundaries of the described video segment. When the VLM detects an incident, it produces anomaly tokens that trigger the Kafka publishing workflow.

## Kafka Alert Pipeline and Incident Publishing

Beyond the SSE caption stream, the RTVI VLM alert tool publishes structured incident notifications to a **Kafka** topic for asynchronous processing. This decouples real-time viewing from alerting and logging systems.

### NvSchema Protobuf Alert Structure

As specified at line 60 in [`skills/rt-vlm/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/rt-vlm/SKILL.md), alerts are published to the `vision-llm-events-incidents` topic by default. The payload uses an **NvSchema protobuf** format containing fields such as `sensorId`, `timestamp`, `isAnomaly`, and `info["triggerPhrase"]`. This binary encoding ensures efficient transmission of metadata-rich alerts.

### Consuming Alerts from the Topic

Downstream consumers can read these alerts using standard Kafka client tools. When using the console consumer inside the VSS container, disable value printing to avoid dumping the binary protobuf data to the terminal, as noted at lines 65–73.

```bash

# Consume alerts from Kafka (run inside the VSS container)

docker exec mdx-kafka kafka-console-consumer \
  --bootstrap-server 127.0.0.1:9092 \
  --topic vision-llm-events-incidents \
  --from-beginning \
  --timeout-ms 5000 \
  --max-messages 10 \
  --property print.timestamp=true \
  --property print.key=true \
  --property print.headers=true \
  --property print.value=false

```

## Stream Lifecycle Management and Teardown

Proper resource cleanup requires a two-step teardown process documented at lines 14–18 in [`skills/rt-vlm/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/rt-vlm/SKILL.md). First, stop the active inference to halt caption generation, then unregister the stream to release the RTSP connection.

### Stopping Inference and Releasing Resources

The `DELETE /v1/generate_captions_alerts/{stream_id}` endpoint stops the caption generation process but maintains the stream registration. To fully release the RTSP source and free associated resources, you must subsequently call `DELETE /v1/streams/delete/{stream_id}`.

```bash

# Stop captioning and unregister the stream

curl -X DELETE "$BASE_URL/v1/generate_captions_alerts/$STREAM_ID" -H "Authorization: Bearer $API_KEY"
curl -X DELETE "$BASE_URL/v1/streams/delete/$STREAM_ID" -H "Authorization: Bearer $API_KEY"

```

## Summary

- **Stream registration** via `/v1/streams/add` establishes a persistent UUID for the RTSP source, enabling subsequent API calls to reference the live feed.
- **Real-time processing** uses Server-Sent Events from `/v1/generate_captions_alerts` with `stream=true`, delivering timestamped caption chunks as the VLM processes video segments.
- **Anomaly detection** triggers automatic publishing of NvSchema protobuf messages to the `vision-llm-events-incidents` Kafka topic, allowing downstream systems to react to incidents immediately.
- **Resource cleanup** requires sequential DELETE requests to stop inference and unregister the stream, preventing resource leaks in long-running deployments.

## Frequently Asked Questions

### What is the RTVI VLM alert tool used for?

The RTVI VLM alert tool is a microservice designed for real-time video analysis that generates dense natural language descriptions of live video streams while simultaneously detecting anomalies and publishing alerts. It bridges computer vision and language models to provide human-readable summaries of video content with sub-second latency.

### How do I consume alerts from the Kafka topic without decoding protobuf?

Use the `kafka-console-consumer` with the `--property print.value=false` flag to inspect message headers, keys, and timestamps while suppressing the binary protobuf payload. This allows you to verify alert flow and metadata without requiring a protobuf decoder during debugging.

### Why does stopping the stream require two separate DELETE calls?

The architecture separates inference lifecycle from stream registration to allow pausing and resuming analysis without re-establishing the RTSP connection. Calling `DELETE /v1/generate_captions_alerts/{stream_id}` stops the VLM processing, while `DELETE /v1/streams/delete/{stream_id}` terminates the underlying RTSP feed and frees network resources.

### What data format does the SSE stream return?

The SSE stream returns text/event-stream data where each event contains a JSON object with `content` (the caption text), `start_ts` (segment start time), and `end_ts` (segment end time). This structure allows clients to synchronize captions precisely with the video timeline.