How the RTVI VLM Alert Tool Processes Video Streams: A Technical Deep Dive
The RTVI VLM alert tool processes video streams by first registering an RTSP source to obtain a unique stream identifier, then establishing a Server-Sent Events (SSE) connection to deliver real-time captions, while simultaneously publishing anomaly alerts to a dedicated Kafka topic for downstream consumers.
The NVIDIA AI Blueprints video search and summarization repository provides a Real-Time Vision-Language Microservice (RTVI VLM) that enables dense captioning and incident detection on live video feeds. Understanding how this microservice ingests, analyzes, and distributes stream data is essential for building production video analytics pipelines that require sub-second latency and reliable event notification. The implementation details are specified in skills/rt-vlm/SKILL.md, which defines the REST API contract and Kafka integration patterns.
Stream Registration and RTSP Source Initialization
Every streaming session begins with registering the live video source. According to skills/rt-vlm/SKILL.md at line 441, clients must POST to /v1/streams/add with a JSON payload containing the RTSP URL and optional metadata.
Registering the Live Stream
The registration endpoint requires the liveStreamUrl field (an rtsp:// URI) and accepts additional parameters such as description, credentials, and location info. The service returns a UUID that uniquely identifies the stream for all subsequent operations.
# Register a live RTSP stream and capture the returned UUID
STREAM_ID=$(curl -fsS -X POST "$BASE_URL/v1/streams/add" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"streams":[{"liveStreamUrl":"rtsp://10.0.0.5:8554/warehouse","description":"warehouse cam"}]}' \
| jq -r '.results[0].id')
Real-Time Caption Generation and SSE Streaming
Once the stream is registered, clients invoke the /v1/generate_captions_alerts endpoint to begin analysis. As documented at line 71 in skills/rt-vlm/SKILL.md, this endpoint accepts parameters that configure the VLM's output style and reasoning capabilities.
Configuring the Caption Request
The request body must include the stream UUID (id), a text prompt describing the desired caption format, and the stream=true flag to enable Server-Sent Events. Optional flags such as enable_reasoning, chunk_duration, and system_prompt allow fine-tuning of the generation behavior.
# Start dense-caption + alert generation with SSE streaming
curl -N -X POST "$BASE_URL/v1/generate_captions_alerts" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"id\": \"$STREAM_ID\",
\"prompt\": \"Describe each event; start each sentence with a timestamp.\",
\"model\": \"cosmos-reason1\",
\"chunk_duration\": 10,
\"stream\": true
}"
Processing the SSE Response
With stream=true, the service returns an SSE stream where each event contains a JSON payload. These payloads include the caption content (content) along with precise timestamps (start_ts, end_ts) indicating the temporal boundaries of the described video segment. When the VLM detects an incident, it produces anomaly tokens that trigger the Kafka publishing workflow.
Kafka Alert Pipeline and Incident Publishing
Beyond the SSE caption stream, the RTVI VLM alert tool publishes structured incident notifications to a Kafka topic for asynchronous processing. This decouples real-time viewing from alerting and logging systems.
NvSchema Protobuf Alert Structure
As specified at line 60 in skills/rt-vlm/SKILL.md, alerts are published to the vision-llm-events-incidents topic by default. The payload uses an NvSchema protobuf format containing fields such as sensorId, timestamp, isAnomaly, and info["triggerPhrase"]. This binary encoding ensures efficient transmission of metadata-rich alerts.
Consuming Alerts from the Topic
Downstream consumers can read these alerts using standard Kafka client tools. When using the console consumer inside the VSS container, disable value printing to avoid dumping the binary protobuf data to the terminal, as noted at lines 65–73.
# Consume alerts from Kafka (run inside the VSS container)
docker exec mdx-kafka kafka-console-consumer \
--bootstrap-server 127.0.0.1:9092 \
--topic vision-llm-events-incidents \
--from-beginning \
--timeout-ms 5000 \
--max-messages 10 \
--property print.timestamp=true \
--property print.key=true \
--property print.headers=true \
--property print.value=false
Stream Lifecycle Management and Teardown
Proper resource cleanup requires a two-step teardown process documented at lines 14–18 in skills/rt-vlm/SKILL.md. First, stop the active inference to halt caption generation, then unregister the stream to release the RTSP connection.
Stopping Inference and Releasing Resources
The DELETE /v1/generate_captions_alerts/{stream_id} endpoint stops the caption generation process but maintains the stream registration. To fully release the RTSP source and free associated resources, you must subsequently call DELETE /v1/streams/delete/{stream_id}.
# Stop captioning and unregister the stream
curl -X DELETE "$BASE_URL/v1/generate_captions_alerts/$STREAM_ID" -H "Authorization: Bearer $API_KEY"
curl -X DELETE "$BASE_URL/v1/streams/delete/$STREAM_ID" -H "Authorization: Bearer $API_KEY"
Summary
- Stream registration via
/v1/streams/addestablishes a persistent UUID for the RTSP source, enabling subsequent API calls to reference the live feed. - Real-time processing uses Server-Sent Events from
/v1/generate_captions_alertswithstream=true, delivering timestamped caption chunks as the VLM processes video segments. - Anomaly detection triggers automatic publishing of NvSchema protobuf messages to the
vision-llm-events-incidentsKafka topic, allowing downstream systems to react to incidents immediately. - Resource cleanup requires sequential DELETE requests to stop inference and unregister the stream, preventing resource leaks in long-running deployments.
Frequently Asked Questions
What is the RTVI VLM alert tool used for?
The RTVI VLM alert tool is a microservice designed for real-time video analysis that generates dense natural language descriptions of live video streams while simultaneously detecting anomalies and publishing alerts. It bridges computer vision and language models to provide human-readable summaries of video content with sub-second latency.
How do I consume alerts from the Kafka topic without decoding protobuf?
Use the kafka-console-consumer with the --property print.value=false flag to inspect message headers, keys, and timestamps while suppressing the binary protobuf payload. This allows you to verify alert flow and metadata without requiring a protobuf decoder during debugging.
Why does stopping the stream require two separate DELETE calls?
The architecture separates inference lifecycle from stream registration to allow pausing and resuming analysis without re-establishing the RTSP connection. Calling DELETE /v1/generate_captions_alerts/{stream_id} stops the VLM processing, while DELETE /v1/streams/delete/{stream_id} terminates the underlying RTSP feed and frees network resources.
What data format does the SSE stream return?
The SSE stream returns text/event-stream data where each event contains a JSON object with content (the caption text), start_ts (segment start time), and end_ts (segment end time). This structure allows clients to synchronize captions precisely with the video timeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →