Architecture of the multi_report_agent for Video Analysis Reports: NVIDIA VSS Agent Design
The multi_report_agent is a deterministic, streaming workflow in the NVIDIA-AI-Blueprints/video-search-and-summarization repository that consolidates multiple video-analysis incidents into unified markdown reports through a three-layer NAT framework architecture.
The multi_report_agent operates within the NVIDIA AI Toolkit (NAT) framework to transform scattered video-analysis events into coherent incident reports. Implemented in the open-source NVIDIA-AI-Blueprints/video-search-and-summarization repository, this agent follows a strict separation of concerns between input validation, configuration management, and execution orchestration. Its architecture enables both synchronous result generation and streaming LLM-compatible output through structured message chunks.
Three-Layer Architecture Overview
The architecture cleanly separates declaration from execution across three distinct layers defined in agent/src/vss_agents/agents/multi_report_agent.py.
Input Validation Layer
The Input Validation Layer enforces schema constraints on incoming requests. The MultiReportAgentInput Pydantic model (lines 43-60) validates the source identifier, source_type, optional ISO8601 time ranges, and ensures positive integers for max_result_size.
Configuration Layer
The Configuration Layer manages static defaults and external tool references. MultiReportAgentConfig extends FunctionBaseConfig (lines 68-84) to store the multi_incident_tool reference and bounds max_incidents between [1, 10000], preventing unbounded result sets.
Execution Layer
The Execution Layer handles the streaming workflow. The multi_report_agent async generator (lines 87-130) orchestrates tool resolution, invocation, and output translation into AgentOutput structures while yielding AgentMessageChunk objects for real-time LLM consumption.
Detailed Execution Flow
The deterministic workflow follows six sequential stages that transform raw video events into polished reports.
1. Tool Resolution
Using the NAT Builder, the agent resolves the multi_incident_formatter tool (or any custom implementation supplied via MultiReportAgentConfig.multi_incident_tool) with the LANGCHAIN wrapper:
multi_incident_tool = await builder.get_tool(
config.multi_incident_tool, wrapper_type=LLMFrameworkEnum.LANGCHAIN
)
2. Effective Size Determination
The request may optionally specify max_result_size; if omitted, the agent falls back to the static config.max_incidents. This ensures a bounded result set while still allowing UI-level pagination:
effective_max_size = max_result_size if max_result_size is not None else config.max_incidents
3. Tool Invocation and Streaming
The agent yields a tool-call chunk (AgentMessageChunkType.TOOL_CALL) that logs the arguments for visibility in the LLM's reasoning trace. It then awaits the formatter's async ainvoke method:
yield AgentMessageChunk(
type=AgentMessageChunkType.TOOL_CALL,
content=f"Tool: multi_incident_formatter\nArgs: {tool_args}"
)
formatter_result = await multi_incident_tool.ainvoke(tool_args)
4. Result Normalisation
The formatter may return either a rich dataclass (with formatted_incidents, total_incidents, chart_html) or a plain dict. The agent normalises these into:
formatted_incidents– Markdown-ready incident listincident_count– Total incidents fetched for summary and metricsside_effects– Optionalchart_htmlthat downstream consumers can embed
This tolerant handling makes the agent robust to future format changes.
5. Final Agent Output
A structured AgentOutput is built, containing human-readable messages, side-effects (e.g., generated charts), status (success/error), and rich metadata (source, source_type, report_type, generation latency, max_result_size). The output is emitted as a final chunk (AgentMessageChunkType.FINAL) with JSON payload for downstream processing:
agent_output = AgentOutput(
messages=[
f"Found {incident_count} incident{'s' if incident_count != 1 else ''} for {source_type} {source}",
formatted_incidents,
],
side_effects=side_effects,
status="success",
metadata={ ... },
)
yield AgentMessageChunk(type=AgentMessageChunkType.FINAL,
content=agent_output.model_dump_json())
6. Error Handling
The agent catches ValueError, KeyError, and AttributeError to surface validation or tool-specific failures, and a generic Exception block to guard against unexpected crashes. In each case, a failure-oriented AgentOutput is streamed with status="error" and an explanatory error_message.
Configuration and Usage Examples
The following examples demonstrate how to construct requests, register the agent in NAT workflows, and interpret the streaming output.
Constructing a Request Payload
from vss_agents.agents.multi_report_agent import MultiReportAgentInput
payload = MultiReportAgentInput(
source="sensor-42",
source_type="sensor",
start_time="2025-09-01T00:00:00.000Z",
end_time="2025-09-01T23:59:59.999Z",
max_result_size=200,
)
Registering the Agent in a NAT Workflow
from nat.builder.builder import Builder
from vss_agents.agents.multi_report_agent import MultiReportAgentConfig, multi_report_agent
builder = Builder() # NAT builder instance
config = MultiReportAgentConfig(
multi_incident_tool="multi_incident_formatter", # could be a custom tool name
max_incidents=5000,
)
# The `multi_report_agent` function returns an async generator that can be streamed
async for chunk in multi_report_agent(config, builder):
if chunk.type == AgentMessageChunkType.TOOL_CALL:
print("LLM sees tool call:", chunk.content)
elif chunk.type == AgentMessageChunkType.FINAL:
result = json.loads(chunk.content) # <-- final AgentOutput JSON
print("Report generated:", result)
Interpreting the Final AgentOutput
import json
from vss_agents.agents.data_models import AgentOutput
final_json = '{"messages":["Found 12 incidents …","<markdown>…"],"side_effects":{"chart_html":"<svg>…"},"status":"success","metadata":{"incident_count":12,"source":"sensor-42","source_type":"sensor","report_type":"multi_incident","generation_time_ms":143,"max_result_size":200}}'
output = AgentOutput.model_validate_json(final_json)
print(output.messages[0]) # Human-readable summary
print(output.side_effects["chart_html"]) # Embedded chart HTML
Key Implementation Files
These files define the deterministic, stream-oriented nature of the Multi-Report Agent:
agent/src/vss_agents/agents/multi_report_agent.py– Core implementation containingMultiReportAgentInput,MultiReportAgentConfig, and themulti_report_agentasync generator function.agent/src/vss_agents/agents/data_models.py– Definitions ofAgentMessageChunk,AgentMessageChunkType, andAgentOutputused by the streaming protocol.agent/tests/unit_test/agents/test_multi_report_agent.py– Unit tests validating schema constraints and default behaviors for the agent.agent/src/vss_agents/tools/multi_incident_formatter.py– The downstream formatter tool that supplies the incident list and optional chart HTML (when present).
Summary
- The multi_report_agent implements a three-layer architecture separating input validation (
MultiReportAgentInput), configuration (MultiReportAgentConfig), and execution (async generator). - It leverages the NAT framework for tool resolution, specifically using
LLMFrameworkEnum.LANGCHAINwrappers to abstract formatter implementations. - The streaming protocol emits tool-call visibility chunks followed by final JSON payloads, enabling real-time LLM interaction and downstream processing.
- Defensive normalization handles both dataclass and dictionary returns from formatter tools, ensuring backward compatibility.
- Bounded result sets are enforced through configurable
max_incidentslimits (default 10000) with optional UI-level override viamax_result_size.
Frequently Asked Questions
What is the primary role of the multi_report_agent in video analysis?
The multi_report_agent synthesizes multiple discrete video-analysis incidents into a single, consolidated markdown report. According to the NVIDIA-AI-Blueprints/video-search-and-summarization source code, it acts as an orchestration layer that coordinates with formatter tools to transform raw event data into human-readable summaries with optional chart visualizations.
How does the multi_report_agent handle streaming output?
The agent implements an async generator pattern that yields AgentMessageChunk objects defined in agent/src/vss_agents/agents/data_models.py. It first emits a TOOL_CALL type chunk to expose the formatter invocation to the LLM's reasoning trace, then emits a FINAL type chunk containing the complete AgentOutput JSON. This design allows LLMs to observe tool usage while receiving structured final results.
Can the multi_report_agent work with custom incident formatter tools?
Yes. The MultiReportAgentConfig class accepts a multi_incident_tool parameter that specifies the tool name to resolve via the NAT Builder. While the default references "multi_incident_formatter", operators can inject custom implementations as long as they conform to the async ainvoke interface and return either dataclass objects with formatted_incidents and chart_html attributes or equivalent dictionaries.
What error handling mechanisms protect the multi_report_agent workflow?
The agent implements tiered exception handling in the execution layer (lines 87-130 of multi_report_agent.py). It specifically catches ValueError, KeyError, and AttributeError to handle validation and schema mismatches, plus a generic Exception block for unexpected failures. All error paths stream an AgentOutput with status="error" and a descriptive error_message, ensuring clients receive structured failure notifications rather than raw stack traces.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →