# Architecture of the multi_report_agent for Video Analysis Reports: NVIDIA VSS Agent Design

> Explore the multi report agent architecture for streaming video analysis and unified markdown reports. Understand the NVIDIA VSS Agent design and its three layer NAT framework.

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: architecture
- Published: 2026-05-15

---

**The multi_report_agent is a deterministic, streaming workflow in the NVIDIA-AI-Blueprints/video-search-and-summarization repository that consolidates multiple video-analysis incidents into unified markdown reports through a three-layer NAT framework architecture.**

The multi_report_agent operates within the NVIDIA AI Toolkit (NAT) framework to transform scattered video-analysis events into coherent incident reports. Implemented in the open-source NVIDIA-AI-Blueprints/video-search-and-summarization repository, this agent follows a strict separation of concerns between input validation, configuration management, and execution orchestration. Its architecture enables both synchronous result generation and streaming LLM-compatible output through structured message chunks.

## Three-Layer Architecture Overview

The architecture cleanly separates declaration from execution across three distinct layers defined in [`agent/src/vss_agents/agents/multi_report_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/agents/multi_report_agent.py).

### Input Validation Layer

The **Input Validation Layer** enforces schema constraints on incoming requests. The `MultiReportAgentInput` Pydantic model (lines 43-60) validates the `source` identifier, `source_type`, optional ISO8601 time ranges, and ensures positive integers for `max_result_size`.

### Configuration Layer

The **Configuration Layer** manages static defaults and external tool references. `MultiReportAgentConfig` extends `FunctionBaseConfig` (lines 68-84) to store the `multi_incident_tool` reference and bounds `max_incidents` between `[1, 10000]`, preventing unbounded result sets.

### Execution Layer

The **Execution Layer** handles the streaming workflow. The `multi_report_agent` async generator (lines 87-130) orchestrates tool resolution, invocation, and output translation into `AgentOutput` structures while yielding `AgentMessageChunk` objects for real-time LLM consumption.

## Detailed Execution Flow

The deterministic workflow follows six sequential stages that transform raw video events into polished reports.

### 1. Tool Resolution

Using the NAT `Builder`, the agent resolves the `multi_incident_formatter` tool (or any custom implementation supplied via `MultiReportAgentConfig.multi_incident_tool`) with the LANGCHAIN wrapper:

```python
multi_incident_tool = await builder.get_tool(
    config.multi_incident_tool, wrapper_type=LLMFrameworkEnum.LANGCHAIN
)

```

### 2. Effective Size Determination

The request may optionally specify `max_result_size`; if omitted, the agent falls back to the static `config.max_incidents`. This ensures a bounded result set while still allowing UI-level pagination:

```python
effective_max_size = max_result_size if max_result_size is not None else config.max_incidents

```

### 3. Tool Invocation and Streaming

The agent yields a **tool-call** chunk (`AgentMessageChunkType.TOOL_CALL`) that logs the arguments for visibility in the LLM's reasoning trace. It then awaits the formatter's async `ainvoke` method:

```python
yield AgentMessageChunk(
    type=AgentMessageChunkType.TOOL_CALL,
    content=f"Tool: multi_incident_formatter\nArgs: {tool_args}"
)
formatter_result = await multi_incident_tool.ainvoke(tool_args)

```

### 4. Result Normalisation

The formatter may return either a rich dataclass (with `formatted_incidents`, `total_incidents`, `chart_html`) or a plain dict. The agent normalises these into:

- `formatted_incidents` – Markdown-ready incident list
- `incident_count` – Total incidents fetched for summary and metrics
- `side_effects` – Optional `chart_html` that downstream consumers can embed

This tolerant handling makes the agent robust to future format changes.

### 5. Final Agent Output

A structured `AgentOutput` is built, containing human-readable messages, side-effects (e.g., generated charts), status (`success`/`error`), and rich metadata (source, source_type, report_type, generation latency, `max_result_size`). The output is emitted as a **final** chunk (`AgentMessageChunkType.FINAL`) with JSON payload for downstream processing:

```python
agent_output = AgentOutput(
    messages=[
      f"Found {incident_count} incident{'s' if incident_count != 1 else ''} for {source_type} {source}",
      formatted_incidents,
    ],
    side_effects=side_effects,
    status="success",
    metadata={ ... },
)
yield AgentMessageChunk(type=AgentMessageChunkType.FINAL,
                        content=agent_output.model_dump_json())

```

### 6. Error Handling

The agent catches `ValueError`, `KeyError`, and `AttributeError` to surface validation or tool-specific failures, and a generic `Exception` block to guard against unexpected crashes. In each case, a failure-oriented `AgentOutput` is streamed with `status="error"` and an explanatory `error_message`.

## Configuration and Usage Examples

The following examples demonstrate how to construct requests, register the agent in NAT workflows, and interpret the streaming output.

### Constructing a Request Payload

```python
from vss_agents.agents.multi_report_agent import MultiReportAgentInput

payload = MultiReportAgentInput(
    source="sensor-42",
    source_type="sensor",
    start_time="2025-09-01T00:00:00.000Z",
    end_time="2025-09-01T23:59:59.999Z",
    max_result_size=200,
)

```

### Registering the Agent in a NAT Workflow

```python
from nat.builder.builder import Builder
from vss_agents.agents.multi_report_agent import MultiReportAgentConfig, multi_report_agent

builder = Builder()                     # NAT builder instance

config = MultiReportAgentConfig(
    multi_incident_tool="multi_incident_formatter",  # could be a custom tool name

    max_incidents=5000,
)

# The `multi_report_agent` function returns an async generator that can be streamed

async for chunk in multi_report_agent(config, builder):
    if chunk.type == AgentMessageChunkType.TOOL_CALL:
        print("LLM sees tool call:", chunk.content)
    elif chunk.type == AgentMessageChunkType.FINAL:
        result = json.loads(chunk.content)   # <-- final AgentOutput JSON

        print("Report generated:", result)

```

### Interpreting the Final AgentOutput

```python
import json
from vss_agents.agents.data_models import AgentOutput

final_json = '{"messages":["Found 12 incidents …","<markdown>…"],"side_effects":{"chart_html":"<svg>…"},"status":"success","metadata":{"incident_count":12,"source":"sensor-42","source_type":"sensor","report_type":"multi_incident","generation_time_ms":143,"max_result_size":200}}'

output = AgentOutput.model_validate_json(final_json)
print(output.messages[0])        # Human-readable summary

print(output.side_effects["chart_html"])  # Embedded chart HTML

```

## Key Implementation Files

These files define the deterministic, stream-oriented nature of the Multi-Report Agent:

- **[`agent/src/vss_agents/agents/multi_report_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/agents/multi_report_agent.py)** – Core implementation containing `MultiReportAgentInput`, `MultiReportAgentConfig`, and the `multi_report_agent` async generator function.
- **[`agent/src/vss_agents/agents/data_models.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/agents/data_models.py)** – Definitions of `AgentMessageChunk`, `AgentMessageChunkType`, and `AgentOutput` used by the streaming protocol.
- **[`agent/tests/unit_test/agents/test_multi_report_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/tests/unit_test/agents/test_multi_report_agent.py)** – Unit tests validating schema constraints and default behaviors for the agent.
- **[`agent/src/vss_agents/tools/multi_incident_formatter.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/tools/multi_incident_formatter.py)** – The downstream formatter tool that supplies the incident list and optional chart HTML (when present).

## Summary

- The **multi_report_agent** implements a three-layer architecture separating input validation (`MultiReportAgentInput`), configuration (`MultiReportAgentConfig`), and execution (async generator).
- It leverages the **NAT framework** for tool resolution, specifically using `LLMFrameworkEnum.LANGCHAIN` wrappers to abstract formatter implementations.
- The streaming protocol emits **tool-call visibility chunks** followed by **final JSON payloads**, enabling real-time LLM interaction and downstream processing.
- **Defensive normalization** handles both dataclass and dictionary returns from formatter tools, ensuring backward compatibility.
- Bounded result sets are enforced through configurable `max_incidents` limits (default 10000) with optional UI-level override via `max_result_size`.

## Frequently Asked Questions

### What is the primary role of the multi_report_agent in video analysis?

The multi_report_agent synthesizes multiple discrete video-analysis incidents into a single, consolidated markdown report. According to the NVIDIA-AI-Blueprints/video-search-and-summarization source code, it acts as an orchestration layer that coordinates with formatter tools to transform raw event data into human-readable summaries with optional chart visualizations.

### How does the multi_report_agent handle streaming output?

The agent implements an async generator pattern that yields `AgentMessageChunk` objects defined in [`agent/src/vss_agents/agents/data_models.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/agents/data_models.py). It first emits a `TOOL_CALL` type chunk to expose the formatter invocation to the LLM's reasoning trace, then emits a `FINAL` type chunk containing the complete `AgentOutput` JSON. This design allows LLMs to observe tool usage while receiving structured final results.

### Can the multi_report_agent work with custom incident formatter tools?

Yes. The `MultiReportAgentConfig` class accepts a `multi_incident_tool` parameter that specifies the tool name to resolve via the NAT `Builder`. While the default references `"multi_incident_formatter"`, operators can inject custom implementations as long as they conform to the async `ainvoke` interface and return either dataclass objects with `formatted_incidents` and `chart_html` attributes or equivalent dictionaries.

### What error handling mechanisms protect the multi_report_agent workflow?

The agent implements tiered exception handling in the execution layer (lines 87-130 of [`multi_report_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/multi_report_agent.py)). It specifically catches `ValueError`, `KeyError`, and `AttributeError` to handle validation and schema mismatches, plus a generic `Exception` block for unexpected failures. All error paths stream an `AgentOutput` with `status="error"` and a descriptive `error_message`, ensuring clients receive structured failure notifications rather than raw stack traces.