Long Video Summarization Chunking and Aggregation: How the LVS Microservice Processes Long-Form Video

The LVS microservice splits long videos into configurable temporal chunks using chunk_duration and chunk_overlap_duration parameters, processes each segment through a VLM, and aggregates results using the summary_aggregation_prompt to produce a unified summary with deduplicated timestamps.

The Long Video Summarization (LVS) microservice within the NVIDIA-AI-Blueprints/video-search-and-summarization repository handles hour-long videos through a two-stage pipeline that chunks content before aggregating results. This architecture prevents token overflow while maintaining temporal continuity across segment boundaries. Understanding how the service configures chunk boundaries and merges per-chunk outputs is essential for optimizing both accuracy and latency in production deployments.

Configuring Video Chunking Parameters

The chunking strategy is controlled through explicit request parameters sent to the /summarize endpoint. These parameters determine whether the video is processed whole or sliced into manageable segments.

Chunk Duration and Overlap

The primary mechanism for controlling segmentation is the chunk_duration field, specified in seconds. According to the implementation in agent/src/vss_agents/tools/lvs_video_understanding.py, setting this value to 0 disables chunking entirely, forcing the microservice to process the complete video in a single VLM call. When chunk_duration is greater than 0, the microservice creates fixed-size temporal slices of the specified length.

To prevent events from being split across chunk boundaries, the optional chunk_overlap_duration parameter adds overlapping seconds between consecutive chunks. This overlap ensures that events occurring at segment edges appear in full within at least one chunk, preserving contextual continuity during the aggregation phase.

Before submitting a summarization request, agents can query the POST /recommended_config endpoint to determine optimal chunk sizing. This endpoint accepts the total video length and a target response time, then returns a computed chunk_size value that balances processing speed against VLM context limits. As documented in skills/video-summarization/references/lvs-api.md, this allows clients to dynamically adjust chunk duration based on video characteristics rather than using static defaults.

The Aggregation Pipeline

After the VLM processes individual chunks, the microservice must reconcile multiple partial outputs into a coherent global summary.

Response Schema and Structure

The LVS microservice returns responses following the OpenAI chat completions schema. The choices[0].message.content field contains a JSON payload with two primary fields: a video_summary string providing the narrative overview, and an events array containing timestamped occurrences. The aggregation logic in agent/src/vss_agents/tools/lvs_video_understanding.py parses this content structure, extracting both elements to construct the final result. If both fields return empty, the tool injects a default note indicating no events were detected.

Prompt-Driven Merge Logic

When multiple chunks are processed, the VLM uses the summary_aggregation_prompt defined in agent/src/vss_agents/prompt.py to concatenate per-chunk captions. This prompt instructs the model to remove duplicate or overlapping timestamps while maintaining chronological order. The aggregation also respects the num_frames_per_chunk parameter, which controls frame sampling density within each segment before the merge operation occurs.

For live streams, the summary_duration parameter similarly governs how much temporal content each VLM invocation receives, with the aggregation prompt ensuring continuity between streaming windows.

Implementation in the Agent Tool

The lvs_video_understanding.py tool demonstrates practical implementation of these concepts through its request construction and response handling logic.

Request Construction

The tool builds the request payload in the lvs_request dictionary, incorporating human-in-the-loop parameters collected via _collect_hitl_parameters. This includes scenario descriptions, target events, and objects of interest, alongside the technical chunking parameters. The implementation constructs an async HTTP POST to {BASE_URL}/summarize, transmitting the configuration including chunk_duration, chunk_overlap_duration, and num_frames_per_chunk.

Response Parsing

Upon receiving the VLM response, the tool extracts the content string and parses it as JSON. The logic checks for the presence of video_summary and events fields, handling edge cases where the VLM returns empty results. This parsing occurs within the response handling block that processes the OpenAI-formatted return from the LVS backend.

Code Examples

The following examples demonstrate configuring chunking parameters when calling the LVS microservice.

Calling the summarization endpoint with explicit chunk sizing:

BASE_URL="http://localhost:38111"
API_KEY="YOUR_TOKEN"

MODEL=$(curl -s "$BASE_URL/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id')

curl -s -X POST "$BASE_URL/v1/summarize" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "'"$MODEL"'",
    "scenario": "warehouse",
    "events": ["safety violation"],
    "url": "https://example.com/warehouse.mp4",
    "chunk_duration": 60,
    "chunk_overlap_duration": 10,
    "prompt": "Summarize the video with timestamps."
  }' | jq '.choices[0].message.content'

Programmatic configuration using the agent tool:

from vss_agents.tools.lvs_video_understanding import lvs_video_understanding, LVSVideoUnderstandingConfig
from nat.builder.builder import Builder

config = LVSVideoUnderstandingConfig(
    lvs_backend_url="http://localhost:38111",
    model="cosmos-reason1",
    chunk_duration=60,               # 1‑minute chunks

    num_frames_per_chunk=20,
    hitl_scenario_template="Describe the video scenario:",
    hitl_events_template="List events to detect:",
    hitl_objects_template="List objects of interest (or type 'skip'):",
)

async def run():
    async for info in lvs_video_understanding(config, Builder()):
        summary = await info.single_fn(LVSVideoUnderstandingInput(sensor_id="my_video"))
        print(summary)

# asyncio.run(run())

Summary

  • Chunking is configurable via chunk_duration: Set to 0 to process entire videos whole, or specify seconds to create temporal segments that prevent VLM context overflow.
  • Overlap prevents boundary loss: The chunk_overlap_duration parameter ensures events crossing chunk boundaries appear complete in at least one segment.
  • Dynamic sizing is available: The /recommended_config endpoint calculates optimal chunk sizes based on video length and latency requirements.
  • Aggregation uses prompt-driven merging: The summary_aggregation_prompt in vss_agents/prompt.py controls how per-chunk outputs concatenate and deduplicate.
  • Response handling follows OpenAI schema: Results arrive in choices[0].message.content as JSON containing video_summary and events fields.

Frequently Asked Questions

What happens if I set chunk_duration to 0?

Setting chunk_duration to 0 disables chunking entirely, causing the LVS microservice to send the complete video to the VLM in a single request. This avoids aggregation overhead but risks hitting token limits or timeout thresholds for long-form content.

How does the microservice handle events that occur across chunk boundaries?

The chunk_overlap_duration parameter creates temporal overlap between consecutive chunks, ensuring that events spanning boundary points appear in full within at least one segment. During aggregation, the summary_aggregation_prompt instructs the VLM to detect and remove duplicate timestamps when merging partial results.

Can I calculate the optimal chunk size automatically?

Yes, the LVS microservice exposes a POST /recommended_config endpoint that accepts total video duration and target response time, returning an optimal chunk_size value. The lvs_video_understanding tool can query this endpoint before submitting the actual summarization request to dynamically configure processing parameters.

Where is the aggregation logic defined in the codebase?

The aggregation behavior is controlled by the summary_aggregation_prompt defined in agent/src/vss_agents/prompt.py, which provides instructions for merging per-chunk outputs. The response parsing and final structure assembly occur in agent/src/vss_agents/tools/lvs_video_understanding.py, specifically within the response handling blocks that process the OpenAI-formatted JSON return.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →