Understanding the video_frame_timestamp Tool and Its Timestamp Mapping in VSS

The video_frame_timestamp tool extracts a specific video frame at a user-defined offset, uses a Vision-Language Model (VLM) to read the embedded timestamp, and returns a Python datetime object mapped to the exact UTC time.

The video_frame_timestamp tool is a core component of the NVIDIA Video Search & Summarization (VSS) blueprint that bridges video frame extraction with vision-language model inference. Located in the NVIDIA-AI-Blueprints/video-search-and-summarization repository, this tool enables precise timestamp extraction from video content using computer vision and AI. It transforms a simple frame offset in seconds into an accurate ISO-8601 datetime object through a sophisticated multi-step pipeline.

How the video_frame_timestamp Tool Works

Input Schema and Configuration

The tool accepts a VideoFrameTimestampInput schema defined in agent/src/vss_agents/tools/video_frame_timestamp.py. This Pydantic model requires two parameters: asset_file_path specifying the video file location, and frame_offset_seconds indicating the temporal position from the video start.

The function is registered to the NAT workflow system via the @register_function decorator (lines 62-63), making it callable as an asynchronous generator that yields a FunctionInfo object for execution.

Frame Extraction Using OpenCV

The implementation uses OpenCV to navigate to the precise frame location. The code converts the user-provided frame_offset_seconds to milliseconds by multiplying by 1000, then seeks to that position using CAP_PROP_POS_MSEC (lines 73-76).


# Conceptual flow based on lines 73-76

cap = cv2.VideoCapture(asset_file_path)
cap.set(cv2.CAP_PROP_POS_MSEC, frame_offset_seconds * 1000)
ret, frame = cap.read()
cap.release()

After positioning the video capture object, it reads a single frame and immediately releases the capture to minimize resource usage.

Vision-Language Model Integration

Once extracted, the frame undergoes JPEG encoding via cv2.imencode followed by Base64 encoding (lines 77-79) to create an inline image payload suitable for LLM consumption. The tool constructs a ChatPromptTemplate using the VIDEO_FRAME_TIMESTAMP_PROMPT constant from agent/src/vss_agents/prompt.py, which explicitly instructs the VLM to return only the timestamp in the format 2024-05-30T01:41:25.000Z without additional commentary (lines 80-95).

The configured LLM (defaulting to openai_llm) is retrieved via the NAT Builder and invoked asynchronously (lines 98-99).

Timestamp Parsing and Mapping

The VLM returns a raw ISO-8601 string, which the tool parses using Python's datetime.strptime with the format specifier "%Y-%m-%dT%H:%M:%S.%fZ" (lines 99-100). This creates a timezone-aware datetime object representing the exact UTC timestamp visible in the video frame.

The mapping completes the transformation from frame offset (seconds) → ISO-8601 string → Python datetime object.

Timestamp Mapping Implementation Details

The timestamp mapping follows a precise conversion chain to ensure frame-accurate temporal alignment:

  1. Input Offset Translation: The tool converts frame_offset_seconds to milliseconds (* 1000) for OpenCV's CAP_PROP_POS_MSEC seek operation.
  2. Visual Recognition: The VLM extracts the embedded timestamp from the frame image and returns it in strict UTC format: YYYY-MM-DDTHH:MM:SS.sssZ.
  3. Python Conversion: The datetime.strptime function parses the string using the "%Y-%m-%dT%H:%M:%S.%fZ" format, producing a timezone-aware datetime object.

This three-stage mapping ensures that the returned timestamp corresponds exactly to the visual frame content at the requested offset, accounting for any timecode burn-in or embedded metadata present in the video itself.

Practical Usage Examples

Direct Invocation of the Registered Function

import asyncio
from vss_agents.tools.video_frame_timestamp import (
    video_frame_timestamp,
    VideoFrameTimestampConfig,
    VideoFrameTimestampInput,
)
from nat.builder.builder import Builder

async def get_timestamp(video_path: str, offset_sec: float):
    cfg = VideoFrameTimestampConfig()
    builder = Builder()
    # The tool yields a FunctionInfo object; we take the first (and only) item.

    gen = video_frame_timestamp(cfg, builder)
    info = await gen.__anext__()          # FunctionInfo

    ts = await info.single_fn(
        VideoFrameTimestampInput(
            asset_file_path=video_path,
            frame_offset_seconds=offset_sec,
        )
    )
    return ts

# Run the coroutine

timestamp = asyncio.run(get_timestamp("samples/demo.mp4", 12.5))
print(timestamp)   # → 2024-05-30 01:41:25+00:00

Parsing the VLM Response into a Datetime Object

from datetime import datetime

# Assume the VLM returned the string "2024-05-30T01:41:25.000Z"

vlm_result = "2024-05-30T01:41:25.000Z"
dt = datetime.strptime(vlm_result, "%Y-%m-%dT%H:%M:%S.%fZ")
print(dt.isoformat())   # 2024-05-30T01:41:25+00:00

Summary

  • The video_frame_timestamp tool in agent/src/vss_agents/tools/video_frame_timestamp.py extracts frames at specific offsets using OpenCV and interprets visual timestamps using a VLM.
  • It maps video frame offsets to Python datetime objects through an intermediate ISO-8601 string format with microsecond precision.
  • The tool relies on the VIDEO_FRAME_TIMESTAMP_PROMPT in agent/src/vss_agents/prompt.py to constrain VLM output to valid timestamp formats only.
  • Implementation details include OpenCV millisecond seeking (CAP_PROP_POS_MSEC), Base64 image encoding, and datetime.strptime parsing with the "%Y-%m-%dT%H:%M:%S.%fZ" format.
  • Unit tests in agent/tests/unit_test/tools/test_video_frame_timestamp.py validate the accuracy of the timestamp extraction and mapping logic.

Frequently Asked Questions

What input parameters does the video_frame_timestamp tool require?

The tool requires a VideoFrameTimestampInput object containing asset_file_path (the video file location) and frame_offset_seconds (the temporal position from video start). These parameters are processed by the video_frame_timestamp function registered in the NAT workflow system to locate and extract the specific frame for VLM analysis.

How does the tool convert video frame offsets to actual timestamps?

The tool converts the frame_offset_seconds to milliseconds and seeks to that position using OpenCV's CAP_PROP_POS_MSEC property (line 74). After extracting the frame, a Vision-Language Model reads the embedded timestamp text, returns it as an ISO-8601 string, and the tool parses it into a Python datetime object using datetime.strptime with the format "%Y-%m-%dT%H:%M:%S.%fZ" (lines 99-100).

What timestamp format does the VLM return?

According to the prompt defined in agent/src/vss_agents/prompt.py, the VLM returns timestamps in strict ISO-8601 UTC format: YYYY-MM-DDTHH:MM:SS.sssZ (for example, 2024-05-30T01:41:25.000Z). This format ensures compatibility with Python's datetime parser and maintains consistent microsecond precision for all VSS operations.

Where can I find the source code and tests for this tool?

The core implementation resides in agent/src/vss_agents/tools/video_frame_timestamp.py, with the system prompt located in agent/src/vss_agents/prompt.py. Comprehensive unit tests validating the timestamp mapping logic are available in agent/tests/unit_test/tools/test_video_frame_timestamp.py and agent/tests/unit_test/tools/test_video_frame_timestamp_coverage.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →