How to Integrate Custom VLMs with video_caption and video_understanding Tools

You can integrate custom Vision-Language Models by registering a LangChain-compatible wrapper with the Nat Builder and referencing it in your tool configuration files, requiring no modifications to the core tool source code.

The NVIDIA-AI-Blueprints/video-search-and-summarization repository provides production-ready agent tools for video analysis. Both the video_caption and video_understanding tools dynamically obtain their language models through the Nat Builder abstraction, enabling you to swap in custom VLMs via configuration rather than code changes.

Architecture Overview

The two primary VLM-powered tools obtain their models through the Nat Builder's get_llm method:

Both tools resolve models using LLMRef names. When builder.get_llm(config.llm_name, wrapper_type=LLMFrameworkEnum.LANGCHAIN) is called, the builder returns your registered wrapper instance.

Step 1: Create a LangChain-Compatible VLM Wrapper

Your custom VLM must implement the LangChain asynchronous interface. Specifically, it needs an ainvoke(messages) → BaseMessage method that accepts a list of messages and returns a response.

Implementing the Async Interface

Place your wrapper under agent/src/vss_agents/llms/. The following example demonstrates a minimal HTTP-based VLM that forwards requests to a remote endpoint:


# agent/src/vss_agents/llms/my_vlm.py

from langchain_core.messages import BaseMessage
from typing import Any, List
import httpx
import json

class HttpVLM:
    def __init__(self, endpoint: str, api_key: str | None = None):
        self.endpoint = endpoint
        self.api_key = api_key

    async def ainvoke(self, messages: List[BaseMessage]) -> BaseMessage:
        payload = {"messages": [msg.dict() for msg in messages]}
        headers = {"Authorization": f"Bearer {self.api_key}"} if self.api_key else {}
        
        async with httpx.AsyncClient() as client:
            resp = await client.post(
                self.endpoint, 
                json=payload, 
                headers=headers, 
                timeout=300
            )
            resp.raise_for_status()
            data = resp.json()
            
        return BaseMessage(content=data["content"])

Step 2: Register the Custom VLM with Nat Builder

Expose your wrapper to the system by registering it in the LLM registry. This creates a named reference that configuration files can use.

Updating the LLM Registry

Add the following to nat/builder/registry.py or an equivalent plugin file:


# nat/builder/registry.py

from nat.builder.llm_registry import register_llm
from vss_agents.llms.my_vlm import HttpVLM

register_llm(
    name="my_custom_vlm",               # Reference name for configs

    factory=lambda cfg: HttpVLM(
        endpoint=cfg["endpoint"],
        api_key=cfg.get("api_key"),
    ),
    config_schema={
        "endpoint": {"type": "string", "required": True},
        "api_key": {"type": "string", "optional": True},
    },
)

Once registered, my_custom_vlm becomes available as a valid LLMRef throughout the application.

Step 3: Configure video_caption to Use Your VLM

The video_caption tool supports two execution modes controlled by the use_vss parameter:

  • VSS backend (use_vss=True): Frames upload to VSS, which then calls your VLM
  • Direct VLM (use_vss=False): The tool samples frames locally and invokes your VLM directly via call_vlm_partition

For custom HTTP endpoints, the direct path is typically simpler to integrate.

YAML Configuration Example

Create or modify your configuration file to reference the registered VLM:


# config/video_caption.yaml

video_caption:
  llm_name: my_custom_vlm                # Your registered reference

  prompt: |
    You are an expert video analyst. Describe the visual content 
    in detail, noting all objects, people, and activities.
  max_frames_per_request: 5
  use_vss: false                         # Forces direct VLM path

  fps: 1.0

Runtime Implementation Details

When use_vss is false, the tool executes the direct VLM path implemented in video_caption.py lines [00309‑00380]. The workflow is:

  1. Video resolution: resolve_video_file locates the source video
  2. Frame sampling: frame_select extracts frames at the specified FPS
  3. Message construction: Lines [61‑78] build a HumanMessage containing the text prompt and base64-encoded frames
  4. VLM invocation: The tool calls llm.ainvoke() using the instance from builder.get_llm(config.llm_name)

Step 4: Configure video_understanding

The video_understanding tool follows a similar pattern but uses config.vlm_name instead of llm_name. It resolves the model via:

base_vlm = await builder.get_llm(
    config.vlm_name, 
    wrapper_type=LLMFrameworkEnum.LANGCHAIN
)

This tool handles both ISO-8601 timestamps and offset seconds, supporting complex reasoning chains. Configure it by setting vlm_name to your registered reference:


# config/video_understanding.yaml

video_understanding:
  vlm_name: my_custom_vlm
  enable_reasoning: true
  max_frames: 8

Summary

  • Create a wrapper class implementing the LangChain ainvoke interface and place it in agent/src/vss_agents/llms/
  • Register the wrapper using register_llm() in nat/builder/registry.py to create a named LLMRef
  • Configure video_caption by setting llm_name to your reference and use_vss: false for direct VLM invocation
  • Configure video_understanding by setting vlm_name to the same reference
  • Verify that frame encoding (lines [61‑78]) and the direct invocation path (lines [309‑380]) in video_caption.py support your VLM's expected input format

Frequently Asked Questions

Do I need to modify video_caption.py or video_understanding.py to use a custom VLM?

No. Both tools obtain their models through the Nat Builder abstraction using builder.get_llm(). As long as your VLM is registered with a unique name in the LLM registry, you only need to update your YAML configuration files to reference that name.

What interface must my custom VLM implement?

Your wrapper must implement the LangChain async interface: specifically, an ainvoke(self, messages: List[BaseMessage]) -> BaseMessage method. The tool will pass a list containing text prompts and base64-encoded image frames, expecting a BaseMessage response with the generated content.

Can I use the VSS backend with a custom VLM?

Yes. If you set use_vss: true in your video_caption configuration, frames upload to the VSS backend service, which then calls your VLM. However, this requires the VSS backend to be configured to recognize and route to your custom model endpoint, whereas use_vss: false routes directly from the agent tool to your VLM.

How do I configure authentication for my custom VLM's API?

Pass sensitive parameters through the config_schema when registering your LLM. The registry captures values from your YAML configuration (such as api_key), and the factory lambda instantiates your wrapper class with these values. Store keys in environment variables or secure vaults referenced by your configuration files rather than hardcoding them.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →