How to Integrate Custom VLMs with video_caption and video_understanding Tools
You can integrate custom Vision-Language Models by registering a LangChain-compatible wrapper with the Nat Builder and referencing it in your tool configuration files, requiring no modifications to the core tool source code.
The NVIDIA-AI-Blueprints/video-search-and-summarization repository provides production-ready agent tools for video analysis. Both the video_caption and video_understanding tools dynamically obtain their language models through the Nat Builder abstraction, enabling you to swap in custom VLMs via configuration rather than code changes.
Architecture Overview
The two primary VLM-powered tools obtain their models through the Nat Builder's get_llm method:
video_caption(agent/src/vss_agents/tools/video_caption.py): Generates timestamped captions using either a direct VLM path or the VSS backendvideo_understanding(agent/src/vss_agents/tools/video_understanding.py): Provides free-form answers about video content with optional reasoning chains
Both tools resolve models using LLMRef names. When builder.get_llm(config.llm_name, wrapper_type=LLMFrameworkEnum.LANGCHAIN) is called, the builder returns your registered wrapper instance.
Step 1: Create a LangChain-Compatible VLM Wrapper
Your custom VLM must implement the LangChain asynchronous interface. Specifically, it needs an ainvoke(messages) → BaseMessage method that accepts a list of messages and returns a response.
Implementing the Async Interface
Place your wrapper under agent/src/vss_agents/llms/. The following example demonstrates a minimal HTTP-based VLM that forwards requests to a remote endpoint:
# agent/src/vss_agents/llms/my_vlm.py
from langchain_core.messages import BaseMessage
from typing import Any, List
import httpx
import json
class HttpVLM:
def __init__(self, endpoint: str, api_key: str | None = None):
self.endpoint = endpoint
self.api_key = api_key
async def ainvoke(self, messages: List[BaseMessage]) -> BaseMessage:
payload = {"messages": [msg.dict() for msg in messages]}
headers = {"Authorization": f"Bearer {self.api_key}"} if self.api_key else {}
async with httpx.AsyncClient() as client:
resp = await client.post(
self.endpoint,
json=payload,
headers=headers,
timeout=300
)
resp.raise_for_status()
data = resp.json()
return BaseMessage(content=data["content"])
Step 2: Register the Custom VLM with Nat Builder
Expose your wrapper to the system by registering it in the LLM registry. This creates a named reference that configuration files can use.
Updating the LLM Registry
Add the following to nat/builder/registry.py or an equivalent plugin file:
# nat/builder/registry.py
from nat.builder.llm_registry import register_llm
from vss_agents.llms.my_vlm import HttpVLM
register_llm(
name="my_custom_vlm", # Reference name for configs
factory=lambda cfg: HttpVLM(
endpoint=cfg["endpoint"],
api_key=cfg.get("api_key"),
),
config_schema={
"endpoint": {"type": "string", "required": True},
"api_key": {"type": "string", "optional": True},
},
)
Once registered, my_custom_vlm becomes available as a valid LLMRef throughout the application.
Step 3: Configure video_caption to Use Your VLM
The video_caption tool supports two execution modes controlled by the use_vss parameter:
- VSS backend (
use_vss=True): Frames upload to VSS, which then calls your VLM - Direct VLM (
use_vss=False): The tool samples frames locally and invokes your VLM directly viacall_vlm_partition
For custom HTTP endpoints, the direct path is typically simpler to integrate.
YAML Configuration Example
Create or modify your configuration file to reference the registered VLM:
# config/video_caption.yaml
video_caption:
llm_name: my_custom_vlm # Your registered reference
prompt: |
You are an expert video analyst. Describe the visual content
in detail, noting all objects, people, and activities.
max_frames_per_request: 5
use_vss: false # Forces direct VLM path
fps: 1.0
Runtime Implementation Details
When use_vss is false, the tool executes the direct VLM path implemented in video_caption.py lines [00309‑00380]. The workflow is:
- Video resolution:
resolve_video_filelocates the source video - Frame sampling:
frame_selectextracts frames at the specified FPS - Message construction: Lines [61‑78] build a HumanMessage containing the text prompt and base64-encoded frames
- VLM invocation: The tool calls
llm.ainvoke()using the instance frombuilder.get_llm(config.llm_name)
Step 4: Configure video_understanding
The video_understanding tool follows a similar pattern but uses config.vlm_name instead of llm_name. It resolves the model via:
base_vlm = await builder.get_llm(
config.vlm_name,
wrapper_type=LLMFrameworkEnum.LANGCHAIN
)
This tool handles both ISO-8601 timestamps and offset seconds, supporting complex reasoning chains. Configure it by setting vlm_name to your registered reference:
# config/video_understanding.yaml
video_understanding:
vlm_name: my_custom_vlm
enable_reasoning: true
max_frames: 8
Summary
- Create a wrapper class implementing the LangChain
ainvokeinterface and place it inagent/src/vss_agents/llms/ - Register the wrapper using
register_llm()innat/builder/registry.pyto create a namedLLMRef - Configure
video_captionby settingllm_nameto your reference anduse_vss: falsefor direct VLM invocation - Configure
video_understandingby settingvlm_nameto the same reference - Verify that frame encoding (lines [61‑78]) and the direct invocation path (lines [309‑380]) in
video_caption.pysupport your VLM's expected input format
Frequently Asked Questions
Do I need to modify video_caption.py or video_understanding.py to use a custom VLM?
No. Both tools obtain their models through the Nat Builder abstraction using builder.get_llm(). As long as your VLM is registered with a unique name in the LLM registry, you only need to update your YAML configuration files to reference that name.
What interface must my custom VLM implement?
Your wrapper must implement the LangChain async interface: specifically, an ainvoke(self, messages: List[BaseMessage]) -> BaseMessage method. The tool will pass a list containing text prompts and base64-encoded image frames, expecting a BaseMessage response with the generated content.
Can I use the VSS backend with a custom VLM?
Yes. If you set use_vss: true in your video_caption configuration, frames upload to the VSS backend service, which then calls your VLM. However, this requires the VSS backend to be configured to recognize and route to your custom model endpoint, whereas use_vss: false routes directly from the agent tool to your VLM.
How do I configure authentication for my custom VLM's API?
Pass sensitive parameters through the config_schema when registering your LLM. The registry captures values from your YAML configuration (such as api_key), and the factory lambda instantiates your wrapper class with these values. Store keys in environment variables or secure vaults referenced by your configuration files rather than hardcoding them.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →