# How to Integrate Custom VLMs with video_caption and video_understanding Tools

> Effortlessly integrate custom VLMs with video_caption and video_understanding tools using Nat Builder. Register compatible wrappers without altering core code. Enhance your video analysis workflow now.

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: how-to-guide
- Published: 2026-05-15

---

**You can integrate custom Vision-Language Models by registering a LangChain-compatible wrapper with the Nat Builder and referencing it in your tool configuration files, requiring no modifications to the core tool source code.**

The NVIDIA-AI-Blueprints/video-search-and-summarization repository provides production-ready agent tools for video analysis. Both the `video_caption` and `video_understanding` tools dynamically obtain their language models through the **Nat Builder** abstraction, enabling you to swap in custom VLMs via configuration rather than code changes.

## Architecture Overview

The two primary VLM-powered tools obtain their models through the Nat Builder's `get_llm` method:

- **`video_caption`** ([`agent/src/vss_agents/tools/video_caption.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/tools/video_caption.py)): Generates timestamped captions using either a direct VLM path or the VSS backend
- **`video_understanding`** ([`agent/src/vss_agents/tools/video_understanding.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/tools/video_understanding.py)): Provides free-form answers about video content with optional reasoning chains

Both tools resolve models using `LLMRef` names. When `builder.get_llm(config.llm_name, wrapper_type=LLMFrameworkEnum.LANGCHAIN)` is called, the builder returns your registered wrapper instance.

## Step 1: Create a LangChain-Compatible VLM Wrapper

Your custom VLM must implement the **LangChain** asynchronous interface. Specifically, it needs an `ainvoke(messages) → BaseMessage` method that accepts a list of messages and returns a response.

### Implementing the Async Interface

Place your wrapper under `agent/src/vss_agents/llms/`. The following example demonstrates a minimal HTTP-based VLM that forwards requests to a remote endpoint:

```python

# agent/src/vss_agents/llms/my_vlm.py

from langchain_core.messages import BaseMessage
from typing import Any, List
import httpx
import json

class HttpVLM:
    def __init__(self, endpoint: str, api_key: str | None = None):
        self.endpoint = endpoint
        self.api_key = api_key

    async def ainvoke(self, messages: List[BaseMessage]) -> BaseMessage:
        payload = {"messages": [msg.dict() for msg in messages]}
        headers = {"Authorization": f"Bearer {self.api_key}"} if self.api_key else {}
        
        async with httpx.AsyncClient() as client:
            resp = await client.post(
                self.endpoint, 
                json=payload, 
                headers=headers, 
                timeout=300
            )
            resp.raise_for_status()
            data = resp.json()
            
        return BaseMessage(content=data["content"])

```

## Step 2: Register the Custom VLM with Nat Builder

Expose your wrapper to the system by registering it in the **LLM registry**. This creates a named reference that configuration files can use.

### Updating the LLM Registry

Add the following to [`nat/builder/registry.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/nat/builder/registry.py) or an equivalent plugin file:

```python

# nat/builder/registry.py

from nat.builder.llm_registry import register_llm
from vss_agents.llms.my_vlm import HttpVLM

register_llm(
    name="my_custom_vlm",               # Reference name for configs

    factory=lambda cfg: HttpVLM(
        endpoint=cfg["endpoint"],
        api_key=cfg.get("api_key"),
    ),
    config_schema={
        "endpoint": {"type": "string", "required": True},
        "api_key": {"type": "string", "optional": True},
    },
)

```

Once registered, `my_custom_vlm` becomes available as a valid `LLMRef` throughout the application.

## Step 3: Configure video_caption to Use Your VLM

The `video_caption` tool supports two execution modes controlled by the `use_vss` parameter:

- **VSS backend** (`use_vss=True`): Frames upload to VSS, which then calls your VLM
- **Direct VLM** (`use_vss=False`): The tool samples frames locally and invokes your VLM directly via `call_vlm_partition`

For custom HTTP endpoints, the direct path is typically simpler to integrate.

### YAML Configuration Example

Create or modify your configuration file to reference the registered VLM:

```yaml

# config/video_caption.yaml

video_caption:
  llm_name: my_custom_vlm                # Your registered reference

  prompt: |
    You are an expert video analyst. Describe the visual content 
    in detail, noting all objects, people, and activities.
  max_frames_per_request: 5
  use_vss: false                         # Forces direct VLM path

  fps: 1.0

```

### Runtime Implementation Details

When `use_vss` is `false`, the tool executes the direct VLM path implemented in [`video_caption.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/video_caption.py) lines [00309‑00380]. The workflow is:

1. **Video resolution**: `resolve_video_file` locates the source video
2. **Frame sampling**: `frame_select` extracts frames at the specified FPS
3. **Message construction**: Lines [61‑78] build a **HumanMessage** containing the text prompt and base64-encoded frames
4. **VLM invocation**: The tool calls `llm.ainvoke()` using the instance from `builder.get_llm(config.llm_name)`

## Step 4: Configure video_understanding

The `video_understanding` tool follows a similar pattern but uses `config.vlm_name` instead of `llm_name`. It resolves the model via:

```python
base_vlm = await builder.get_llm(
    config.vlm_name, 
    wrapper_type=LLMFrameworkEnum.LANGCHAIN
)

```

This tool handles both ISO-8601 timestamps and offset seconds, supporting complex reasoning chains. Configure it by setting `vlm_name` to your registered reference:

```yaml

# config/video_understanding.yaml

video_understanding:
  vlm_name: my_custom_vlm
  enable_reasoning: true
  max_frames: 8

```

## Summary

- **Create** a wrapper class implementing the LangChain `ainvoke` interface and place it in `agent/src/vss_agents/llms/`
- **Register** the wrapper using `register_llm()` in [`nat/builder/registry.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/nat/builder/registry.py) to create a named `LLMRef`
- **Configure** `video_caption` by setting `llm_name` to your reference and `use_vss: false` for direct VLM invocation
- **Configure** `video_understanding` by setting `vlm_name` to the same reference
- **Verify** that frame encoding (lines [61‑78]) and the direct invocation path (lines [309‑380]) in [`video_caption.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/video_caption.py) support your VLM's expected input format

## Frequently Asked Questions

### Do I need to modify video_caption.py or video_understanding.py to use a custom VLM?

No. Both tools obtain their models through the Nat Builder abstraction using `builder.get_llm()`. As long as your VLM is registered with a unique name in the LLM registry, you only need to update your YAML configuration files to reference that name.

### What interface must my custom VLM implement?

Your wrapper must implement the LangChain async interface: specifically, an `ainvoke(self, messages: List[BaseMessage]) -> BaseMessage` method. The tool will pass a list containing text prompts and base64-encoded image frames, expecting a `BaseMessage` response with the generated content.

### Can I use the VSS backend with a custom VLM?

Yes. If you set `use_vss: true` in your `video_caption` configuration, frames upload to the VSS backend service, which then calls your VLM. However, this requires the VSS backend to be configured to recognize and route to your custom model endpoint, whereas `use_vss: false` routes directly from the agent tool to your VLM.

### How do I configure authentication for my custom VLM's API?

Pass sensitive parameters through the `config_schema` when registering your LLM. The registry captures values from your YAML configuration (such as `api_key`), and the factory lambda instantiates your wrapper class with these values. Store keys in environment variables or secure vaults referenced by your configuration files rather than hardcoding them.