How the VSS Agent Architecture Integrates with NVIDIA NIM Microservices

The VSS agent integrates with NVIDIA NIM microservices through environment-driven service discovery, OpenAI-compatible HTTP endpoints, and LangChain-style tool abstractions that route requests to LLM and VLM containers.

The NVIDIA Video Search & Summarization (VSS) blueprint provides a Python-based orchestrator that connects to NVIDIA NIM (NVIDIA Inference Microservices) containers for large language model (LLM) and vision language model (VLM) inference. Understanding how the VSS agent architecture interfaces with these microservices is essential for deploying scalable video analysis pipelines. This article examines the integration patterns found in the NVIDIA-AI-Blueprints/video-search-and-summarization repository, from environment variable configuration to HTTP request handling.

Environment-Driven Service Discovery

The VSS agent relies on environment variables to discover and connect to running NIM containers. During deployment, Docker Compose injects variables such as LLM_BASE_URL, VLM_BASE_URL, and HOST_IP into the agent's runtime environment. These variables point to the network locations of the microservices, allowing the agent to resolve endpoints dynamically without hardcoding addresses.

In deployments/agents/vss-agent/vss-agent-docker-compose.yml, the service definition maps these variables directly to the NIM containers:

VLM_BASE_URL: ${VLM_BASE_URL:-http://${HOST_IP}:${VLM_PORT}}

This pattern enables flexible deployments where operators can run the agent on bare metal while pointing it at remote NIM endpoints hosted on NVIDIA AI Enterprise cloud, or colocate everything via local Docker networking.

OpenAI-Compatible API Contract

Each NIM container exposes an OpenAI-compatible /v1 API, specifically implementing endpoints like POST /v1/chat/completions. The VSS agent never communicates with NIM containers through proprietary protocols; instead, it treats every NIM as an OpenAI-style HTTP service. The agent constructs target URLs by appending /v1 to the base URLs defined in the environment variables.

According to the skill specifications in skills/video-summarization/SKILL.md, short-video requests follow this exact pattern:


POST ${VLM_BASE_URL}/v1/chat/completions

This standardization allows the VSS agent to swap between different NIM models (such as Cosmos Reason 2 or Nemotron Nano) without changing the underlying HTTP client logic.

Agent-Side Tool Abstraction and Orchestration

Tool Wrapping and Reasoning Metadata

Inside the agent codebase, every NIM call is wrapped as a LangChain-style tool (e.g., video_understanding, search_agent). The top-level agent defined in agent/src/vss_agents/agents/top_agent.py builds the request payload, injects the base URL from the environment, and forwards the call to the appropriate NIM. The tool layer additionally inserts reasoning metadata—specifically llm_reasoning and vlm_reasoning parameters—that the NIM consumes to guide inference.

The top_agent.py file handles the orchestration logic that maps user queries to these tool calls, ensuring the correct NIM endpoint receives the properly formatted request body.

Handling Empty Payloads

Certain NIM containers, such as the local Nemotron Nano 9b v2, reject empty HTTP payloads. To prevent request failures, agent/src/vss_agents/agents/top_agent.py implements a safeguard that substitutes a placeholder string ("Agent wants to call tools.") whenever the generated content is empty. This ensures the NIM's HTTP endpoint accepts the request even during intermediate reasoning steps where content is still being constructed.

Configuration and Deployment Wiring

Runtime configuration files bridge the gap between environment variables and the agent's HTTP client. The file agent/app/video_search_frag/configs/config.yml contains entries like:

base_url: ${VLM_BASE_URL}/v1

When the container starts, Docker Compose resolves ${VLM_BASE_URL} to the actual NIM address, allowing the agent's internal client to target the correct microservice immediately. The root-level docker-compose.yml orchestrates the full stack, including the VSS agent, NIM containers, and supporting services like Elasticsearch and Milvus. The depends_on clauses for NIM containers specify required: false, allowing the agent to start even when referencing remote NIM endpoints that are not managed by the local Compose project.

Practical Integration Examples

Configuring Environment Variables for Local Deployment

Set the following variables before launching the stack to point the VSS agent at local NIM instances:


# .env for a local development profile (dev-profile-search)

export HOST_IP=$(hostname -I | awk '{print $1}')
export VLM_PORT=30082                # Port exposed by the VLM NIM container

export VLM_BASE_URL="http://${HOST_IP}:${VLM_PORT}"
export LLM_PORT=30081
export LLM_BASE_URL="http://${HOST_IP}:${LLM_PORT}"

# Start the stack

docker compose -f deployments/developer-workflow/dev-profile-search/docker-compose.yml up -d

Directly Invoking the VSS Agent's Generate Endpoint

Once running, send requests to the agent's REST API, which forwards them to the underlying VLM NIM:


# Assuming the VSS agent container exposes port 8000

curl -X POST http://localhost:8000/generate \
     -H "Content-Type: application/json" \
     -d '{
           "input": "Summarize the following 2‑minute video clip:",
           "source_type": "video_file",
           "video_path": "/data/sample.mp4"
         }'

The agent translates this into a call to ${VLM_BASE_URL}/v1/chat/completions and returns the generated caption with timestamps.

Using the Python SDK to Call NIM-Backed Tools

You can also invoke the functionality programmatically through the VSS agent's tool abstraction:

from vss_agents.video_understanding.tools import VideoUnderstandingToolConfig, video_understanding

cfg = VideoUnderstandingToolConfig(
    base_url = "http://localhost:8000/v1",   # VSS agent’s base (will proxy to VLM NIM)

    model_name = "cosmos-reason2-8b"
)

result = video_understanding(
    cfg,
    prompt="Describe what is happening in the first 30 seconds."
)
print(result)

The video_understanding tool constructs the final HTTP request to the VLM NIM via the agent's internal client.

Summary

  • Environment variables (VLM_BASE_URL, LLM_BASE_URL) drive dynamic service discovery, allowing the VSS agent to locate NIM containers without code changes.
  • The agent communicates via OpenAI-compatible /v1 endpoints, standardizing interactions across different LLM and VLM NIM implementations.
  • Tool abstraction in agent/src/vss_agents/agents/top_agent.py wraps NIM calls and injects reasoning metadata (llm_reasoning, vlm_reasoning).
  • Placeholder handling prevents empty-payload rejections from strict NIM containers like Nemotron Nano 9b v2.
  • Docker Compose wiring in deployments/agents/vss-agent/vss-agent-docker-compose.yml and root-level compose files supports both colocated and remote NIM deployment scenarios.

Frequently Asked Questions

How does the VSS agent locate NVIDIA NIM microservices at runtime?

The VSS agent reads environment variables such as VLM_BASE_URL and LLM_BASE_URL that are injected by Docker Compose during container startup. These variables resolve to the network addresses of running NIM containers, enabling the agent to discover services dynamically whether they are colocated on the same host or hosted remotely on NVIDIA AI Enterprise cloud.

What API format does the VSS agent use to communicate with NIM containers?

The agent uses an OpenAI-compatible REST API, specifically targeting endpoints like POST /v1/chat/completions. Each NIM microservices container exposes this standard interface, allowing the VSS agent to treat all models—whether LLMs or VLMs—as interchangeable OpenAI-style services without implementing custom protocols for each backend.

Why does the VSS agent send a placeholder string to the NIM microservice?

Certain NIM implementations, including the local Nemotron Nano 9b v2, reject HTTP requests with empty payloads. To prevent these rejections during intermediate processing steps, the agent logic in agent/src/vss_agents/agents/top_agent.py substitutes a minimal non-empty string ("Agent wants to call tools.") when the generated content is empty, ensuring the request lifecycle remains robust.

Can the VSS agent operate without locally deployed NIM containers?

Yes. The depends_on clauses in the Docker Compose configuration specify required: false for NIM services, allowing the agent container to start independently. Operators can configure VLM_BASE_URL and LLM_BASE_URL to point at remote NIM endpoints, enabling the VSS agent to run on bare metal or edge devices while offloading inference to centralized NVIDIA NIM microservices.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →