How the VSS Agent Architecture Integrates with NVIDIA NIM Microservices
The VSS agent integrates with NVIDIA NIM microservices through environment-driven service discovery, OpenAI-compatible HTTP endpoints, and LangChain-style tool abstractions that route requests to LLM and VLM containers.
The NVIDIA Video Search & Summarization (VSS) blueprint provides a Python-based orchestrator that connects to NVIDIA NIM (NVIDIA Inference Microservices) containers for large language model (LLM) and vision language model (VLM) inference. Understanding how the VSS agent architecture interfaces with these microservices is essential for deploying scalable video analysis pipelines. This article examines the integration patterns found in the NVIDIA-AI-Blueprints/video-search-and-summarization repository, from environment variable configuration to HTTP request handling.
Environment-Driven Service Discovery
The VSS agent relies on environment variables to discover and connect to running NIM containers. During deployment, Docker Compose injects variables such as LLM_BASE_URL, VLM_BASE_URL, and HOST_IP into the agent's runtime environment. These variables point to the network locations of the microservices, allowing the agent to resolve endpoints dynamically without hardcoding addresses.
In deployments/agents/vss-agent/vss-agent-docker-compose.yml, the service definition maps these variables directly to the NIM containers:
VLM_BASE_URL: ${VLM_BASE_URL:-http://${HOST_IP}:${VLM_PORT}}
This pattern enables flexible deployments where operators can run the agent on bare metal while pointing it at remote NIM endpoints hosted on NVIDIA AI Enterprise cloud, or colocate everything via local Docker networking.
OpenAI-Compatible API Contract
Each NIM container exposes an OpenAI-compatible /v1 API, specifically implementing endpoints like POST /v1/chat/completions. The VSS agent never communicates with NIM containers through proprietary protocols; instead, it treats every NIM as an OpenAI-style HTTP service. The agent constructs target URLs by appending /v1 to the base URLs defined in the environment variables.
According to the skill specifications in skills/video-summarization/SKILL.md, short-video requests follow this exact pattern:
POST ${VLM_BASE_URL}/v1/chat/completions
This standardization allows the VSS agent to swap between different NIM models (such as Cosmos Reason 2 or Nemotron Nano) without changing the underlying HTTP client logic.
Agent-Side Tool Abstraction and Orchestration
Tool Wrapping and Reasoning Metadata
Inside the agent codebase, every NIM call is wrapped as a LangChain-style tool (e.g., video_understanding, search_agent). The top-level agent defined in agent/src/vss_agents/agents/top_agent.py builds the request payload, injects the base URL from the environment, and forwards the call to the appropriate NIM. The tool layer additionally inserts reasoning metadata—specifically llm_reasoning and vlm_reasoning parameters—that the NIM consumes to guide inference.
The top_agent.py file handles the orchestration logic that maps user queries to these tool calls, ensuring the correct NIM endpoint receives the properly formatted request body.
Handling Empty Payloads
Certain NIM containers, such as the local Nemotron Nano 9b v2, reject empty HTTP payloads. To prevent request failures, agent/src/vss_agents/agents/top_agent.py implements a safeguard that substitutes a placeholder string ("Agent wants to call tools.") whenever the generated content is empty. This ensures the NIM's HTTP endpoint accepts the request even during intermediate reasoning steps where content is still being constructed.
Configuration and Deployment Wiring
Runtime configuration files bridge the gap between environment variables and the agent's HTTP client. The file agent/app/video_search_frag/configs/config.yml contains entries like:
base_url: ${VLM_BASE_URL}/v1
When the container starts, Docker Compose resolves ${VLM_BASE_URL} to the actual NIM address, allowing the agent's internal client to target the correct microservice immediately. The root-level docker-compose.yml orchestrates the full stack, including the VSS agent, NIM containers, and supporting services like Elasticsearch and Milvus. The depends_on clauses for NIM containers specify required: false, allowing the agent to start even when referencing remote NIM endpoints that are not managed by the local Compose project.
Practical Integration Examples
Configuring Environment Variables for Local Deployment
Set the following variables before launching the stack to point the VSS agent at local NIM instances:
# .env for a local development profile (dev-profile-search)
export HOST_IP=$(hostname -I | awk '{print $1}')
export VLM_PORT=30082 # Port exposed by the VLM NIM container
export VLM_BASE_URL="http://${HOST_IP}:${VLM_PORT}"
export LLM_PORT=30081
export LLM_BASE_URL="http://${HOST_IP}:${LLM_PORT}"
# Start the stack
docker compose -f deployments/developer-workflow/dev-profile-search/docker-compose.yml up -d
Directly Invoking the VSS Agent's Generate Endpoint
Once running, send requests to the agent's REST API, which forwards them to the underlying VLM NIM:
# Assuming the VSS agent container exposes port 8000
curl -X POST http://localhost:8000/generate \
-H "Content-Type: application/json" \
-d '{
"input": "Summarize the following 2‑minute video clip:",
"source_type": "video_file",
"video_path": "/data/sample.mp4"
}'
The agent translates this into a call to ${VLM_BASE_URL}/v1/chat/completions and returns the generated caption with timestamps.
Using the Python SDK to Call NIM-Backed Tools
You can also invoke the functionality programmatically through the VSS agent's tool abstraction:
from vss_agents.video_understanding.tools import VideoUnderstandingToolConfig, video_understanding
cfg = VideoUnderstandingToolConfig(
base_url = "http://localhost:8000/v1", # VSS agent’s base (will proxy to VLM NIM)
model_name = "cosmos-reason2-8b"
)
result = video_understanding(
cfg,
prompt="Describe what is happening in the first 30 seconds."
)
print(result)
The video_understanding tool constructs the final HTTP request to the VLM NIM via the agent's internal client.
Summary
- Environment variables (
VLM_BASE_URL,LLM_BASE_URL) drive dynamic service discovery, allowing the VSS agent to locate NIM containers without code changes. - The agent communicates via OpenAI-compatible
/v1endpoints, standardizing interactions across different LLM and VLM NIM implementations. - Tool abstraction in
agent/src/vss_agents/agents/top_agent.pywraps NIM calls and injects reasoning metadata (llm_reasoning,vlm_reasoning). - Placeholder handling prevents empty-payload rejections from strict NIM containers like Nemotron Nano 9b v2.
- Docker Compose wiring in
deployments/agents/vss-agent/vss-agent-docker-compose.ymland root-level compose files supports both colocated and remote NIM deployment scenarios.
Frequently Asked Questions
How does the VSS agent locate NVIDIA NIM microservices at runtime?
The VSS agent reads environment variables such as VLM_BASE_URL and LLM_BASE_URL that are injected by Docker Compose during container startup. These variables resolve to the network addresses of running NIM containers, enabling the agent to discover services dynamically whether they are colocated on the same host or hosted remotely on NVIDIA AI Enterprise cloud.
What API format does the VSS agent use to communicate with NIM containers?
The agent uses an OpenAI-compatible REST API, specifically targeting endpoints like POST /v1/chat/completions. Each NIM microservices container exposes this standard interface, allowing the VSS agent to treat all models—whether LLMs or VLMs—as interchangeable OpenAI-style services without implementing custom protocols for each backend.
Why does the VSS agent send a placeholder string to the NIM microservice?
Certain NIM implementations, including the local Nemotron Nano 9b v2, reject HTTP requests with empty payloads. To prevent these rejections during intermediate processing steps, the agent logic in agent/src/vss_agents/agents/top_agent.py substitutes a minimal non-empty string ("Agent wants to call tools.") when the generated content is empty, ensuring the request lifecycle remains robust.
Can the VSS agent operate without locally deployed NIM containers?
Yes. The depends_on clauses in the Docker Compose configuration specify required: false for NIM services, allowing the agent container to start independently. Operators can configure VLM_BASE_URL and LLM_BASE_URL to point at remote NIM endpoints, enabling the VSS agent to run on bare metal or edge devices while offloading inference to centralized NVIDIA NIM microservices.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →