# How the VSS Agent Architecture Integrates with NVIDIA NIM Microservices

> Discover how the VSS agent architecture integrates with NVIDIA NIM microservices via service discovery, OpenAI-compatible endpoints, and LangChain abstractions. Learn more!

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: architecture
- Published: 2026-05-15

---

**The VSS agent integrates with NVIDIA NIM microservices through environment-driven service discovery, OpenAI-compatible HTTP endpoints, and LangChain-style tool abstractions that route requests to LLM and VLM containers.**

The NVIDIA Video Search & Summarization (VSS) blueprint provides a Python-based orchestrator that connects to NVIDIA NIM (NVIDIA Inference Microservices) containers for large language model (LLM) and vision language model (VLM) inference. Understanding how the VSS agent architecture interfaces with these microservices is essential for deploying scalable video analysis pipelines. This article examines the integration patterns found in the `NVIDIA-AI-Blueprints/video-search-and-summarization` repository, from environment variable configuration to HTTP request handling.

## Environment-Driven Service Discovery

The VSS agent relies on **environment variables** to discover and connect to running NIM containers. During deployment, Docker Compose injects variables such as `LLM_BASE_URL`, `VLM_BASE_URL`, and `HOST_IP` into the agent's runtime environment. These variables point to the network locations of the microservices, allowing the agent to resolve endpoints dynamically without hardcoding addresses.

In [`deployments/agents/vss-agent/vss-agent-docker-compose.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/deployments/agents/vss-agent/vss-agent-docker-compose.yml), the service definition maps these variables directly to the NIM containers:

```yaml
VLM_BASE_URL: ${VLM_BASE_URL:-http://${HOST_IP}:${VLM_PORT}}

```

This pattern enables flexible deployments where operators can run the agent on bare metal while pointing it at remote NIM endpoints hosted on NVIDIA AI Enterprise cloud, or colocate everything via local Docker networking.

## OpenAI-Compatible API Contract

Each NIM container exposes an **OpenAI-compatible `/v1` API**, specifically implementing endpoints like `POST /v1/chat/completions`. The VSS agent never communicates with NIM containers through proprietary protocols; instead, it treats every NIM as an OpenAI-style HTTP service. The agent constructs target URLs by appending `/v1` to the base URLs defined in the environment variables.

According to the skill specifications in [`skills/video-summarization/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/SKILL.md), short-video requests follow this exact pattern:

```

POST ${VLM_BASE_URL}/v1/chat/completions

```

This standardization allows the VSS agent to swap between different NIM models (such as Cosmos Reason 2 or Nemotron Nano) without changing the underlying HTTP client logic.

## Agent-Side Tool Abstraction and Orchestration

### Tool Wrapping and Reasoning Metadata

Inside the agent codebase, every NIM call is wrapped as a **LangChain-style tool** (e.g., `video_understanding`, `search_agent`). The top-level agent defined in [`agent/src/vss_agents/agents/top_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/agents/top_agent.py) builds the request payload, injects the base URL from the environment, and forwards the call to the appropriate NIM. The tool layer additionally inserts **reasoning metadata**—specifically `llm_reasoning` and `vlm_reasoning` parameters—that the NIM consumes to guide inference.

The [`top_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/top_agent.py) file handles the orchestration logic that maps user queries to these tool calls, ensuring the correct NIM endpoint receives the properly formatted request body.

### Handling Empty Payloads

Certain NIM containers, such as the local Nemotron Nano 9b v2, reject empty HTTP payloads. To prevent request failures, [`agent/src/vss_agents/agents/top_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/agents/top_agent.py) implements a safeguard that substitutes a **placeholder string** (`"Agent wants to call tools."`) whenever the generated content is empty. This ensures the NIM's HTTP endpoint accepts the request even during intermediate reasoning steps where content is still being constructed.

## Configuration and Deployment Wiring

Runtime configuration files bridge the gap between environment variables and the agent's HTTP client. The file [`agent/app/video_search_frag/configs/config.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/app/video_search_frag/configs/config.yml) contains entries like:

```yaml
base_url: ${VLM_BASE_URL}/v1

```

When the container starts, Docker Compose resolves `${VLM_BASE_URL}` to the actual NIM address, allowing the agent's internal client to target the correct microservice immediately. The root-level [`docker-compose.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/docker-compose.yml) orchestrates the full stack, including the VSS agent, NIM containers, and supporting services like Elasticsearch and Milvus. The `depends_on` clauses for NIM containers specify `required: false`, allowing the agent to start even when referencing remote NIM endpoints that are not managed by the local Compose project.

## Practical Integration Examples

### Configuring Environment Variables for Local Deployment

Set the following variables before launching the stack to point the VSS agent at local NIM instances:

```bash

# .env for a local development profile (dev-profile-search)

export HOST_IP=$(hostname -I | awk '{print $1}')
export VLM_PORT=30082                # Port exposed by the VLM NIM container

export VLM_BASE_URL="http://${HOST_IP}:${VLM_PORT}"
export LLM_PORT=30081
export LLM_BASE_URL="http://${HOST_IP}:${LLM_PORT}"

# Start the stack

docker compose -f deployments/developer-workflow/dev-profile-search/docker-compose.yml up -d

```

### Directly Invoking the VSS Agent's Generate Endpoint

Once running, send requests to the agent's REST API, which forwards them to the underlying VLM NIM:

```bash

# Assuming the VSS agent container exposes port 8000

curl -X POST http://localhost:8000/generate \
     -H "Content-Type: application/json" \
     -d '{
           "input": "Summarize the following 2‑minute video clip:",
           "source_type": "video_file",
           "video_path": "/data/sample.mp4"
         }'

```

The agent translates this into a call to `${VLM_BASE_URL}/v1/chat/completions` and returns the generated caption with timestamps.

### Using the Python SDK to Call NIM-Backed Tools

You can also invoke the functionality programmatically through the VSS agent's tool abstraction:

```python
from vss_agents.video_understanding.tools import VideoUnderstandingToolConfig, video_understanding

cfg = VideoUnderstandingToolConfig(
    base_url = "http://localhost:8000/v1",   # VSS agent’s base (will proxy to VLM NIM)

    model_name = "cosmos-reason2-8b"
)

result = video_understanding(
    cfg,
    prompt="Describe what is happening in the first 30 seconds."
)
print(result)

```

The `video_understanding` tool constructs the final HTTP request to the VLM NIM via the agent's internal client.

## Summary

- **Environment variables** (`VLM_BASE_URL`, `LLM_BASE_URL`) drive dynamic service discovery, allowing the VSS agent to locate NIM containers without code changes.
- The agent communicates via **OpenAI-compatible `/v1` endpoints**, standardizing interactions across different LLM and VLM NIM implementations.
- **Tool abstraction** in [`agent/src/vss_agents/agents/top_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/agents/top_agent.py) wraps NIM calls and injects reasoning metadata (`llm_reasoning`, `vlm_reasoning`).
- **Placeholder handling** prevents empty-payload rejections from strict NIM containers like Nemotron Nano 9b v2.
- **Docker Compose wiring** in [`deployments/agents/vss-agent/vss-agent-docker-compose.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/deployments/agents/vss-agent/vss-agent-docker-compose.yml) and root-level compose files supports both colocated and remote NIM deployment scenarios.

## Frequently Asked Questions

### How does the VSS agent locate NVIDIA NIM microservices at runtime?

The VSS agent reads environment variables such as `VLM_BASE_URL` and `LLM_BASE_URL` that are injected by Docker Compose during container startup. These variables resolve to the network addresses of running NIM containers, enabling the agent to discover services dynamically whether they are colocated on the same host or hosted remotely on NVIDIA AI Enterprise cloud.

### What API format does the VSS agent use to communicate with NIM containers?

The agent uses an **OpenAI-compatible REST API**, specifically targeting endpoints like `POST /v1/chat/completions`. Each NIM microservices container exposes this standard interface, allowing the VSS agent to treat all models—whether LLMs or VLMs—as interchangeable OpenAI-style services without implementing custom protocols for each backend.

### Why does the VSS agent send a placeholder string to the NIM microservice?

Certain NIM implementations, including the local Nemotron Nano 9b v2, reject HTTP requests with empty payloads. To prevent these rejections during intermediate processing steps, the agent logic in [`agent/src/vss_agents/agents/top_agent.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/agents/top_agent.py) substitutes a minimal non-empty string (`"Agent wants to call tools."`) when the generated content is empty, ensuring the request lifecycle remains robust.

### Can the VSS agent operate without locally deployed NIM containers?

Yes. The `depends_on` clauses in the Docker Compose configuration specify `required: false` for NIM services, allowing the agent container to start independently. Operators can configure `VLM_BASE_URL` and `LLM_BASE_URL` to point at remote NIM endpoints, enabling the VSS agent to run on bare metal or edge devices while offloading inference to centralized NVIDIA NIM microservices.