# Security Considerations for Deploying VSS with NVIDIA NIM: Best Practices and Checklist

> Discover essential security considerations for deploying VSS with NVIDIA NIM. Learn best practices for API key isolation network configuration and GPU memory limits to secure your deployment.

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: best-practices
- Published: 2026-05-15

---

**Securing a Video Search & Summarization (VSS) deployment with NVIDIA NIM requires strict isolation of API keys in `.env` files, careful configuration of host-mode networking on ports 38111-38113, and explicit GPU memory limits to prevent resource exhaustion attacks.**

The Video Search & Summarization (VSS) blueprint from the `NVIDIA-AI-Blueprints/video-search-and-summarization` repository orchestrates multiple microservices—including the VSS Agent, VST ingestion service, and Long-Video Summarization (LVS)—alongside NVIDIA NIM containers pulled from NGC. Because the stack uses **host networking** and relies on external LLM/VLM endpoints, security hinges on credential management, network exposure controls, and GPU resource isolation.

## Secure Credential Management for NGC and NIM Access

The VSS stack requires several authentication tokens that must never be committed to version control. According to [[`SECURITY.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/SECURITY.md)](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/SECURITY.md), all secrets belong in a runtime `.env` file excluded from Git.

**Critical secrets include:**

- **`NGC_API_KEY`** – Authenticates pulls from `nvcr.io` to retrieve private NIM images
- **`NVIDIA_API_KEY`** or **`OPENAI_API_KEY`** – Provides access to LLM endpoints with a fallback chain (`OPENAI_API_KEY` → `NVIDIA_API_KEY`)
- **`BREV_ENV_ID`** / **`BREV_LINK_PREFIX`** – Secures video URL tokens for the VST ↔ VSS bridge

The reference template in [`skills/video-summarization/references/lvs.env.example`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/references/lvs.env.example) demonstrates the required variables. Always copy this template to `$MDX_SAMPLE_APPS_DIR/lvs/.env` and populate values offline—never embed credentials in [`compose.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/compose.yml) or Dockerfiles.

## Network Exposure Risks in Host-Mode Docker Compose

The LVS microservice uses `network_mode: host` as implemented in [[`deployments/lvs/compose.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/deployments/lvs/compose.yml)](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/deployments/lvs/compose.yml). This configuration binds containers directly to the host network stack, bypassing Docker’s bridge isolation.

**Security implications:**

- Services bind to **ports 38111, 38112, and 38113** directly on the host interface
- The `ports:` mapping in Compose is ignored, making firewall rules critical
- Any process on the host or connected network can reach these APIs if unprotected

Override default ports via environment variables to avoid collisions and reduce predictability:

```bash
cat >> "$MDX_SAMPLE_APPS_DIR/lvs/.env" <<EOF
BACKEND_PORT=48111
LVS_MCP_PORT=48112
FRONTEND_PORT=48113
EOF

```

Apply host-level firewall rules to restrict access to these ports after deployment, as documented in [[`skills/video-summarization/references/deploy-lvs-service.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/references/deploy-lvs-service.md)](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/references/deploy-lvs-service.md).

## GPU Isolation and Resource Constraints

LVS and NIM sidecars request GPUs through the modern Compose `deploy.resources.reservations.devices` syntax (lines 71-76 of [`compose.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/compose.yml)). Unconstrained GPU access can lead to **OOM kills (Exit 137)** or cross-service resource starvation.

**Controls to implement:**

- **`GPU_DEVICES`** – Restricts visibility to specific GPUs (e.g., `GPU_DEVICES=0,1`)
- **`VLM_BATCH_SIZE`** – Auto-tuned based on VRAM, but should be capped manually:
  - ≤ 46 GB VRAM → 1-3 batches
  - 46-80 GB → 2-16 batches  
  - \> 80 GB (int4) → up to 128 batches
- **`VLLM_GPU_MEMORY_UTILIZATION`** – Limits VRAM fraction per container (e.g., `0.5` for 50%)

Set conservative limits to prevent denial-of-service via resource exhaustion:

```bash
echo "VLM_BATCH_SIZE=2" >> "$MDX_SAMPLE_APPS_DIR/lvs/.env"
echo "VLLM_GPU_MEMORY_UTILIZATION=0.5" >> "$MDX_SAMPLE_APPS_DIR/lvs/.env"

```

These settings are detailed in the troubleshooting guide [[`skills/video-summarization/references/lvs-debugging.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/references/lvs-debugging.md)](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/references/lvs-debugging.md).

## Securing NIM Endpoint Authentication

NIM containers expose OpenAI-compatible endpoints (`/v1/chat/completions`, `/v1/models`) but **do not embed API keys**. Instead, they rely on environment variables injected by the host runtime:

- **LLM endpoints** – Require `LLM_BASE_URL` (or `HOST_IP` + `LLM_PORT`) and `LVS_LLM_MODEL_NAME`
- **VLM endpoints** – Require `VLM_BASE_URL` (or `HOST_IP` + `VLM_PORT`) and `VIA_VLM_API_KEY`

If authentication is missing, the NIM returns HTTP 401, which surfaces as a healthcheck failure in LVS after the 120-second grace period. Verify endpoint reachability before starting the stack:

```bash
curl -s $LLM_BASE_URL/v1/models
curl -s $VLM_BASE_URL/v1/models

```

## Secure Deployment Checklist

Follow this sequence to deploy VSS with security controls enforced:

```bash

# 1. Set blueprint root for config binds

export MDX_SAMPLE_APPS_DIR=~/met-blueprints/deployments

# 2. Authenticate with NGC (requires NGC_API_KEY)

export NGC_API_KEY="nvapi-..."
docker login nvcr.io -u '$oauthtoken' -p "$NGC_API_KEY"

# 3. Pull images using the mandatory profile flag

docker compose \
  -f "$MDX_SAMPLE_APPS_DIR/lvs/compose.yml" \
  --profile bp_developer_lvs_2d pull

# 4. Initialize environment from template

cp skills/video-summarization/references/lvs.env.example \
   "$MDX_SAMPLE_APPS_DIR/lvs/.env"

# Edit to set: NGC_API_KEY, LLM_BASE_URL, VLM_BASE_URL, MODEL names, GPU limits

# 5. Deploy with profile

docker compose \
  -f "$MDX_SAMPLE_APPS_DIR/lvs/compose.yml" \
  --profile bp_developer_lvs_2d up -d

# 6. Verify health (120s start period)

until curl -sf http://localhost:${BACKEND_PORT:-38111}/v1/ready; do
  sleep 5
done
echo "LVS ready"

```

Never omit the `--profile bp_developer_lvs_2d` flag; without it, Compose creates no containers, which can lead to misconfiguration errors if an operator attempts manual workarounds.

## Common Security Failure Modes

| Symptom | Root Cause | Mitigation |
|---------|------------|------------|
| **No containers created** | Missing `--profile` flag | Append `--profile bp_developer_lvs_2d` to all compose commands |
| **`bind source path does not exist`** | `MDX_SAMPLE_APPS_DIR` unset | Export the variable pointing to your deployment directory |
| **Unauthorized (401) on pull** | Using `nvstaging` tag without access | Switch to public tag (`nvcr.io/nvidia/vss-core/vss-long-video-summarization:3.1.0`) via `CONTAINER_IMAGE` |
| **Exit 137 (OOM)** | GPU VRAM exhausted | Lower `VLM_BATCH_SIZE` and `VLLM_GPU_MEMORY_UTILIZATION` |
| **Healthcheck timeout** | Upstream LLM/VLM unreachable | Verify firewall rules and endpoint URLs from host |
| **Port binding error** | Collision on 38111-38113 | Override `BACKEND_PORT`, `LVS_MCP_PORT`, `FRONTEND_PORT` in `.env` |

All failure modes are documented in [[`skills/video-summarization/references/lvs-debugging.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/references/lvs-debugging.md)](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/references/lvs-debugging.md).

## Summary

- **Store secrets exclusively in `.env`** – Never commit `NGC_API_KEY` or `NVIDIA_API_KEY` to the repository; follow the template in `lvs.env.example`.
- **Use public NIM image tags** – Avoid `nvstaging` prefixes unless you have explicit NGC staging access.
- **Always specify the compose profile** – Invoke every `docker compose` command with `--profile bp_developer_lvs_2d`.
- **Lock down host ports** – Host networking exposes services directly; customize ports via environment variables and apply strict firewall rules.
- **Constrain GPU resources** – Set `VLM_BATCH_SIZE` and `VLLM_GPU_MEMORY_UTILIZATION` to prevent OOM conditions and ensure fair resource sharing.

## Frequently Asked Questions

### Where should I store the NGC_API_KEY when deploying VSS?

Store the `NGC_API_KEY` only in the runtime `.env` file at `$MDX_SAMPLE_APPS_DIR/lvs/.env`, never in source control or shell history. The file is pre-configured in `.gitignore` to prevent accidental commits, and the template at `skills/video-summarization/references/lvs.env.example` shows the required format.

### Why does VSS use host networking mode instead of Docker bridge networking?

The LVS microservice uses `network_mode: host` to bind directly to host ports 38111-38113 for performance and compatibility with GPU passthrough requirements. This bypasses Docker’s NAT layer but requires explicit firewall configuration on the host since container-level port mapping is ignored.

### How do I prevent GPU out-of-memory errors in NIM containers?

Set explicit limits via environment variables: reduce `VLM_BATCH_SIZE` (e.g., to 2 for 48GB cards) and lower `VLLM_GPU_MEMORY_UTILIZATION` (e.g., to 0.5). The entrypoint auto-tunes based on detected VRAM, but manual caps prevent the "Exit 137" OOM kills documented in the debugging guide.

### What causes the "unauthorized" error when pulling VSS containers?

This occurs when using a staging container tag (`nvstaging`) without proper NGC credentials, or when the `NGC_API_KEY` is invalid. Switch to the public production tag by setting `CONTAINER_IMAGE=nvcr.io/nvidia/vss-core/vss-long-video-summarization:3.1.0` in your `.env` file, or verify your API key has pull permissions for private images.