Security Considerations for Deploying VSS with NVIDIA NIM: Best Practices and Checklist
Securing a Video Search & Summarization (VSS) deployment with NVIDIA NIM requires strict isolation of API keys in .env files, careful configuration of host-mode networking on ports 38111-38113, and explicit GPU memory limits to prevent resource exhaustion attacks.
The Video Search & Summarization (VSS) blueprint from the NVIDIA-AI-Blueprints/video-search-and-summarization repository orchestrates multiple microservices—including the VSS Agent, VST ingestion service, and Long-Video Summarization (LVS)—alongside NVIDIA NIM containers pulled from NGC. Because the stack uses host networking and relies on external LLM/VLM endpoints, security hinges on credential management, network exposure controls, and GPU resource isolation.
Secure Credential Management for NGC and NIM Access
The VSS stack requires several authentication tokens that must never be committed to version control. According to [SECURITY.md](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/SECURITY.md), all secrets belong in a runtime .env file excluded from Git.
Critical secrets include:
NGC_API_KEY– Authenticates pulls fromnvcr.ioto retrieve private NIM imagesNVIDIA_API_KEYorOPENAI_API_KEY– Provides access to LLM endpoints with a fallback chain (OPENAI_API_KEY→NVIDIA_API_KEY)BREV_ENV_ID/BREV_LINK_PREFIX– Secures video URL tokens for the VST ↔ VSS bridge
The reference template in skills/video-summarization/references/lvs.env.example demonstrates the required variables. Always copy this template to $MDX_SAMPLE_APPS_DIR/lvs/.env and populate values offline—never embed credentials in compose.yml or Dockerfiles.
Network Exposure Risks in Host-Mode Docker Compose
The LVS microservice uses network_mode: host as implemented in [deployments/lvs/compose.yml](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/deployments/lvs/compose.yml). This configuration binds containers directly to the host network stack, bypassing Docker’s bridge isolation.
Security implications:
- Services bind to ports 38111, 38112, and 38113 directly on the host interface
- The
ports:mapping in Compose is ignored, making firewall rules critical - Any process on the host or connected network can reach these APIs if unprotected
Override default ports via environment variables to avoid collisions and reduce predictability:
cat >> "$MDX_SAMPLE_APPS_DIR/lvs/.env" <<EOF
BACKEND_PORT=48111
LVS_MCP_PORT=48112
FRONTEND_PORT=48113
EOF
Apply host-level firewall rules to restrict access to these ports after deployment, as documented in [skills/video-summarization/references/deploy-lvs-service.md](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/references/deploy-lvs-service.md).
GPU Isolation and Resource Constraints
LVS and NIM sidecars request GPUs through the modern Compose deploy.resources.reservations.devices syntax (lines 71-76 of compose.yml). Unconstrained GPU access can lead to OOM kills (Exit 137) or cross-service resource starvation.
Controls to implement:
GPU_DEVICES– Restricts visibility to specific GPUs (e.g.,GPU_DEVICES=0,1)VLM_BATCH_SIZE– Auto-tuned based on VRAM, but should be capped manually:- ≤ 46 GB VRAM → 1-3 batches
- 46-80 GB → 2-16 batches
- > 80 GB (int4) → up to 128 batches
VLLM_GPU_MEMORY_UTILIZATION– Limits VRAM fraction per container (e.g.,0.5for 50%)
Set conservative limits to prevent denial-of-service via resource exhaustion:
echo "VLM_BATCH_SIZE=2" >> "$MDX_SAMPLE_APPS_DIR/lvs/.env"
echo "VLLM_GPU_MEMORY_UTILIZATION=0.5" >> "$MDX_SAMPLE_APPS_DIR/lvs/.env"
These settings are detailed in the troubleshooting guide [skills/video-summarization/references/lvs-debugging.md](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/references/lvs-debugging.md).
Securing NIM Endpoint Authentication
NIM containers expose OpenAI-compatible endpoints (/v1/chat/completions, /v1/models) but do not embed API keys. Instead, they rely on environment variables injected by the host runtime:
- LLM endpoints – Require
LLM_BASE_URL(orHOST_IP+LLM_PORT) andLVS_LLM_MODEL_NAME - VLM endpoints – Require
VLM_BASE_URL(orHOST_IP+VLM_PORT) andVIA_VLM_API_KEY
If authentication is missing, the NIM returns HTTP 401, which surfaces as a healthcheck failure in LVS after the 120-second grace period. Verify endpoint reachability before starting the stack:
curl -s $LLM_BASE_URL/v1/models
curl -s $VLM_BASE_URL/v1/models
Secure Deployment Checklist
Follow this sequence to deploy VSS with security controls enforced:
# 1. Set blueprint root for config binds
export MDX_SAMPLE_APPS_DIR=~/met-blueprints/deployments
# 2. Authenticate with NGC (requires NGC_API_KEY)
export NGC_API_KEY="nvapi-..."
docker login nvcr.io -u '$oauthtoken' -p "$NGC_API_KEY"
# 3. Pull images using the mandatory profile flag
docker compose \
-f "$MDX_SAMPLE_APPS_DIR/lvs/compose.yml" \
--profile bp_developer_lvs_2d pull
# 4. Initialize environment from template
cp skills/video-summarization/references/lvs.env.example \
"$MDX_SAMPLE_APPS_DIR/lvs/.env"
# Edit to set: NGC_API_KEY, LLM_BASE_URL, VLM_BASE_URL, MODEL names, GPU limits
# 5. Deploy with profile
docker compose \
-f "$MDX_SAMPLE_APPS_DIR/lvs/compose.yml" \
--profile bp_developer_lvs_2d up -d
# 6. Verify health (120s start period)
until curl -sf http://localhost:${BACKEND_PORT:-38111}/v1/ready; do
sleep 5
done
echo "LVS ready"
Never omit the --profile bp_developer_lvs_2d flag; without it, Compose creates no containers, which can lead to misconfiguration errors if an operator attempts manual workarounds.
Common Security Failure Modes
| Symptom | Root Cause | Mitigation |
|---|---|---|
| No containers created | Missing --profile flag |
Append --profile bp_developer_lvs_2d to all compose commands |
bind source path does not exist |
MDX_SAMPLE_APPS_DIR unset |
Export the variable pointing to your deployment directory |
| Unauthorized (401) on pull | Using nvstaging tag without access |
Switch to public tag (nvcr.io/nvidia/vss-core/vss-long-video-summarization:3.1.0) via CONTAINER_IMAGE |
| Exit 137 (OOM) | GPU VRAM exhausted | Lower VLM_BATCH_SIZE and VLLM_GPU_MEMORY_UTILIZATION |
| Healthcheck timeout | Upstream LLM/VLM unreachable | Verify firewall rules and endpoint URLs from host |
| Port binding error | Collision on 38111-38113 | Override BACKEND_PORT, LVS_MCP_PORT, FRONTEND_PORT in .env |
All failure modes are documented in [skills/video-summarization/references/lvs-debugging.md](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/video-summarization/references/lvs-debugging.md).
Summary
- Store secrets exclusively in
.env– Never commitNGC_API_KEYorNVIDIA_API_KEYto the repository; follow the template inlvs.env.example. - Use public NIM image tags – Avoid
nvstagingprefixes unless you have explicit NGC staging access. - Always specify the compose profile – Invoke every
docker composecommand with--profile bp_developer_lvs_2d. - Lock down host ports – Host networking exposes services directly; customize ports via environment variables and apply strict firewall rules.
- Constrain GPU resources – Set
VLM_BATCH_SIZEandVLLM_GPU_MEMORY_UTILIZATIONto prevent OOM conditions and ensure fair resource sharing.
Frequently Asked Questions
Where should I store the NGC_API_KEY when deploying VSS?
Store the NGC_API_KEY only in the runtime .env file at $MDX_SAMPLE_APPS_DIR/lvs/.env, never in source control or shell history. The file is pre-configured in .gitignore to prevent accidental commits, and the template at skills/video-summarization/references/lvs.env.example shows the required format.
Why does VSS use host networking mode instead of Docker bridge networking?
The LVS microservice uses network_mode: host to bind directly to host ports 38111-38113 for performance and compatibility with GPU passthrough requirements. This bypasses Docker’s NAT layer but requires explicit firewall configuration on the host since container-level port mapping is ignored.
How do I prevent GPU out-of-memory errors in NIM containers?
Set explicit limits via environment variables: reduce VLM_BATCH_SIZE (e.g., to 2 for 48GB cards) and lower VLLM_GPU_MEMORY_UTILIZATION (e.g., to 0.5). The entrypoint auto-tunes based on detected VRAM, but manual caps prevent the "Exit 137" OOM kills documented in the debugging guide.
What causes the "unauthorized" error when pulling VSS containers?
This occurs when using a staging container tag (nvstaging) without proper NGC credentials, or when the NGC_API_KEY is invalid. Switch to the public production tag by setting CONTAINER_IMAGE=nvcr.io/nvidia/vss-core/vss-long-video-summarization:3.1.0 in your .env file, or verify your API key has pull permissions for private images.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →