How to Debug VSS Deployment Issues Using the Deploy Skill and Docker Compose Logs

To debug VSS (Video Search & Summarization) deployment issues, use the deploy skill's built-in workflow to verify container health via endpoint checks, validate the ingestion pipeline with an end-to-end test, and inspect specific service logs using the generated resolved.yml file with Docker Compose commands.

The NVIDIA-AI-Blueprints/video-search-and-summarization repository orchestrates the VSS platform using Docker Compose, with the deploy skill managing the full lifecycle from image builds to stack initialization. When deployments fail to start or services become unresponsive, systematic debugging prevents guesswork. This guide walks through the exact diagnostic commands and source files referenced in skills/deploy/SKILL.md to help you debug VSS deployment issues efficiently.

Understanding the Deploy Skill Architecture

The deploy skill automates the entire deployment pipeline: building container images, generating a resolved Docker Compose file, and bringing up the service stack. Located at skills/deploy/SKILL.md, this skill also embeds a structured debugging workflow that targets the most common failure points in containerized environments.

Step-by-Step Debugging Workflow

Verify Container Status and Core Health Endpoints

Begin with quick sanity checks to confirm all containers are running and critical API endpoints respond. According to the debugging section in skills/deploy/SKILL.md, execute these commands:


# Verify all containers are running

docker ps --format 'table {{.Names}}\t{{.Status}}\t{{.Ports}}'

# Confirm the Agent REST API is reachable (Swagger UI)

curl -sf http://localhost:8000/docs >/dev/null && echo "agent OK"

# Confirm the UI is reachable

curl -sf http://localhost:3000/ >/dev/null && echo "ui OK"

# (Base/LVS profiles) Verify the VLM NIM service

curl -sf http://localhost:30082/v1/models | python3 -m json.tool

# Verify the LLM NIM service

curl -sf http://localhost:30081/v1/models | python3 -m json.tool

These commands validate that the agent, UI, and NIM services (VLM and LLM) have started successfully and are accepting network connections.

Run an End-to-End Video Sanity Check

Once basic connectivity passes, validate the full data pipeline by uploading a video through the VST UI or calling the agent's chat endpoint directly. A non-empty response indicates the ingestion-to-answer path is functional; failures here signal the need for log inspection.


# End-to-end sanity query (after uploading a video)

curl -sS -X POST "http://localhost:8000/v1/chat" \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"Describe the uploaded video"}]}' | python3 -m json.tool

If this request returns a valid JSON response with content, the deployment is healthy. Empty responses or timeouts indicate pipeline failures requiring log analysis.

Inspect Docker Compose Logs for Specific Services

When services fail health checks or requests timeout, examine container logs using the resolved.yml file generated by the deploy skill. This resolved Compose file defines the runtime configuration for your specific deployment profile.


# Show the last 50 lines of logs for the vss-agent container

docker compose -f $REPO/deployments/resolved.yml logs --tail 50 vss-agent

Replace vss-agent with the specific service name (e.g., vlm-nim, llm-nim) to isolate logs for individual components. The resolved.yml file path is critical—do not use the base docker-compose.yml for log inspection, as the deploy skill generates the resolved version with environment-specific variables interpolated.

Resolve Common Deployment Failures

The deploy skill documentation highlights recurring issues and remedies. Common failures include missing NVIDIA Container Toolkit (preventing GPU access), NGC authentication errors blocking image pulls, or a missing resolved.yml file indicating an incomplete dry-run step. Address these prerequisites before debugging runtime errors.

If docker ps shows containers in Restarting or Exited states, verify GPU drivers and toolkit installation first. Authentication errors when pulling images require valid NGC API keys configured in your environment.

Key Source Files for Advanced Troubleshooting

For deeper investigation, reference these implementation files:

Summary

  • The deploy skill in skills/deploy/SKILL.md provides the canonical debugging workflow for VSS deployments.
  • Use docker ps and curl commands to verify container status and service health at ports 8000, 3000, 30081, and 30082.
  • Validate the ingestion pipeline with an end-to-end video test before investigating logs.
  • Inspect service-specific logs using docker compose -f deployments/resolved.yml logs.
  • Check for NVIDIA Container Toolkit and NGC authentication issues when containers fail to start.

Frequently Asked Questions

Where is the resolved.yml file located?

The deploy skill generates resolved.yml inside the deployments/ directory during the build process. If this file is missing, run the dry-run step of the deploy skill to regenerate the resolved Docker Compose configuration specific to your deployment profile.

How do I check if the VLM and LLM NIM services are healthy?

For Base and LVS deployment profiles, query the model endpoints directly. Execute curl -sf http://localhost:30082/v1/models for the VLM NIM service and curl -sf http://localhost:30081/v1/models for the LLM NIM service. Valid JSON responses confirm the services are operational.

What should I do if containers exit immediately with GPU errors?

Immediate container exits typically indicate the NVIDIA Container Toolkit is not installed or configured correctly on the host. Verify that the Docker runtime supports GPUs with docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi before attempting to deploy VSS services.

Can I view logs for all services simultaneously?

Yes, omit the service name from the Docker Compose logs command. Run docker compose -f $REPO/deployments/resolved.yml logs without specifying a container to stream logs from all services in the stack, though filtering by specific service names like vss-agent reduces noise during targeted debugging.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →