# How to Debug VSS Deployment Issues Using the Deploy Skill and Docker Compose Logs

> Debug VSS deployment issues effectively. Use the deploy skill for container health checks, pipeline validation, and inspect Docker Compose logs with resolved.yml for quick troubleshooting.

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: how-to-guide
- Published: 2026-05-15

---

**To debug VSS (Video Search & Summarization) deployment issues, use the `deploy` skill's built-in workflow to verify container health via endpoint checks, validate the ingestion pipeline with an end-to-end test, and inspect specific service logs using the generated [`resolved.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/resolved.yml) file with Docker Compose commands.**

The NVIDIA-AI-Blueprints/video-search-and-summarization repository orchestrates the VSS platform using Docker Compose, with the **`deploy` skill** managing the full lifecycle from image builds to stack initialization. When deployments fail to start or services become unresponsive, systematic debugging prevents guesswork. This guide walks through the exact diagnostic commands and source files referenced in [`skills/deploy/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/deploy/SKILL.md) to help you debug VSS deployment issues efficiently.

## Understanding the Deploy Skill Architecture

The **`deploy` skill** automates the entire deployment pipeline: building container images, generating a resolved Docker Compose file, and bringing up the service stack. Located at [`skills/deploy/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/deploy/SKILL.md), this skill also embeds a structured debugging workflow that targets the most common failure points in containerized environments.

## Step-by-Step Debugging Workflow

### Verify Container Status and Core Health Endpoints

Begin with quick sanity checks to confirm all containers are running and critical API endpoints respond. According to the debugging section in [`skills/deploy/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/deploy/SKILL.md), execute these commands:

```bash

# Verify all containers are running

docker ps --format 'table {{.Names}}\t{{.Status}}\t{{.Ports}}'

# Confirm the Agent REST API is reachable (Swagger UI)

curl -sf http://localhost:8000/docs >/dev/null && echo "agent OK"

# Confirm the UI is reachable

curl -sf http://localhost:3000/ >/dev/null && echo "ui OK"

# (Base/LVS profiles) Verify the VLM NIM service

curl -sf http://localhost:30082/v1/models | python3 -m json.tool

# Verify the LLM NIM service

curl -sf http://localhost:30081/v1/models | python3 -m json.tool

```

These commands validate that the **agent**, **UI**, and **NIM services** (VLM and LLM) have started successfully and are accepting network connections.

### Run an End-to-End Video Sanity Check

Once basic connectivity passes, validate the full data pipeline by uploading a video through the VST UI or calling the agent's chat endpoint directly. A non-empty response indicates the ingestion-to-answer path is functional; failures here signal the need for log inspection.

```bash

# End-to-end sanity query (after uploading a video)

curl -sS -X POST "http://localhost:8000/v1/chat" \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"Describe the uploaded video"}]}' | python3 -m json.tool

```

If this request returns a valid JSON response with content, the deployment is healthy. Empty responses or timeouts indicate pipeline failures requiring log analysis.

### Inspect Docker Compose Logs for Specific Services

When services fail health checks or requests timeout, examine container logs using the **[`resolved.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/resolved.yml)** file generated by the deploy skill. This resolved Compose file defines the runtime configuration for your specific deployment profile.

```bash

# Show the last 50 lines of logs for the vss-agent container

docker compose -f $REPO/deployments/resolved.yml logs --tail 50 vss-agent

```

Replace `vss-agent` with the specific service name (e.g., `vlm-nim`, `llm-nim`) to isolate logs for individual components. The [`resolved.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/resolved.yml) file path is critical—do not use the base [`docker-compose.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/docker-compose.yml) for log inspection, as the deploy skill generates the resolved version with environment-specific variables interpolated.

### Resolve Common Deployment Failures

The `deploy` skill documentation highlights recurring issues and remedies. Common failures include missing **NVIDIA Container Toolkit** (preventing GPU access), **NGC authentication errors** blocking image pulls, or a missing [`resolved.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/resolved.yml) file indicating an incomplete dry-run step. Address these prerequisites before debugging runtime errors.

If `docker ps` shows containers in `Restarting` or `Exited` states, verify GPU drivers and toolkit installation first. Authentication errors when pulling images require valid NGC API keys configured in your environment.

## Key Source Files for Advanced Troubleshooting

For deeper investigation, reference these implementation files:

- **`agent/docker/Dockerfile`**: Defines the base VSS agent image; modifications here require rebuilding the stack.
- **[`agent/src/vss_agents/video_analytics/interface.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/video_analytics/interface.py)**: Contains the HTTP API entry point; exceptions here surface as health-check failures on port 8000.
- **[`agent/src/vss_agents/utils/retry.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/utils/retry.py)**: Controls retry logic for downstream services; consult this when investigating intermittent connectivity or timeout issues.
- **[`agent/app/video_search_frag/docker-compose.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/app/video_search_frag/docker-compose.yml)**: Example compose configuration for the frag extension, useful when debugging services that extend the base deployment.

## Summary

- The **`deploy` skill** in [`skills/deploy/SKILL.md`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/skills/deploy/SKILL.md) provides the canonical debugging workflow for VSS deployments.
- Use **`docker ps`** and **`curl`** commands to verify container status and service health at ports **8000**, **3000**, **30081**, and **30082**.
- Validate the ingestion pipeline with an **end-to-end video test** before investigating logs.
- Inspect service-specific logs using **`docker compose -f deployments/resolved.yml logs`**.
- Check for **NVIDIA Container Toolkit** and **NGC authentication** issues when containers fail to start.

## Frequently Asked Questions

### Where is the resolved.yml file located?

The deploy skill generates [`resolved.yml`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/resolved.yml) inside the `deployments/` directory during the build process. If this file is missing, run the dry-run step of the deploy skill to regenerate the resolved Docker Compose configuration specific to your deployment profile.

### How do I check if the VLM and LLM NIM services are healthy?

For Base and LVS deployment profiles, query the model endpoints directly. Execute `curl -sf http://localhost:30082/v1/models` for the VLM NIM service and `curl -sf http://localhost:30081/v1/models` for the LLM NIM service. Valid JSON responses confirm the services are operational.

### What should I do if containers exit immediately with GPU errors?

Immediate container exits typically indicate the NVIDIA Container Toolkit is not installed or configured correctly on the host. Verify that the Docker runtime supports GPUs with `docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi` before attempting to deploy VSS services.

### Can I view logs for all services simultaneously?

Yes, omit the service name from the Docker Compose logs command. Run `docker compose -f $REPO/deployments/resolved.yml logs` without specifying a container to stream logs from all services in the stack, though filtering by specific service names like `vss-agent` reduces noise during targeted debugging.