What Is the Model Context Protocol (MCP) in VSS for Video Analytics?

The Model Context Protocol (MCP) in NVIDIA Video Search & Summarization (VSS) is a JSON-RPC over HTTP service that exposes video analytics data—incidents, alerts, and sensor metrics—as callable tools for large language models, operating on default ports 9901 (Video Analytics) and 38112 (Long Video Summarization).

Video analytics pipelines require a clean separation between real-time feature extraction and downstream AI reasoning. The Model Context Protocol (MCP) bridges this gap by providing a deterministic, language-model-friendly interface that allows agents to query stored analytics without embedding complex SDK logic directly into their reasoning layers.

What Is the Model Context Protocol (MCP)?

The Model Context Protocol (MCP) is a lightweight JSON-RPC service that runs alongside the VSS stack as implemented in the NVIDIA-AI-Blueprints/video-search-and-summarization repository. It exposes video analytics capabilities through standard HTTP endpoints, converting complex queries into structured tool calls that any compatible language model can invoke.

According to the top-level documentation, the agent "leverages the Model Context Protocol (MCP) to access video analytics data, incident records, and vision processing capabilities through a unified tool interface" (see README.md). This design allows the agentic layer to treat video data as first-class functions rather than raw database connections.

Why VSS Uses MCP for Video Analytics

NVIDIA VSS separates real-time video intelligence (embedding generation, feature extraction) from offline analytics and agentic processing. The MCP exists to solve a specific architectural need: providing the LLM with a single, deterministic interface to Elasticsearch-stored analytics without requiring the model to understand query syntax or database schemas.

Key benefits of this approach include:

  • Decoupling: The agent runtime remains agnostic to how video analytics are stored or indexed.
  • Standardization: All data access follows the JSON-RPC 2.0 specification, producing predictable JSON responses that the LLM can parse and reason over.
  • Security: The protocol operates over HTTP with session-based authentication, preventing unauthorized access to sensitive incident data.

Architecture and Data Flow

In the VSS architecture, the MCP server sits between the Elasticsearch store (which holds incidents, alerts, and sensor metrics) and the agent runtime. When the LLM determines it needs video analytics data, the agent dispatches a tool call that translates into an HTTP request to the MCP endpoint.

The Two-Step MCP Workflow

As documented in skills/video-analytics/SKILL.md, every interaction follows a strict initialization-then-invocation pattern:

  1. Initialize a session – The client POSTs an initialize method to the MCP endpoint. The server returns a header mcp-session-id that must be included in all subsequent requests. Omitting this header results in a "Bad Request: Missing session ID" error.
  2. Call a tool – Each capability is addressed as tools/call with the concrete tool name (e.g., video_analytics__get_incidents) in the request payload.
  3. Receive JSON result – The MCP streams back a JSON-RPC response containing the requested data structure, which the agent incorporates into its final output.

This pattern ensures stateful session management while maintaining compatibility with stateless HTTP infrastructure.

Configuring and Enabling MCP

MCP functionality is optional and controlled via environment variables. According to skills/video-summarization/references/deploy-lvs-service.md, you enable the service by setting:

LVS_ENABLE_MCP=true

Port configuration is flexible to avoid clashes with existing services:

  • Video Analytics MCP: Defaults to port 9901 (configurable via VA_MCP_PORT)
  • Long Video Summarization MCP: Defaults to port 38112 (configurable via LVS_MCP_PORT)

These settings allow multiple MCP instances to run concurrently on the same host without conflict.

Practical Example: Querying Incidents via MCP

The following shell commands demonstrate the complete MCP workflow for retrieving the ten most recent incidents from the Video Analytics service. This implements the exact pattern described in the skill documentation:


# Step 1 – Initialize the MCP session and capture the session ID

SESSION_ID=$(curl -si -X POST http://localhost:9901/mcp \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"cli","version":"1.0"}},"id":0}' \
  | grep -i "mcp-session-id" | awk '{print $2}' | tr -d '\r')

# Step 2 – Invoke the get_incidents tool with the captured session ID

curl -s -X POST http://localhost:9901/mcp \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -H "mcp-session-id: $SESSION_ID" \
  -d '{"jsonrpc":"2.0","method":"tools/call","params":{"name":"video_analytics__get_incidents","arguments":{"max_count":10}},"id":1}' \
  | grep '^data:' | sed 's/^data: //' | jq -r '.result.content[0].text'

To query sensor IDs instead, change the name parameter to video_analytics__get_sensor_ids and adjust the arguments object accordingly. The available tools and their schemas are defined in the tool reference section of skills/video-analytics/SKILL.md.

Integration with the VSS Agent

The VSS agent treats MCP-exposed tools exactly like standard LangChain tools. In agent/src/vss_agents/agents/top_agent.py, the top-level agent discovers and dispatches MCP tools through Python wrappers defined in agent/src/vss_agents/video_analytics/tools.py.

When the LLM emits a tool call such as video_analytics_mcp.video_analytics.get_sensor_ids, the framework:

  1. Looks up the corresponding wrapper function in tools.py
  2. Executes the HTTP JSON-RPC flow against the MCP server
  3. Parses the JSON-RPC response
  4. Returns the structured data to the LLM context window

This implementation allows developers to extend video analytics capabilities by adding new tool definitions to the MCP server without modifying the core agent logic in top_agent.py.

Summary

  • The Model Context Protocol (MCP) is VSS's JSON-RPC bridge that turns video analytics data into invocable tools for large language models.
  • It operates on port 9901 (Video Analytics) and 38112 (Long Video Summarization) by default, configured via VA_MCP_PORT and LVS_MCP_PORT.
  • All interactions require a two-step workflow: initialize to obtain an mcp-session-id, then call specific tools like video_analytics__get_incidents.
  • The protocol is optional; enable it by setting LVS_ENABLE_MCP=true.
  • Agent integration happens through LangChain-compatible wrappers in agent/src/vss_agents/video_analytics/tools.py, driven by the top agent in agent/src/vss_agents/agents/top_agent.py.

Frequently Asked Questions

What is the default port for the Video Analytics MCP in VSS?

The Video Analytics MCP listens on port 9901 by default. For Long Video Summarization MCP, the default is port 38112. You can override these by setting the VA_MCP_PORT or LVS_MCP_PORT environment variables before starting the service.

How do I enable MCP in NVIDIA VSS?

Set the environment variable LVS_ENABLE_MCP=true when deploying the Long Video Summarization service. For the Video Analytics MCP, launch the dedicated VA-MCP container. Refer to skills/video-summarization/references/deploy-lvs-service.md for detailed deployment instructions.

What tools are exposed through the Model Context Protocol?

The MCP exposes video analytics tools including video_analytics__get_incidents, video_analytics__get_sensor_ids, and video_analytics__analyze. Each tool accepts specific arguments defined in the schema documentation within skills/video-analytics/SKILL.md.

Why does the MCP require a two-step initialization process?

The initialization step establishes a session state identified by the mcp-session-id header. This design maintains stateful security and transaction context across multiple HTTP requests while keeping the underlying transport protocol stateless. Skipping initialization results in a "Missing session ID" error because the server cannot authenticate or track the request context without this identifier.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →