MCP Tool Poisoning Detection Patterns in NVIDIA SkillSpector: A Complete Guide to TP1-TP4

NVIDIA SkillSpector implements four distinct detection patterns (TP1-TP4) in src/skillspector/nodes/analyzers/mcp_tool_poisoning.py to identify hidden instructions, Unicode deception, parameter injection attacks, and LLM-based behavioral mismatches in AI skill manifests.

The SkillSpector framework provides static analysis capabilities for NVIDIA Skill Manifests, protecting against malicious capability poisoning (MCP) through specialized analyzer nodes. The MCP tool poisoning analyzer examines every textual field within a skill manifest—ranging from names and descriptions to parameter metadata—to identify four specific attack vectors that threat actors use to compromise AI tool behavior.

Understanding the Four MCP Tool Poisoning Detection Patterns

The analyzer located at src/skillspector/nodes/analyzers/mcp_tool_poisoning.py implements a multi-layered defense strategy through four specialized detection patterns. Each pattern targets distinct attack methodologies used to inject malicious instructions into tool definitions.

TP1: Hidden Instructions Detection

TP1 identifies concealed payloads embedded within manifest text fields that remain invisible during casual inspection but executable by parsing systems. The _check_tp1() function employs multiple regex patterns to detect:

  • HTML comments (<!--...-->) using _HTML_COMMENT_RE
  • Markdown comments ([//]: #(...)) using _MARKDOWN_COMMENT_RE
  • Zero-width Unicode characters (invisible joiners and non-printable code points) using _ZERO_WIDTH_RE
  • Base64-encoded blobs using _BASE64_RE to catch obfuscated data payloads
  • Data-URI schemes using _DATA_URI_RE to detect embedded executable content

When the analyzer detects any of these patterns in metadata fields, it generates a Finding object documenting the specific match location and threat type.

TP2: Unicode Deception and Homoglyph Attacks

TP2 targets visual spoofing attacks that use Unicode confusables to masquerade malicious parameters as legitimate system variables. The _check_tp2() function in mcp_tool_poisoning.py performs script analysis using:

  • _CONFUSABLES sets containing Cyrillic/Greek homoglyphs that visually mimic Latin characters
  • _RTL_CHARS collections of right-to-left directional override characters that can reverse text display
  • _INVISIBLE_CHARS formatting characters that modify rendering without visible glyphs
  • _get_script_prefix() function to detect mixed-script identifiers where different Unicode blocks combine suspiciously

This detection prevents attacks where a parameter like "system" might use Cyrillic 'е' (U+0435) instead of Latin 'e' (U+0065) to bypass string matching while appearing identical to human reviewers.

TP3: Parameter Description Injection

TP3 focuses specifically on the parameters list within skill manifests, scanning for instruction overrides and data exfiltration attempts. The _check_tp3() function applies targeted regex patterns to parameter descriptions:

  • _TP3_INSTRUCTION_OVERRIDE_RE detects phrases attempting to override system prompts or modify tool behavior
  • _TP3_SYSTEM_TOKEN_RE identifies embedded system-level command tokens
  • _TP3_EXFILTRATION_RE recognizes patterns suggesting unauthorized data transmission
  • _TP3_MALICIOUS_URL_RE flags suspicious external URLs embedded in descriptions
  • _TP3_SHELL_CMD_RE catches shell command injections disguised as documentation

The analyzer also enforces length constraints to identify overly verbose descriptions that may contain hidden payloads, and validates default values for executable content.

TP4: LLM-Based Description Verification

TP4 represents the most sophisticated detection layer, using an LLM to perform semantic analysis comparing declared manifest descriptions against actual code implementation. The _check_tp4() function:

  1. Extracts code snippets from the state’s file_cache
  2. Constructs a structured prompt containing both the manifest description and executable code
  3. Invokes chat_completion() from src/skillspector/llm_utils.py to analyze behavioral consistency
  4. Parses the JSON response to identify mismatches between claimed functionality and implemented behavior
  5. Emits a Finding when the LLM reports discrepancies with sufficient confidence thresholds

This pattern catches sophisticated attacks where the manifest declares benign functionality while the underlying code implements malicious operations.

How the Analyzer Node Orchestrates Detection

The node(state) function serves as the central orchestrator, coordinating all four detection patterns across the manifest structure:

def node(state: SkillspectorState) -> AnalyzerNodeResponse:
    manifest = state.get("manifest") or {}
    if not manifest:
        return {"findings": []}
    
    # Extract all textual fields from the manifest

    metadata_texts = _extract_metadata_texts(manifest)
    
    # TP1-TP2: Scan every text field for hidden instructions and Unicode deception

    for text, field, is_id in metadata_texts:
        findings.extend(_check_tp1(text, field))
        findings.extend(_check_tp2(text, field, is_id))
    
    # TP3: Deep scan of parameter definitions

    findings.extend(_check_tp3(manifest.get("parameters") or []))
    
    # TP4: Optional LLM verification (requires external model)

    if state.get("use_llm", True):
        findings.extend(_check_tp4(state))
    
    return {"findings": findings}

The node first validates the manifest presence, then extracts all textual content using _extract_metadata_texts(). TP1 and TP2 run against every text field universally, while TP3 specifically targets the parameters array. TP4 executes only when LLM validation is enabled in the state configuration.

Practical Implementation and Code Examples

Running the Analyzer via CLI

Execute the MCP tool poisoning analyzer directly from the command line:

skillSpector analyze --analyzer mcp_tool_poisoning path/to/skill/

Programmatic Integration

Integrate the analyzer into custom Python workflows:

from skillspector.graph import build_graph
from skillspector.nodes.analyzers.mcp_tool_poisoning import node as mcp_node

# Load a skill directory into a SkillspectorState

graph = build_graph("path/to/skill")
state = graph.state

# Run only the MCP-tool-poisoning analyzer

result = mcp_node(state)
print(result["findings"])   # list of Finding objects

Inspecting Detection Results

Analyze specific findings to extract threat details:

finding = result["findings"][0]
print(f"Rule: {finding.rule_id}")
print(f"Message: {finding.message}")
print(f"Severity: {finding.severity}")
print(f"Location: {finding.matched_text}")

Key Source Files and Architecture

File Role
src/skillspector/nodes/analyzers/mcp_tool_poisoning.py Core analyzer implementing TP1-TP4 detection logic
src/skillspector/models.py Definition of Finding dataclass for structured reporting
src/skillspector/llm_utils.py Helper chat_completion() function powering TP4 analysis
src/skillspector/nodes/__init__.py Registry entry point for mcp_tool_poisoning_node
tests/test_mcp_tool_poisoning.py Comprehensive test suite verifying each detection path

Summary

  • TP1 detects hidden instructions through HTML/Markdown comments, zero-width characters, Base64 encoding, and Data-URI payloads using regex patterns in _check_tp1()
  • TP2 prevents visual spoofing via Unicode confusables, RTL overrides, and mixed-script identifiers analyzed in _check_tp2()
  • TP3 scans parameter descriptions for instruction overrides, exfiltration cues, and shell commands using targeted regexes in _check_tp3()
  • TP4 employs LLM reasoning to verify behavioral consistency between manifest descriptions and actual code implementation via _check_tp4()
  • The node() function orchestrates these checks sequentially across all manifest text fields and parameters

Frequently Asked Questions

What is MCP tool poisoning in AI skill manifests?

MCP (Malicious Capability Poisoning) tool poisoning refers to injection attacks where threat actors embed hidden instructions, deceptive Unicode characters, or malicious payloads within AI tool definitions. These manipulations can override intended tool behavior, exfiltrate data, or execute unauthorized commands when the AI processes the poisoned manifest.

How does SkillSpector detect hidden Base64 instructions?

The _check_tp1() function utilizes the _BASE64_RE regex pattern to identify Base64-encoded blobs within manifest text fields. When detected, the analyzer generates a Finding object with severity ratings indicating the presence of potentially obfuscated executable content hidden in descriptions or metadata fields.

What Unicode deception techniques does TP2 identify?

TP2 specifically targets confusable homoglyphs (Cyrillic/Greek characters visually identical to Latin), right-to-left override characters that manipulate text display direction, invisible formatting characters that alter rendering without visible presence, and mixed-script identifiers where multiple Unicode blocks combine suspiciously within single tokens.

When should I enable the LLM-based TP4 detection?

Enable TP4 when analyzing skill manifests with complex executable components where static analysis might miss behavioral inconsistencies. TP4 consumes LLM tokens via chat_completion(), making it suitable for high-security environments requiring semantic verification that the manifest description accurately reflects the underlying code implementation, particularly for tools handling sensitive data or system operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →