# MCP Tool Poisoning Detection Patterns in NVIDIA SkillSpector: A Complete Guide to TP1-TP4

> Explore MCP tool poisoning detection patterns TP1-TP4 in NVIDIA SkillSpector. Learn to identify hidden instructions, Unicode deception, parameter injection, and LLM mismatches in AI skill manifests.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: deep-dive
- Published: 2026-07-09

---

**NVIDIA SkillSpector implements four distinct detection patterns (TP1-TP4) in [`src/skillspector/nodes/analyzers/mcp_tool_poisoning.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/mcp_tool_poisoning.py) to identify hidden instructions, Unicode deception, parameter injection attacks, and LLM-based behavioral mismatches in AI skill manifests.**

The SkillSpector framework provides static analysis capabilities for NVIDIA Skill Manifests, protecting against malicious capability poisoning (MCP) through specialized analyzer nodes. The MCP tool poisoning analyzer examines every textual field within a skill manifest—ranging from names and descriptions to parameter metadata—to identify four specific attack vectors that threat actors use to compromise AI tool behavior.

## Understanding the Four MCP Tool Poisoning Detection Patterns

The analyzer located at [`src/skillspector/nodes/analyzers/mcp_tool_poisoning.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/mcp_tool_poisoning.py) implements a multi-layered defense strategy through four specialized detection patterns. Each pattern targets distinct attack methodologies used to inject malicious instructions into tool definitions.

### TP1: Hidden Instructions Detection

**TP1** identifies concealed payloads embedded within manifest text fields that remain invisible during casual inspection but executable by parsing systems. The `_check_tp1()` function employs multiple regex patterns to detect:

- **HTML comments** (`<!--...-->`) using `_HTML_COMMENT_RE`
- **Markdown comments** (`[//]: #(...)`) using `_MARKDOWN_COMMENT_RE`
- **Zero-width Unicode characters** (invisible joiners and non-printable code points) using `_ZERO_WIDTH_RE`
- **Base64-encoded blobs** using `_BASE64_RE` to catch obfuscated data payloads
- **Data-URI schemes** using `_DATA_URI_RE` to detect embedded executable content

When the analyzer detects any of these patterns in metadata fields, it generates a `Finding` object documenting the specific match location and threat type.

### TP2: Unicode Deception and Homoglyph Attacks

**TP2** targets visual spoofing attacks that use Unicode confusables to masquerade malicious parameters as legitimate system variables. The `_check_tp2()` function in [`mcp_tool_poisoning.py`](https://github.com/NVIDIA/SkillSpector/blob/main/mcp_tool_poisoning.py) performs script analysis using:

- **`_CONFUSABLES`** sets containing Cyrillic/Greek homoglyphs that visually mimic Latin characters
- **`_RTL_CHARS`** collections of right-to-left directional override characters that can reverse text display
- **`_INVISIBLE_CHARS`** formatting characters that modify rendering without visible glyphs
- **`_get_script_prefix()`** function to detect mixed-script identifiers where different Unicode blocks combine suspiciously

This detection prevents attacks where a parameter like "system" might use Cyrillic 'е' (U+0435) instead of Latin 'e' (U+0065) to bypass string matching while appearing identical to human reviewers.

### TP3: Parameter Description Injection

**TP3** focuses specifically on the `parameters` list within skill manifests, scanning for instruction overrides and data exfiltration attempts. The `_check_tp3()` function applies targeted regex patterns to parameter descriptions:

- **`_TP3_INSTRUCTION_OVERRIDE_RE`** detects phrases attempting to override system prompts or modify tool behavior
- **`_TP3_SYSTEM_TOKEN_RE`** identifies embedded system-level command tokens
- **`_TP3_EXFILTRATION_RE`** recognizes patterns suggesting unauthorized data transmission
- **`_TP3_MALICIOUS_URL_RE`** flags suspicious external URLs embedded in descriptions
- **`_TP3_SHELL_CMD_RE`** catches shell command injections disguised as documentation

The analyzer also enforces length constraints to identify overly verbose descriptions that may contain hidden payloads, and validates default values for executable content.

### TP4: LLM-Based Description Verification

**TP4** represents the most sophisticated detection layer, using an LLM to perform semantic analysis comparing declared manifest descriptions against actual code implementation. The `_check_tp4()` function:

1. Extracts code snippets from the state’s `file_cache`
2. Constructs a structured prompt containing both the manifest description and executable code
3. Invokes `chat_completion()` from [`src/skillspector/llm_utils.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_utils.py) to analyze behavioral consistency
4. Parses the JSON response to identify mismatches between claimed functionality and implemented behavior
5. Emits a `Finding` when the LLM reports discrepancies with sufficient confidence thresholds

This pattern catches sophisticated attacks where the manifest declares benign functionality while the underlying code implements malicious operations.

## How the Analyzer Node Orchestrates Detection

The `node(state)` function serves as the central orchestrator, coordinating all four detection patterns across the manifest structure:

```python
def node(state: SkillspectorState) -> AnalyzerNodeResponse:
    manifest = state.get("manifest") or {}
    if not manifest:
        return {"findings": []}
    
    # Extract all textual fields from the manifest

    metadata_texts = _extract_metadata_texts(manifest)
    
    # TP1-TP2: Scan every text field for hidden instructions and Unicode deception

    for text, field, is_id in metadata_texts:
        findings.extend(_check_tp1(text, field))
        findings.extend(_check_tp2(text, field, is_id))
    
    # TP3: Deep scan of parameter definitions

    findings.extend(_check_tp3(manifest.get("parameters") or []))
    
    # TP4: Optional LLM verification (requires external model)

    if state.get("use_llm", True):
        findings.extend(_check_tp4(state))
    
    return {"findings": findings}

```

The node first validates the manifest presence, then extracts all textual content using `_extract_metadata_texts()`. TP1 and TP2 run against every text field universally, while TP3 specifically targets the parameters array. TP4 executes only when LLM validation is enabled in the state configuration.

## Practical Implementation and Code Examples

### Running the Analyzer via CLI

Execute the MCP tool poisoning analyzer directly from the command line:

```bash
skillSpector analyze --analyzer mcp_tool_poisoning path/to/skill/

```

### Programmatic Integration

Integrate the analyzer into custom Python workflows:

```python
from skillspector.graph import build_graph
from skillspector.nodes.analyzers.mcp_tool_poisoning import node as mcp_node

# Load a skill directory into a SkillspectorState

graph = build_graph("path/to/skill")
state = graph.state

# Run only the MCP-tool-poisoning analyzer

result = mcp_node(state)
print(result["findings"])   # list of Finding objects

```

### Inspecting Detection Results

Analyze specific findings to extract threat details:

```python
finding = result["findings"][0]
print(f"Rule: {finding.rule_id}")
print(f"Message: {finding.message}")
print(f"Severity: {finding.severity}")
print(f"Location: {finding.matched_text}")

```

## Key Source Files and Architecture

| File | Role |
|------|------|
| [`src/skillspector/nodes/analyzers/mcp_tool_poisoning.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/mcp_tool_poisoning.py) | Core analyzer implementing TP1-TP4 detection logic |
| [`src/skillspector/models.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/models.py) | Definition of `Finding` dataclass for structured reporting |
| [`src/skillspector/llm_utils.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_utils.py) | Helper `chat_completion()` function powering TP4 analysis |
| [`src/skillspector/nodes/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/__init__.py) | Registry entry point for `mcp_tool_poisoning_node` |
| [`tests/test_mcp_tool_poisoning.py`](https://github.com/NVIDIA/SkillSpector/blob/main/tests/test_mcp_tool_poisoning.py) | Comprehensive test suite verifying each detection path |

## Summary

- **TP1** detects hidden instructions through HTML/Markdown comments, zero-width characters, Base64 encoding, and Data-URI payloads using regex patterns in `_check_tp1()`
- **TP2** prevents visual spoofing via Unicode confusables, RTL overrides, and mixed-script identifiers analyzed in `_check_tp2()`
- **TP3** scans parameter descriptions for instruction overrides, exfiltration cues, and shell commands using targeted regexes in `_check_tp3()`
- **TP4** employs LLM reasoning to verify behavioral consistency between manifest descriptions and actual code implementation via `_check_tp4()`
- The `node()` function orchestrates these checks sequentially across all manifest text fields and parameters

## Frequently Asked Questions

### What is MCP tool poisoning in AI skill manifests?

MCP (Malicious Capability Poisoning) tool poisoning refers to injection attacks where threat actors embed hidden instructions, deceptive Unicode characters, or malicious payloads within AI tool definitions. These manipulations can override intended tool behavior, exfiltrate data, or execute unauthorized commands when the AI processes the poisoned manifest.

### How does SkillSpector detect hidden Base64 instructions?

The `_check_tp1()` function utilizes the `_BASE64_RE` regex pattern to identify Base64-encoded blobs within manifest text fields. When detected, the analyzer generates a `Finding` object with severity ratings indicating the presence of potentially obfuscated executable content hidden in descriptions or metadata fields.

### What Unicode deception techniques does TP2 identify?

TP2 specifically targets **confusable homoglyphs** (Cyrillic/Greek characters visually identical to Latin), **right-to-left override characters** that manipulate text display direction, **invisible formatting characters** that alter rendering without visible presence, and **mixed-script identifiers** where multiple Unicode blocks combine suspiciously within single tokens.

### When should I enable the LLM-based TP4 detection?

Enable TP4 when analyzing skill manifests with complex executable components where static analysis might miss behavioral inconsistencies. TP4 consumes LLM tokens via `chat_completion()`, making it suitable for high-security environments requiring semantic verification that the manifest description accurately reflects the underlying code implementation, particularly for tools handling sensitive data or system operations.