MCP Tool Poisoning Detection Patterns in NVIDIA SkillSpector: A Complete Guide to TP1-TP4
NVIDIA SkillSpector implements four distinct detection patterns (TP1-TP4) in src/skillspector/nodes/analyzers/mcp_tool_poisoning.py to identify hidden instructions, Unicode deception, parameter injection attacks, and LLM-based behavioral mismatches in AI skill manifests.
The SkillSpector framework provides static analysis capabilities for NVIDIA Skill Manifests, protecting against malicious capability poisoning (MCP) through specialized analyzer nodes. The MCP tool poisoning analyzer examines every textual field within a skill manifest—ranging from names and descriptions to parameter metadata—to identify four specific attack vectors that threat actors use to compromise AI tool behavior.
Understanding the Four MCP Tool Poisoning Detection Patterns
The analyzer located at src/skillspector/nodes/analyzers/mcp_tool_poisoning.py implements a multi-layered defense strategy through four specialized detection patterns. Each pattern targets distinct attack methodologies used to inject malicious instructions into tool definitions.
TP1: Hidden Instructions Detection
TP1 identifies concealed payloads embedded within manifest text fields that remain invisible during casual inspection but executable by parsing systems. The _check_tp1() function employs multiple regex patterns to detect:
- HTML comments (
<!--...-->) using_HTML_COMMENT_RE - Markdown comments (
[//]: #(...)) using_MARKDOWN_COMMENT_RE - Zero-width Unicode characters (invisible joiners and non-printable code points) using
_ZERO_WIDTH_RE - Base64-encoded blobs using
_BASE64_REto catch obfuscated data payloads - Data-URI schemes using
_DATA_URI_REto detect embedded executable content
When the analyzer detects any of these patterns in metadata fields, it generates a Finding object documenting the specific match location and threat type.
TP2: Unicode Deception and Homoglyph Attacks
TP2 targets visual spoofing attacks that use Unicode confusables to masquerade malicious parameters as legitimate system variables. The _check_tp2() function in mcp_tool_poisoning.py performs script analysis using:
_CONFUSABLESsets containing Cyrillic/Greek homoglyphs that visually mimic Latin characters_RTL_CHARScollections of right-to-left directional override characters that can reverse text display_INVISIBLE_CHARSformatting characters that modify rendering without visible glyphs_get_script_prefix()function to detect mixed-script identifiers where different Unicode blocks combine suspiciously
This detection prevents attacks where a parameter like "system" might use Cyrillic 'е' (U+0435) instead of Latin 'e' (U+0065) to bypass string matching while appearing identical to human reviewers.
TP3: Parameter Description Injection
TP3 focuses specifically on the parameters list within skill manifests, scanning for instruction overrides and data exfiltration attempts. The _check_tp3() function applies targeted regex patterns to parameter descriptions:
_TP3_INSTRUCTION_OVERRIDE_REdetects phrases attempting to override system prompts or modify tool behavior_TP3_SYSTEM_TOKEN_REidentifies embedded system-level command tokens_TP3_EXFILTRATION_RErecognizes patterns suggesting unauthorized data transmission_TP3_MALICIOUS_URL_REflags suspicious external URLs embedded in descriptions_TP3_SHELL_CMD_REcatches shell command injections disguised as documentation
The analyzer also enforces length constraints to identify overly verbose descriptions that may contain hidden payloads, and validates default values for executable content.
TP4: LLM-Based Description Verification
TP4 represents the most sophisticated detection layer, using an LLM to perform semantic analysis comparing declared manifest descriptions against actual code implementation. The _check_tp4() function:
- Extracts code snippets from the state’s
file_cache - Constructs a structured prompt containing both the manifest description and executable code
- Invokes
chat_completion()fromsrc/skillspector/llm_utils.pyto analyze behavioral consistency - Parses the JSON response to identify mismatches between claimed functionality and implemented behavior
- Emits a
Findingwhen the LLM reports discrepancies with sufficient confidence thresholds
This pattern catches sophisticated attacks where the manifest declares benign functionality while the underlying code implements malicious operations.
How the Analyzer Node Orchestrates Detection
The node(state) function serves as the central orchestrator, coordinating all four detection patterns across the manifest structure:
def node(state: SkillspectorState) -> AnalyzerNodeResponse:
manifest = state.get("manifest") or {}
if not manifest:
return {"findings": []}
# Extract all textual fields from the manifest
metadata_texts = _extract_metadata_texts(manifest)
# TP1-TP2: Scan every text field for hidden instructions and Unicode deception
for text, field, is_id in metadata_texts:
findings.extend(_check_tp1(text, field))
findings.extend(_check_tp2(text, field, is_id))
# TP3: Deep scan of parameter definitions
findings.extend(_check_tp3(manifest.get("parameters") or []))
# TP4: Optional LLM verification (requires external model)
if state.get("use_llm", True):
findings.extend(_check_tp4(state))
return {"findings": findings}
The node first validates the manifest presence, then extracts all textual content using _extract_metadata_texts(). TP1 and TP2 run against every text field universally, while TP3 specifically targets the parameters array. TP4 executes only when LLM validation is enabled in the state configuration.
Practical Implementation and Code Examples
Running the Analyzer via CLI
Execute the MCP tool poisoning analyzer directly from the command line:
skillSpector analyze --analyzer mcp_tool_poisoning path/to/skill/
Programmatic Integration
Integrate the analyzer into custom Python workflows:
from skillspector.graph import build_graph
from skillspector.nodes.analyzers.mcp_tool_poisoning import node as mcp_node
# Load a skill directory into a SkillspectorState
graph = build_graph("path/to/skill")
state = graph.state
# Run only the MCP-tool-poisoning analyzer
result = mcp_node(state)
print(result["findings"]) # list of Finding objects
Inspecting Detection Results
Analyze specific findings to extract threat details:
finding = result["findings"][0]
print(f"Rule: {finding.rule_id}")
print(f"Message: {finding.message}")
print(f"Severity: {finding.severity}")
print(f"Location: {finding.matched_text}")
Key Source Files and Architecture
| File | Role |
|---|---|
src/skillspector/nodes/analyzers/mcp_tool_poisoning.py |
Core analyzer implementing TP1-TP4 detection logic |
src/skillspector/models.py |
Definition of Finding dataclass for structured reporting |
src/skillspector/llm_utils.py |
Helper chat_completion() function powering TP4 analysis |
src/skillspector/nodes/__init__.py |
Registry entry point for mcp_tool_poisoning_node |
tests/test_mcp_tool_poisoning.py |
Comprehensive test suite verifying each detection path |
Summary
- TP1 detects hidden instructions through HTML/Markdown comments, zero-width characters, Base64 encoding, and Data-URI payloads using regex patterns in
_check_tp1() - TP2 prevents visual spoofing via Unicode confusables, RTL overrides, and mixed-script identifiers analyzed in
_check_tp2() - TP3 scans parameter descriptions for instruction overrides, exfiltration cues, and shell commands using targeted regexes in
_check_tp3() - TP4 employs LLM reasoning to verify behavioral consistency between manifest descriptions and actual code implementation via
_check_tp4() - The
node()function orchestrates these checks sequentially across all manifest text fields and parameters
Frequently Asked Questions
What is MCP tool poisoning in AI skill manifests?
MCP (Malicious Capability Poisoning) tool poisoning refers to injection attacks where threat actors embed hidden instructions, deceptive Unicode characters, or malicious payloads within AI tool definitions. These manipulations can override intended tool behavior, exfiltrate data, or execute unauthorized commands when the AI processes the poisoned manifest.
How does SkillSpector detect hidden Base64 instructions?
The _check_tp1() function utilizes the _BASE64_RE regex pattern to identify Base64-encoded blobs within manifest text fields. When detected, the analyzer generates a Finding object with severity ratings indicating the presence of potentially obfuscated executable content hidden in descriptions or metadata fields.
What Unicode deception techniques does TP2 identify?
TP2 specifically targets confusable homoglyphs (Cyrillic/Greek characters visually identical to Latin), right-to-left override characters that manipulate text display direction, invisible formatting characters that alter rendering without visible presence, and mixed-script identifiers where multiple Unicode blocks combine suspiciously within single tokens.
When should I enable the LLM-based TP4 detection?
Enable TP4 when analyzing skill manifests with complex executable components where static analysis might miss behavioral inconsistencies. TP4 consumes LLM tokens via chat_completion(), making it suitable for high-security environments requiring semantic verification that the manifest description accurately reflects the underlying code implementation, particularly for tools handling sensitive data or system operations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →