MCP Tool Poisoning Patterns Detected by NVIDIA SkillSpector: TP1-TP4 Analysis
NVIDIA SkillSpector detects four distinct MCP tool poisoning patterns—TP1 (hidden instructions), TP2 (Unicode deception), TP3 (parameter injection), and TP4 (semantic mismatch)—using a dedicated static analyzer that scans skill manifests for malicious payloads embedded in text fields, Unicode characters, parameter descriptions, and code-to-description inconsistencies.
The MCP tool poisoning analyzer in NVIDIA/SkillSpector provides static analysis capabilities specifically designed to identify malicious capability poisoning (MCP) attempts within NVIDIA Skill Manifests. Located at src/skillspector/nodes/analyzers/mcp_tool_poisoning.py, this analyzer extracts all textual fields—including names, descriptions, triggers, and parameter metadata—and applies four specialized detection heuristics to identify hidden instructions, visual spoofing attacks, injection attempts, and semantic drift.
TP1 – Hidden Instructions in Manifest Text
TP1 targets covert payloads concealed within standard manifest text fields. The _check_tp1() function scans for steganographic techniques attackers use to embed instructions invisible to casual review but potentially executable by LLM parsers.
The analyzer detects:
- HTML comments (
<!--...-->) using_HTML_COMMENT_RE - Markdown comments (
[//]: #(...)) via_MARKDOWN_COMMENT_RE - Zero-width Unicode characters (invisible control characters) matched by
_ZERO_WIDTH_RE - Base64 encoded blobs identified with
_BASE64_RE - Data URI payloads caught by
_DATA_URI_RE
When any pattern matches, the analyzer generates a Finding object (defined in src/skillspector/models.py) that includes the matched text, field location, and severity rating.
TP2 – Unicode Deception and Homoglyph Attacks
TP2 identifies visual spoofing attacks that rely on Unicode confusables and directional overrides to trick developers or automated systems. The _check_tp2() function implements multi-layered Unicode analysis to catch homoglyph attacks and bidirectional text exploits.
The detection covers:
- Confusable Cyrillic/Greek homoglyphs matched against
_CONFUSABLES - RTL (right-to-left) directional override characters detected via
_RTL_CHARS - Invisible formatting characters identified by
_INVISIBLE_CHARS - Mixed-script identifiers detected through
_get_script_prefix(), which flags when a single identifier mixes Latin with non-Latin scripts
This check runs across all metadata fields, with special handling for identifier fields (where homoglyph attacks are most dangerous).
TP3 – Parameter Description Injection
TP3 focuses on the parameters array within skill manifests, where attackers may inject instruction overrides or malicious defaults. The _check_tp3() function iterates through parameter names and descriptions looking for active injection patterns.
Specifically, it applies regex patterns to detect:
- Instruction override phrases (
_TP3_INSTRUCTION_OVERRIDE_RE) - System prompt tokens (
_TP3_SYSTEM_TOKEN_RE) - Data exfiltration cues (
_TP3_EXFILTRATION_RE) - Malicious URLs embedded in defaults (
_TP3_MALICIOUS_URL_RE) - Shell command injection (
_TP3_SHELL_CMD_RE)
Unlike TP1 and TP2, this check specifically targets the parameters list extracted from the manifest, as these fields often get passed directly to LLM context windows or execution environments.
TP4 – LLM-Based Semantic Analysis
TP4 represents the semantic layer of detection, using an LLM to verify that code behavior matches declared descriptions. The _check_tp4() function collects executable code snippets from the state's file_cache, constructs a structured prompt comparing the manifest description against the actual implementation, and invokes chat_completion() from src/skillspector/llm_utils.py.
The LLM evaluates whether:
- Tool descriptions accurately reflect the code's functionality
- Trigger conditions align with actual execution paths
- Permission requirements match the scope of operations performed
Results are parsed from JSON responses, and findings are emitted only when the model reports a mismatch with sufficient confidence. This check is optional and controlled by the use_llm state flag (defaults to True).
Running the MCP Tool Poisoning Analyzer
Execute the analyzer from the command line to scan a skill directory:
skillSpector analyze --analyzer mcp_tool_poisoning path/to/skill/
For programmatic integration within the SkillSpector graph framework:
from skillspector.graph import build_graph
from skillspector.nodes.analyzers.mcp_tool_poisoning import node as mcp_node
# Load skill into state
graph = build_graph("path/to/skill")
state = graph.state
# Execute TP1-TP4 checks
result = mcp_node(state)
# Inspect findings
for finding in result["findings"]:
print(f"Rule: {finding.rule_id}")
print(f"Message: {finding.message}")
print(f"Severity: {finding.severity}")
print(f"Location: {finding.matched_text}")
The analyzer returns a list of Finding objects, each containing rule_id, message, severity, and matched_text attributes that pinpoint the exact location and nature of the detected MCP pattern.
Summary
- TP1 detects hidden instructions via HTML comments, Markdown comments, zero-width characters, Base64 blobs, and Data URIs using regex patterns in
_check_tp1(). - TP2 identifies Unicode deception through confusable homoglyphs, RTL overrides, invisible characters, and mixed-script identifiers in
_check_tp2(). - TP3 scans parameter descriptions for injection attempts including instruction overrides, system tokens, exfiltration cues, and malicious URLs/shell commands via
_check_tp3(). - TP4 leverages LLM semantic analysis to detect description-to-code mismatches using
chat_completion()in_check_tp4(). - All four patterns are orchestrated by the
node()function insrc/skillspector/nodes/analyzers/mcp_tool_poisoning.py, which extracts textual fields and routes them to appropriate check functions.
Frequently Asked Questions
What is the difference between TP1 and TP3 detection in SkillSpector?
TP1 focuses on steganographic techniques—hidden data embedded in any manifest text field using comments, zero-width characters, or encoding schemes. TP3 specifically targets active injection attacks within parameter descriptions, such as instruction overrides or malicious default values that could execute during tool invocation. While TP1 finds concealed payloads, TP3 finds explicit manipulation attempts in parameter metadata.
How does SkillSpector detect homoglyph attacks in skill manifests?
The analyzer uses _check_tp2() to scan for confusable characters (Cyrillic/Greek letters that look like Latin), RTL directional overrides that can reorder text display, and invisible formatting characters. It also employs _get_script_prefix() to detect mixed-script identifiers where a single word combines Latin with non-Latin scripts— a common technique for creating visually deceptive function or parameter names.
Can I run the MCP tool poisoning analyzer without LLM capabilities?
Yes. Set use_llm: False in the SkillspectorState before invoking the node. This disables TP4 (semantic analysis) while preserving TP1-TP3 static checks. The code explicitly checks state.get("use_llm", True) before calling _check_tp4(), allowing the analyzer to run in environments without LLM access or when you need faster, rule-based-only scanning.
Where does SkillSpector store the code snippets used for TP4 validation?
The analyzer retrieves code from the state's file_cache, which contains executable components discovered during the graph building phase. These snippets are passed to chat_completion() in src/skillspector/llm_utils.py along with the manifest description to perform the semantic comparison that identifies description-to-behavior mismatches.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →