SkillSpector Vulnerability Categories and Patterns: Complete Guide to 17 Categories and 68 Detection Rules

SkillSpector detects 17 vulnerability categories defined in the PatternCategory enum and 68 distinct detection patterns (rule IDs) that map to specific security risks in LLM skills, ranging from prompt injection to supply chain attacks.

NVIDIA's SkillSpector is an open-source security scanner designed to audit LLM skills and agents. Understanding the complete taxonomy of SkillSpector vulnerability categories and patterns is essential for developers auditing their code and security teams interpreting SARIF reports.

The 17 PatternCategory Enum Values

SkillSpector tags each finding with a category taken from the PatternCategory enum defined in [src/skillspector/nodes/analyzers/pattern_defaults.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/pattern_defaults.py#L24-L44). The 17 categories cover the full spectrum of LLM security risks:

  • Prompt Injection – Skill tries to make the LLM ignore or override its own safety prompts.
  • Data Exfiltration – Skill leaks data (e.g., env vars, files, conversation context).
  • Privilege Escalation – Skill requests or uses higher-privilege capabilities than needed.
  • Supply Chain – Vulnerable or abandoned third-party dependencies.
  • Excessive Agency – Unrestricted tool access or autonomous high-impact decisions.
  • Output Handling – Unsanitised model output fed into another security context.
  • System Prompt Leakage – Skill reveals or extracts the system prompt.
  • Memory Poisoning – Skill injects content that persists in the LLM's memory.
  • Tool Misuse – Unsafe parameters or chaining of tools.
  • Rogue Agent – Self-modifying code or persistence mechanisms.
  • Trigger Abuse – Over-broad or conflicting trigger patterns.
  • YARA Match – Known malware signatures detected by YARA rules.
  • MCP Least Privilege – Capability-declared permission mismatches.
  • MCP Tool Poisoning – Hidden or deceptive metadata in the skill manifest.
  • Agent Snooping – Reads agent configuration or other skill files.
  • Anti-Refusal – Instructions that force the LLM never to refuse.
  • Server-Side Request Forgery – Network requests that can reach internal services/metadata.

The 68 Detection Patterns (Rule IDs)

Each pattern is a rule ID (e.g., P1, E2, LP3) defined across three dictionaries in pattern_defaults.py: DEFAULT_EXPLANATIONS (lines 46-138), RULE_ID_TO_CATEGORY (lines 141-213), and PATTERN_NAMES (lines 216-288). The 68 patterns are grouped by category as follows:

Prompt Injection (P1-P5)

  • P1 – Override Instructions: Attempts to override system instructions or ignore safety constraints.
  • P2 – Hidden Instructions: Hidden instructions in comments or invisible text.
  • P3 – External Transmission Instructions: Directs the agent to send conversation context or user data to external services.
  • P4 – Subtle Steering: Subtle instructions that may alter agent decision-making.
  • P5 – Harmful Content: Content that could cause physical harm if followed.

System Prompt Leakage (P6-P8)

  • P6 – System Prompt Leakage: Direct exposure of system prompts or internal rules.
  • P7 – System Prompt Leakage: Indirect extraction via rephrasing, translation, or summarisation.
  • P8 – System Prompt Leakage: Exfiltration of system prompts via tool calls.

Data Exfiltration (E1-E4)

  • E1 – External Transmission: Data sent to an external URL; may be telemetry or exfiltration.
  • E2 – Env Variable Harvesting: Access to environment variables that may contain secrets.
  • E3 – File System Enumeration: Scans directories for sensitive files.
  • E4 – Conversation Context Leak: Leaks agent conversation context to external services.

Privilege Escalation (PE1-PE3)

  • PE1 – Excessive Permissions: Skill requests more permissions than needed.
  • PE2 – Sudo/Root Invocation: Uses sudo or root privileges.
  • PE3 – Credential File Access: Accesses credential files (SSH keys, AWS credentials, etc.).

Supply Chain (SC1-SC6)

  • SC1 – Unpinned Dependencies: Dependencies lack version pinning.
  • SC2 – Remote Code Execution: Remote code downloaded and executed.
  • SC3 – Obfuscated Code: Base64/hex encoded code with execution.
  • SC4 – Known Vulnerable Dependency: Dependency has known CVEs.
  • SC5 – Abandoned Dependency: Dependency appears unmaintained.
  • SC6 – Typosquatting Dependency: Package name resembles a popular one (possible typo-squatting).

Excessive Agency (EA1-EA4)

  • EA1 – Unrestricted Tool Access: Skill grants unrestricted tool access.
  • EA2 – Autonomous Decision Making: High-impact decisions without human-in-the-loop.
  • EA3 – Scope Creep: Skill performs actions outside its stated purpose.
  • EA4 – Unbounded Resource Access: No limits on API calls, storage, compute, etc.

Output Handling (OH1-OH3)

  • OH1 – Unvalidated Output Injection: Model output used without validation/sanitisation.
  • OH2 – Cross-Context Output: Output moved between security contexts without checks.
  • OH3 – Unbounded Output: No limits on output size or generation rate.

Memory Poisoning (MP1-MP3)

  • MP1 – Persistent Context Injection: Injects content that persists across interactions.
  • MP2 – Context Window Stuffing: Filler content displaces legitimate instructions.
  • MP3 – Memory Manipulation: Manipulates agent memory/state.

Tool Misuse (TM1-TM3)

  • TM1 – Tool Parameter Abuse: Unsafe tool parameters (e.g., shell=True).
  • TM2 – Chaining Abuse: Chains tools to bypass safety checks.
  • TM3 – Unsafe Defaults: Default tool settings that are insecure.

Rogue Agent (RA1-RA2)

  • RA1 – Self-Modification: Skill modifies its own code or config at runtime.
  • RA2 – Session Persistence: Creates cron jobs, startup scripts, or state files.

MCP Least Privilege (LP1-LP4)

  • LP1 – Underdeclared Capability: Uses capabilities not covered by declared permissions.
  • LP2 – Wildcard Permission: Permission list contains a wildcard (*).
  • LP3 – Missing Permission Declaration: No permissions field but capabilities are used.
  • LP4 – Overdeclared Permission: Permission declared but no corresponding code capability.

MCP Tool Poisoning (TP1-TP4)

  • TP1 – Hidden Instructions: Hidden instructions in skill metadata.
  • TP2 – Unicode Deception: Homoglyphs, RTL overrides, invisible characters.
  • TP3 – Parameter Description Injection: Malicious content in parameter descriptions/defaults.
  • TP4 – Description-Behavior Mismatch: Skill description does not match actual code behavior.

Agent Snooping (AS1-AS3)

  • AS1 – Agent Config Directory Access: Reads .claude/, .codex/, .gemini/ directories.
  • AS2 – MCP Config Access: Reads mcp.json (server URLs, tokens, tool definitions).
  • AS3 – Skill Enumeration: Reads other installed skills' SKILL.md files.

Anti-Refusal (AR1-AR3)

  • AR1 – Refusal Suppression: Instructs the agent to never refuse.
  • AR2 – Disclaimer Suppression: Instructs the agent to omit warnings/disclaimers.
  • AR3 – Safety Policy Nullification: Directly attempts to nullify safety policies.

Server-Side Request Forgery (SSRF1-SSRF3)

  • SSRF1 – Cloud Metadata Access: Accesses cloud instance metadata (e.g., 169.254.169.254).
  • SSRF2 – Internal Network Request: Requests to loopback, link-local or private-range hosts.
  • SSRF3 – Dynamic Request Target: Builds request target from untrusted data.

YARA Match (YR1-YR4)

  • YR1 – Malware Signature: YARA match for known malware (reverse shell, backdoor, etc.).
  • YR2 – Webshell Detected: YARA match for known webshell patterns.
  • YR3 – Crypto Miner Detected: YARA match for cryptocurrency mining indicators.
  • YR4 – Hack Tool / Exploit Detected: YARA match for offensive tools or exploit frameworks.

Behavioral AST Patterns (AST1-AST9)

The [src/skillspector/nodes/analyzers/behavioral_ast.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/behavioral_ast.py) analyzer detects 9 additional patterns (AST1 through AST9) covering execution-related evasion techniques. These are counted within the 68 total patterns but are typically collapsed into a single "AST" family in public documentation.

How SkillSpector Maps Rules to Categories

Static pattern analyzers (e.g., static_patterns_prompt_injection.py, static_patterns_ssrf.py) import PatternCategory and call get_category(rule_id) to tag findings. Semantic analyzers (e.g., semantic_developer_intent.py) also map detections to the same categories via helper functions in pattern_defaults.py.

The Meta-Analyzer in [src/skillspector/nodes/meta_analyzer.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/meta_analyzer.py#L262-L298) collects every finding, reads the rule's category, explanation, and remediation, and writes them into the SARIF report.

Practical Usage and SARIF Output

Run SkillSpector on a local skill directory:

skill-spector analyze ./my_skill --output report.sarif

Inspect the SARIF report for a specific pattern:

jq '.runs[0].results[] | select(.ruleId=="P1")' report.sarif

The output contains structured entries:

{
  "ruleId": "P1",
  "level": "error",
  "message": {
    "text": "This pattern attempts to override system instructions or ignore safety constraints..."
  },
  "properties": {
    "category": "Prompt Injection",
    "remediation": "Remove or rewrite any text that instructs the agent to ignore prompts..."
  }
}

Summary

  • SkillSpector recognizes 17 vulnerability categories defined in the PatternCategory enum at lines 24-44 of pattern_defaults.py.
  • The tool detects 68 distinct pattern IDs mapped through DEFAULT_EXPLANATIONS, RULE_ID_TO_CATEGORY, and PATTERN_NAMES dictionaries.
  • Patterns span critical security domains including prompt injection, data exfiltration, supply chain risks, and MCP-specific vulnerabilities.
  • Static analyzers and the behavioral AST analyzer generate findings that the Meta-Analyzer aggregates into SARIF format with full remediation guidance.

Frequently Asked Questions

What file contains the PatternCategory enum in SkillSpector?

The PatternCategory enum is defined in [src/skillspector/nodes/analyzers/pattern_defaults.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/pattern_defaults.py#L24-L44) at lines 24-44. This file serves as the central registry for all 17 vulnerability categories used to classify detections.

How do I map a SkillSpector rule ID to its vulnerability category?

SkillSpector uses the RULE_ID_TO_CATEGORY dictionary in pattern_defaults.py (lines 141-213) to map rule IDs like P1 or E2 to their respective PatternCategory values. Static analyzers call get_category(rule_id) to resolve this mapping automatically during analysis.

What is the difference between static patterns and behavioral AST patterns in SkillSpector?

Static patterns (P1-YR4) are detected through signature-based analysis of source code and metadata, while behavioral AST patterns (AST1-AST9) are identified by the behavioral_ast.py analyzer through execution-related evasion detection in the abstract syntax tree. Both use the same category taxonomy but AST patterns focus on runtime code behavior.

How does SkillSpector output pattern detections in SARIF format?

The Meta-Analyzer in meta_analyzer.py (lines 262-298) aggregates findings from all analyzers, resolves each rule ID to its category, explanation, and remediation strings from pattern_defaults.py, and writes them to a SARIF-formatted JSON file with the ruleId, level, and properties fields populated.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →