# How SkillSpector Detects Vulnerabilities in AI Agent Skills: A Three-Stage Pipeline

> Discover how SkillSpector detects vulnerabilities in AI agent skills using a three-stage pipeline combining static analysis and LLM meta-analysis for actionable insights.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: deep-dive
- Published: 2026-07-11

---

**SkillSpector detects vulnerabilities in AI agent skills by combining deterministic static pattern matching with LLM-driven meta-analysis to filter false positives and enrich findings with actionable remediation guidance.**

NVIDIA's SkillSpector is an open-source security scanner designed specifically to detect vulnerabilities in AI agent skills. It employs a hybrid analysis pipeline that blends regex-based static analysis with generative AI validation, ensuring comprehensive coverage while minimizing noise through strict structured output schemas.

## The Three-Stage Detection Pipeline

SkillSpector implements a directed acyclic graph (DAG) that orchestrates three tightly coupled stages. Each stage refines the previous output, moving from broad pattern matching to nuanced vulnerability assessment.

### Stage 1: Static Pattern Analysis

The first stage scans every file in a skill—code, markdown, prompts, and configuration—using regular-expression-based rules. These rules target known risky constructs such as prompt-injection directives, hidden Unicode tags, insecure tool usage, SSRF URLs, and privilege-escalation calls.

Each rule lives in a dedicated analyzer module under `src/skillspector/nodes/analyzers/`. For example, prompt-injection detection is handled in [`src/skillspector/nodes/analyzers/static_patterns_prompt_injection.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/static_patterns_prompt_injection.py), which defines regex patterns P1 through P4 alongside confidence scores. The `analyze` function walks file content and produces `AnalyzerFinding` objects containing the rule ID, location, confidence level, and contextual snippet.

### Stage 2: LLM-Driven Meta-Analysis

After static analysis, raw findings are passed to a large language model (LLM) that acts as a meta-analyzer. This stage filters false positives and enriches true vulnerabilities with intent classification, impact severity, explanation, and remediation steps.

The implementation resides in [`src/skillspector/nodes/meta_analyzer.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/meta_analyzer.py). This node constructs a structured prompt using `PER_FILE_ANALYSIS_PROMPT`, injecting skill metadata, file content, and static findings. It calls the LLM through `LLMAnalyzerBase` with a Pydantic schema `MetaAnalyzerResult` to enforce type-safe JSON output. The schema guarantees that each returned finding includes boolean `is_vulnerability`, string fields for `intent` (malicious/negligent/benign) and `impact` (critical/high/medium/low), plus concise remediation guidance.

### Stage 3: Graph Orchestration and Reporting

The final stage aggregates outputs from all analyzers into a unified report. The orchestration code in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) creates a `SkillspectorState`, registers each analyzer node (including `MetaAnalyzer`), and executes `await state.run()` to traverse the DAG. After execution, the `Report` node formats consolidated results into SARIF or human-readable console output, accessible via the `skill-spector scan <path>` command defined in [`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py).

## How the Detection Flow Works in Practice

The complete workflow executes in four concrete steps when scanning a skill directory:

1. **Skill ingestion** – The CLI reads [`manifest.json`](https://github.com/NVIDIA/SkillSpector/blob/main/manifest.json) or [`SKILL.md`](https://github.com/NVIDIA/SkillSpector/blob/main/SKILL.md) and walks the directory, loading each file into memory.
2. **Static analysis** – For every file, the suite of static analyzers runs in parallel. Each analyzer contributes an `AnalyzerFinding` with a category such as `PatternCategory.PROMPT_INJECTION` defined in [`src/skillspector/nodes/analyzers/pattern_defaults.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/pattern_defaults.py).
3. **LLM meta-analysis** – All findings for a single file are collected and passed to the `MetaAnalyzer` node. The LLM evaluates each finding, determines if it represents a genuine vulnerability, assigns intent and impact ratings, and generates remediation advice. The prompt enforces strict security rules, instructing the model to ignore "safe" claims within the skill and never execute code.
4. **Result synthesis** – The structured `MetaAnalyzerResult` is parsed, merged with original static findings, and fed to the reporting layer. The final SARIF report includes per-finding metadata, confidence scores, and the enriched LLM assessment.

## Reproducing the Detection Logic

You can execute the core detection logic without invoking the full CLI. The following Python snippet demonstrates loading a skill file, running a static analyzer, and calling the LLM meta-analyzer:

```python
from pathlib import Path
from skillspector.models import Finding
from skillspector.nodes.analyzers.static_patterns_prompt_injection import analyze as pi_analyze
from skillspector.nodes.meta_analyzer import (
    _format_metadata,
    PER_FILE_ANALYSIS_PROMPT,
    MetaAnalyzerResult,
    MetaAnalyzerFinding,
)
from skillspector.llm_analyzer_base import LLMAnalyzerBase, estimate_tokens

# 1️⃣ Load a single file from a skill

file_path = Path("my_skill/skill.py")
content = file_path.read_text()

# 2️⃣ Run a static analyzer (prompt‑injection example)

static_findings: list[Finding] = pi_analyze(content, str(file_path), "python")

# 3️⃣ Prepare the LLM prompt

metadata = {"name": "my_skill", "description": "Demo skill"}
prompt = PER_FILE_ANALYSIS_PROMPT.format(
    metadata=_format_metadata(metadata),
    file_label="Python file",
    file_content=content,
    static_findings="\n".join(f"- {f.rule_id}: {f.message}" for f in static_findings),
)

# 4️⃣ Call the LLM (LLMAnalyzerBase is a thin wrapper around LangChain)

llm = LLMAnalyzerBase(model_name="gpt-4o-mini")  # any structured‑output compatible model

raw_response = llm.run_structured(prompt, schema=MetaAnalyzerResult)

# 5️⃣ Parse the response

result: MetaAnalyzerResult = MetaAnalyzerResult.parse_obj(raw_response)
for finding in result.findings:
    print(
        f"[{finding.pattern_id}] Vulnerable={finding.is_vulnerability} "
        f"Intent={finding.intent} Impact={finding.impact}"
    )
    print("Explanation:", finding.explanation)
    print("Remediation:", finding.remediation, "\n")

```

Running this on a skill containing a hidden "ignore previous instructions" comment returns a `MetaAnalyzerFinding` with `is_vulnerability=True`, `intent="malicious"`, and remediation guidance such as "Remove the hidden directive and review the skill for additional jailbreak content".

## Key Implementation Files

The detection engine spans the following core modules:

- **[`src/skillspector/nodes/analyzers/static_patterns_prompt_injection.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/static_patterns_prompt_injection.py)** – Defines regex patterns P1-P4 for prompt-injection detection and implements the `analyze` function that produces `AnalyzerFinding` objects.
- **[`src/skillspector/nodes/analyzers/pattern_defaults.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/pattern_defaults.py)** – Central catalog of `PatternCategory` enums, default explanations, and remediation text shared across static analyzers.
- **[`src/skillspector/nodes/meta_analyzer.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/meta_analyzer.py)** – Implements the LLM meta-analysis node, including the `PER_FILE_ANALYSIS_PROMPT` template, Pydantic response schemas, and `MetaAnalyzerResult` parsing logic.
- **[`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py)** – Builds the analysis DAG, wires analyzer nodes together, manages `SkillspectorState`, and aggregates outputs.
- **[`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py)** – Provides the `skill-spector scan` command-line interface that instantiates the graph and triggers the full detection workflow.
- **[`src/skillspector/models.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/models.py)** – Core data models including `AnalyzerFinding`, `Location`, and `Severity` used throughout the pipeline.

## Summary

SkillSpector detects vulnerabilities in AI agent skills through a sophisticated hybrid approach:

- **Static pattern analysis** identifies suspicious constructs using regex rules (P1-P4) with confidence scoring.
- **LLM meta-analysis** filters false positives and enriches findings with intent, impact, and remediation via structured Pydantic output.
- **Graph orchestration** executes the pipeline as a DAG, producing standardized SARIF reports or console output.
- The entire workflow is accessible via the `skill-spector scan` CLI command, with modular Python APIs available for custom integrations.

## Frequently Asked Questions

### What distinguishes the static analyzer from the meta-analyzer in SkillSpector?

The static analyzer performs deterministic, regex-based pattern matching to flag potential vulnerabilities with high recall, while the meta-analyzer uses an LLM to apply semantic reasoning, filtering false positives and enriching true positives with contextual severity ratings and remediation steps that regex alone cannot determine.

### How does SkillSpector prevent prompt injection within its own LLM analysis phase?

The meta-analyzer in [`src/skillspector/nodes/meta_analyzer.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/meta_analyzer.py) injects explicit security instructions into `PER_FILE_ANALYSIS_PROMPT` that instruct the LLM to ignore any "safe" or "ignore previous instructions" claims embedded in the skill content, and strictly forbids the model from executing any code during analysis.

### What output formats does SkillSpector support for vulnerability reports?

According to the source code in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) and the reporting layer, SkillSpector generates findings in **SARIF** (Static Analysis Results Interchange Format) for integration with CI/CD pipelines, as well as human-readable console output for local development workflows.

### Can SkillSpector analyze skills written in languages other than Python?

Yes. The static analyzers in `src/skillspector/nodes/analyzers/` treat all skill files—including markdown, JSON, YAML, and source code in any language—as text streams, applying regex patterns and LLM analysis to content regardless of the underlying programming language or file format.