How SkillSpector Detects Vulnerabilities in AI Agent Skills: A Three-Stage Pipeline
SkillSpector detects vulnerabilities in AI agent skills by combining deterministic static pattern matching with LLM-driven meta-analysis to filter false positives and enrich findings with actionable remediation guidance.
NVIDIA's SkillSpector is an open-source security scanner designed specifically to detect vulnerabilities in AI agent skills. It employs a hybrid analysis pipeline that blends regex-based static analysis with generative AI validation, ensuring comprehensive coverage while minimizing noise through strict structured output schemas.
The Three-Stage Detection Pipeline
SkillSpector implements a directed acyclic graph (DAG) that orchestrates three tightly coupled stages. Each stage refines the previous output, moving from broad pattern matching to nuanced vulnerability assessment.
Stage 1: Static Pattern Analysis
The first stage scans every file in a skill—code, markdown, prompts, and configuration—using regular-expression-based rules. These rules target known risky constructs such as prompt-injection directives, hidden Unicode tags, insecure tool usage, SSRF URLs, and privilege-escalation calls.
Each rule lives in a dedicated analyzer module under src/skillspector/nodes/analyzers/. For example, prompt-injection detection is handled in src/skillspector/nodes/analyzers/static_patterns_prompt_injection.py, which defines regex patterns P1 through P4 alongside confidence scores. The analyze function walks file content and produces AnalyzerFinding objects containing the rule ID, location, confidence level, and contextual snippet.
Stage 2: LLM-Driven Meta-Analysis
After static analysis, raw findings are passed to a large language model (LLM) that acts as a meta-analyzer. This stage filters false positives and enriches true vulnerabilities with intent classification, impact severity, explanation, and remediation steps.
The implementation resides in src/skillspector/nodes/meta_analyzer.py. This node constructs a structured prompt using PER_FILE_ANALYSIS_PROMPT, injecting skill metadata, file content, and static findings. It calls the LLM through LLMAnalyzerBase with a Pydantic schema MetaAnalyzerResult to enforce type-safe JSON output. The schema guarantees that each returned finding includes boolean is_vulnerability, string fields for intent (malicious/negligent/benign) and impact (critical/high/medium/low), plus concise remediation guidance.
Stage 3: Graph Orchestration and Reporting
The final stage aggregates outputs from all analyzers into a unified report. The orchestration code in src/skillspector/graph.py creates a SkillspectorState, registers each analyzer node (including MetaAnalyzer), and executes await state.run() to traverse the DAG. After execution, the Report node formats consolidated results into SARIF or human-readable console output, accessible via the skill-spector scan <path> command defined in src/skillspector/cli.py.
How the Detection Flow Works in Practice
The complete workflow executes in four concrete steps when scanning a skill directory:
- Skill ingestion – The CLI reads
manifest.jsonorSKILL.mdand walks the directory, loading each file into memory. - Static analysis – For every file, the suite of static analyzers runs in parallel. Each analyzer contributes an
AnalyzerFindingwith a category such asPatternCategory.PROMPT_INJECTIONdefined insrc/skillspector/nodes/analyzers/pattern_defaults.py. - LLM meta-analysis – All findings for a single file are collected and passed to the
MetaAnalyzernode. The LLM evaluates each finding, determines if it represents a genuine vulnerability, assigns intent and impact ratings, and generates remediation advice. The prompt enforces strict security rules, instructing the model to ignore "safe" claims within the skill and never execute code. - Result synthesis – The structured
MetaAnalyzerResultis parsed, merged with original static findings, and fed to the reporting layer. The final SARIF report includes per-finding metadata, confidence scores, and the enriched LLM assessment.
Reproducing the Detection Logic
You can execute the core detection logic without invoking the full CLI. The following Python snippet demonstrates loading a skill file, running a static analyzer, and calling the LLM meta-analyzer:
from pathlib import Path
from skillspector.models import Finding
from skillspector.nodes.analyzers.static_patterns_prompt_injection import analyze as pi_analyze
from skillspector.nodes.meta_analyzer import (
_format_metadata,
PER_FILE_ANALYSIS_PROMPT,
MetaAnalyzerResult,
MetaAnalyzerFinding,
)
from skillspector.llm_analyzer_base import LLMAnalyzerBase, estimate_tokens
# 1️⃣ Load a single file from a skill
file_path = Path("my_skill/skill.py")
content = file_path.read_text()
# 2️⃣ Run a static analyzer (prompt‑injection example)
static_findings: list[Finding] = pi_analyze(content, str(file_path), "python")
# 3️⃣ Prepare the LLM prompt
metadata = {"name": "my_skill", "description": "Demo skill"}
prompt = PER_FILE_ANALYSIS_PROMPT.format(
metadata=_format_metadata(metadata),
file_label="Python file",
file_content=content,
static_findings="\n".join(f"- {f.rule_id}: {f.message}" for f in static_findings),
)
# 4️⃣ Call the LLM (LLMAnalyzerBase is a thin wrapper around LangChain)
llm = LLMAnalyzerBase(model_name="gpt-4o-mini") # any structured‑output compatible model
raw_response = llm.run_structured(prompt, schema=MetaAnalyzerResult)
# 5️⃣ Parse the response
result: MetaAnalyzerResult = MetaAnalyzerResult.parse_obj(raw_response)
for finding in result.findings:
print(
f"[{finding.pattern_id}] Vulnerable={finding.is_vulnerability} "
f"Intent={finding.intent} Impact={finding.impact}"
)
print("Explanation:", finding.explanation)
print("Remediation:", finding.remediation, "\n")
Running this on a skill containing a hidden "ignore previous instructions" comment returns a MetaAnalyzerFinding with is_vulnerability=True, intent="malicious", and remediation guidance such as "Remove the hidden directive and review the skill for additional jailbreak content".
Key Implementation Files
The detection engine spans the following core modules:
src/skillspector/nodes/analyzers/static_patterns_prompt_injection.py– Defines regex patterns P1-P4 for prompt-injection detection and implements theanalyzefunction that producesAnalyzerFindingobjects.src/skillspector/nodes/analyzers/pattern_defaults.py– Central catalog ofPatternCategoryenums, default explanations, and remediation text shared across static analyzers.src/skillspector/nodes/meta_analyzer.py– Implements the LLM meta-analysis node, including thePER_FILE_ANALYSIS_PROMPTtemplate, Pydantic response schemas, andMetaAnalyzerResultparsing logic.src/skillspector/graph.py– Builds the analysis DAG, wires analyzer nodes together, managesSkillspectorState, and aggregates outputs.src/skillspector/cli.py– Provides theskill-spector scancommand-line interface that instantiates the graph and triggers the full detection workflow.src/skillspector/models.py– Core data models includingAnalyzerFinding,Location, andSeverityused throughout the pipeline.
Summary
SkillSpector detects vulnerabilities in AI agent skills through a sophisticated hybrid approach:
- Static pattern analysis identifies suspicious constructs using regex rules (P1-P4) with confidence scoring.
- LLM meta-analysis filters false positives and enriches findings with intent, impact, and remediation via structured Pydantic output.
- Graph orchestration executes the pipeline as a DAG, producing standardized SARIF reports or console output.
- The entire workflow is accessible via the
skill-spector scanCLI command, with modular Python APIs available for custom integrations.
Frequently Asked Questions
What distinguishes the static analyzer from the meta-analyzer in SkillSpector?
The static analyzer performs deterministic, regex-based pattern matching to flag potential vulnerabilities with high recall, while the meta-analyzer uses an LLM to apply semantic reasoning, filtering false positives and enriching true positives with contextual severity ratings and remediation steps that regex alone cannot determine.
How does SkillSpector prevent prompt injection within its own LLM analysis phase?
The meta-analyzer in src/skillspector/nodes/meta_analyzer.py injects explicit security instructions into PER_FILE_ANALYSIS_PROMPT that instruct the LLM to ignore any "safe" or "ignore previous instructions" claims embedded in the skill content, and strictly forbids the model from executing any code during analysis.
What output formats does SkillSpector support for vulnerability reports?
According to the source code in src/skillspector/graph.py and the reporting layer, SkillSpector generates findings in SARIF (Static Analysis Results Interchange Format) for integration with CI/CD pipelines, as well as human-readable console output for local development workflows.
Can SkillSpector analyze skills written in languages other than Python?
Yes. The static analyzers in src/skillspector/nodes/analyzers/ treat all skill files—including markdown, JSON, YAML, and source code in any language—as text streams, applying regex patterns and LLM analysis to content regardless of the underlying programming language or file format.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →