Understanding the SkillSpector Two-Stage Analysis Pipeline: Static Pattern Matching and LLM Semantic Evaluation
SkillSpector analyzes agent skills by first running deterministic static pattern matching to flag potential issues, then using an LLM to semantically validate and enrich those findings while ensuring critical alerts are never silently dropped.
NVIDIA's SkillSpector implements a hybrid security analysis approach that combines the speed of deterministic scanning with the reasoning capabilities of large language models. This SkillSpector two-stage analysis pipeline first examines skill files using lightweight rule-based scanners, then invokes an LLM to perform semantic validation and filter false positives. The architecture ensures fast feedback during development while maintaining rigorous security standards through a "fail-closed" design that preserves high-severity findings even when the LLM disagrees.
Stage One: Static Pattern Matching
Before any LLM credentials are consulted, SkillSpector executes a collection of light-weight, rule-based scanners including YARA rules, regular-expression matchers, and custom pattern definitions. Each scanner inspects raw skill files and emits a list of Finding objects that serve as the deterministic foundation for the entire analysis.
How Static Scanners Identify Potential Issues
The static phase is deliberately deterministic and fast, operating entirely without LLM dependencies. According to the test suite in tests/nodes/analyzers/test_static_patterns.py, scanners apply patterns to source files and construct Finding objects containing:
rule_id(e.g.,E2,P1)- Precise location (
file,start_line,end_line) - Severity and confidence scores
- Matched text snippets
- Optional explanation and remediation drawn from the default pattern library
This stage functions as a broad-spectrum filter, identifying syntactic patterns that might indicate security anti-patterns, hardcoded credentials, or dangerous API usage.
The Finding Data Model
The Finding class defined in src/skillspector/models.py serves as the universal data model shared between both pipeline stages. Static analyzers populate these objects using rule definitions from src/skillspector/nodes/analyzers/pattern_defaults.py, which provides default security explanations when the LLM has not yet enriched the result.
Stage Two: LLM-Based Semantic Evaluation
After the static stage produces a per-file list of findings, the graph invokes the meta_analyzer node to perform semantic assessment. This stage leverages the LLMMetaAnalyzer class—a subclass of LLMAnalyzerBase defined in src/skillspector/llm_analyzer_base.py—to reason about the context and intent behind each flagged pattern.
The Meta Analyzer Node
The core orchestration lives in src/skillspector/nodes/meta_analyzer.py within the meta_analyzer(state) function (lines 497–525). This node:
- Retrieves static
findingsfrom the graph state - Builds per-file prompts using
PER_FILE_ANALYSIS_PROMPT - Invokes the LLM via
LLMMetaAnalyzerwith batching support for large files - Gathers structured
MetaAnalyzerResultobjects
The LLM is called once per file (or in batches if token limits require splitting), returning a structured Pydantic schema containing semantic verdicts for each finding.
Prompt Engineering and Structured Output
The PER_FILE_ANALYSIS_PROMPT template submitted to the LLM includes:
- Skill metadata (name, description, triggers, permissions)
- Raw file content (
{file_content}) - Static findings formatted by
_format_findings_for_prompt - Explicit instructions forbidding the LLM to ignore "safe-skill" messages or dismiss findings without analysis
The LLM returns a MetaAnalyzerResult containing for each finding:
is_vulnerability(boolean)confidence(0–1 float)intent(malicious / negligent / benign)impact(critical / high / medium / low)- Enriched
explanationandremediation
Filtering Logic and Safety Invariants
The apply_filter method merges LLM verdicts back into the static findings while enforcing strict safety invariants:
- Confirmed vulnerabilities: Findings with
confidence >= 0.6are retained and enriched with LLM-generated explanations and remediations. - High-severity protection: Static findings marked
CRITICALorHIGHare never dropped. If the LLM fails to confirm them, they are retained with the special tag"llm-unconfirmed"to preserve safety. - False-positive suppression: Lower-severity findings are kept only when the LLM explicitly validates them; otherwise, they are filtered out as likely false positives.
This "fail-closed" behavior ensures that the deterministic static stage acts as a safety net that the LLM cannot override.
Pipeline Orchestration in Code
You can invoke the complete pipeline from the command line:
# Analyze a skill manifest using the full two-stage pipeline
$ skillspector analyze my_skill.yaml
To run only the static stage (useful for CI/CD speed):
from skillspector.nodes.analyzers import static_patterns
findings = static_patterns.run_on_file("scripts/greet.py")
print([f.rule_id for f in findings])
The LLM filtering step can be instantiated programmatically:
from skillspector.nodes.meta_analyzer import LLMMetaAnalyzer, PER_FILE_ANALYSIS_PROMPT
llm = LLMMetaAnalyzer(model="gpt-4o-mini")
prompt = llm.build_prompt(
batch=llm.Batch(
file_label="greet.py",
content=open("scripts/greet.py").read(),
findings=findings,
),
metadata_text="Name: GreetSkill\nDescription: Simple greeting skill"
)
response = llm.call(prompt) # Returns structured MetaAnalyzerResult
enriched = llm.apply_filter(findings, [(batch, response)])
Summary
- Stage One uses deterministic YARA rules and regex patterns to generate
Findingobjects quickly without LLM dependencies. - Stage Two invokes the
meta_analyzernode to semantically validate findings, enriching them with intent classification and impact assessment via structured LLM output. - Safety invariants prevent high-severity static findings from being dropped, tagging them
"llm-unconfirmed"rather than removing them when the LLM disagrees. - Key files include
src/skillspector/nodes/meta_analyzer.pyfor orchestration,tests/nodes/analyzers/test_static_patterns.pyfor pattern logic, andsrc/skillspector/models.pyfor the shared data model.
Frequently Asked Questions
What makes the SkillSpector two-stage pipeline different from pure LLM analysis?
SkillSpector combines deterministic static analysis with LLM reasoning rather than relying solely on probabilistic models. According to the NVIDIA/SkillSpector source code, the static stage provides guaranteed coverage of known dangerous patterns, while the LLM stage adds semantic understanding to reduce false positives. This hybrid approach ensures the tool remains functional without API credentials while still benefiting from LLM intelligence when available.
How does SkillSpector ensure critical security findings are not missed?
The apply_filter logic in src/skillspector/nodes/meta_analyzer.py enforces a strict safety invariant: any static finding with CRITICAL or HIGH severity is never removed from the output. If the LLM fails to confirm such a finding or returns low confidence, SkillSpector retains the alert with an "llm-unconfirmed" tag rather than filtering it out, ensuring potentially dangerous code is always flagged for review.
What happens if the LLM returns a low confidence score for a high-severity static finding?
When the LLM returns a confidence score below 0.6 for a high-severity finding, SkillSpector preserves the original static alert and appends the metadata tag "llm-unconfirmed". The finding remains in the final report with its original severity and location data, but without the enriched explanation or remediation that the LLM typically provides. This prevents the LLM from inadvertently suppressing serious security issues due to context misunderstanding.
Can SkillSpector run without LLM credentials?
Yes. The static pattern-matching stage is completely independent of LLM services and functions deterministically using only local rules defined in src/skillspector/nodes/analyzers/pattern_defaults.py. Users can run static analysis in air-gapped environments or CI pipelines without API keys, though they will receive unfiltered findings without the semantic validation and false-positive reduction that the LLM stage provides.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →