How SkillSpector's Two-Stage Static + LLM Analysis Pipeline Works
SkillSpector processes agent skills through a LangGraph workflow that first runs deterministic static analyzers to generate raw findings, then employs a LLM meta-analyzer to filter false positives, enrich vulnerability context, and produce structured assessments with fail-closed safety guarantees.
NVIDIA's SkillSpector implements a defense-in-depth approach to skill security scanning by combining fast static pattern matching with contextual large language model reasoning. This two-stage analysis pipeline ensures deterministic detection of known vulnerability signatures while leveraging LLMs to reduce noise and provide human-readable explanations. The entire workflow is orchestrated through a directed graph defined in src/skillspector/graph.py that wires together context building, static analysis, and intelligent filtering nodes.
Stage One: Static Analysis and Context Building
The pipeline begins with the build_context node located in src/skillspector/nodes/build_context.py. This node reads the skill manifest and materializes all skill files into a file cache, creating an expanded representation that includes surrounding code context for each file.
Once the context is built, the workflow dispatches files to independent static analyzer nodes. These analyzers include:
static_yara.py– YARA signature matching for known malicious patternsstatic_patterns_privilege_escalation.py– Pattern-based detection of privilege escalation vectorsosv_client.py– Open Source Vulnerability database client for dependency scanning
Each analyzer returns a list of Finding objects defined in src/skillspector/models.py. These Pydantic models encapsulate:
rule_id– The identifier of the triggered rulemessage– Human-readable description of the issueseverity– Risk level (INFO, LOW, MEDIUM, HIGH, CRITICAL)confidence– Certainty score of the matchline_numbersandmatched_text– Precise location data- Optional
remediation– Suggested fixes when available
Stage Two: LLM Filtering and Enrichment
After all static analyzers complete, the meta_analyzer node executes from src/skillspector/nodes/meta_analyzer.py. This stage introduces contextual reasoning through the LLMMetaAnalyzer class, which extends LLMAnalyzerBase from src/skillspector/llm_analyzer_base.py.
For every file containing at least one static finding, the meta-analyzer constructs a per-file LLM request using the PER_FILE_ANALYSIS_PROMPT template. The prompt injects:
- Skill metadata (name, description, intended purpose)
- Full file contents with syntax highlighting
- Raw static findings from stage one
- Security-focused instructions that explicitly forbid the model from trusting self-declared "safe" statements
The LLM response must conform to the MetaAnalyzerResult Pydantic schema, ensuring structured JSON output containing:
findings– Enriched entries with booleanis_vulnerability, recalibratedconfidence,intentclassification,impactassessment, detailedexplanation, and specificremediationstepsoverall_assessment– A file-level risk summary categorizing the aggregate threat level
The apply_filter routine then merges the LLM response with the original static findings using a safety-gated logic: it preserves every HIGH or CRITICAL static finding unconditionally, adding an "llm-unconfirmed" tag when the LLM disagrees with the severity. Lower-severity findings are discarded or downgraded based on the LLM verdict, significantly reducing false positives without risking silent suppression of serious vulnerabilities.
Workflow Orchestration in LangGraph
The complete pipeline is assembled in src/skillspector/graph.py using LangGraph's state machine architecture. The workflow follows this directed acyclic graph:
resolve_input– Parses CLI arguments and manifestsbuild_context– Creates the file cache and analyzer inputs- Static analyzer nodes – Parallel execution of all analyzers listed in
ANALYZER_NODE_IDS meta_analyzer– Sequential LLM processing of accumulated findingsreport– Serialization to SARIF or JSON viasrc/skillspector/report.py
The SkillspectorState typed dictionary (defined in src/skillspector/state.py) carries immutable state through the graph, accumulating Finding objects at each stage until the final filtered results are produced.
Fail-Closed Safety Mechanisms
SkillSpector implements fail-closed semantics to ensure pipeline reliability. If LLM calls are disabled via use_llm=False or if API calls fail, the system activates fallback routines:
_fallback_filtered– Applies confidence-based heuristics to static findings and attaches default remediations when LLM enrichment is unavailable_passthrough_with_defaults– Ensures critical findings are never dropped by supplying conservative default values for missing LLM assessments
The credential resolution and model instantiation are handled by src/skillspector/llm_utils.py, which manages API keys and constructs LangChain ChatModel objects compatible with various provider endpoints.
Running the Pipeline
Command-line execution (default configuration with LLM):
skillspector scan path/to/skill \
--model meta_analyzer=gpt-4o-mini \
--output results.sarif
Programmatic invocation using the compiled LangGraph:
from skillspector.graph import graph
from skillspector.state import SkillspectorState
# Initialize state with manifest and configuration
state = SkillspectorState(
manifest={"name": "demo-skill", "description": "example"},
file_cache={},
use_llm=True,
model_config={"meta_analyzer": "gpt-4o-mini"},
)
# Execute the complete workflow
final_state = graph.invoke(state)
# Access filtered findings
for finding in final_state["filtered_findings"]:
print(f"{finding.rule_id}: {finding.message} (confidence={finding.confidence:.2f})")
CI/CD execution without LLM dependencies:
skillspector scan path/to/skill --no-llm
Setting --no-llm triggers the heuristic fallback path, ensuring security scanning continues uninterrupted in environments without API access.
Summary
- Two-stage architecture combines deterministic static analysis in
build_contextand analyzer nodes with contextual LLM filtering inmeta_analyzer - Static phase uses YARA signatures, pattern matchers, and OSV clients to generate structured
Findingobjects with precise line numbers and severity ratings - LLM phase employs the
LLMMetaAnalyzerclass with thePER_FILE_ANALYSIS_PROMPTtemplate to validate findings against theMetaAnalyzerResultschema, enriching true positives with impact assessments while filtering false positives - Fail-closed guarantees ensure HIGH and CRITICAL static findings are always preserved regardless of LLM availability, with
_fallback_filteredproviding sensible defaults during outages - LangGraph orchestration in
graph.pymanages the workflow throughSkillspectorState, enabling both CLI usage viaskillspector scanand programmatic integration throughgraph.invoke()
Frequently Asked Questions
What happens if the LLM service is unavailable or rate-limited?
The pipeline activates failure-resistant fallback mechanisms. The meta_analyzer node detects the failure and invokes _fallback_filtered or _passthrough_with_defaults, which apply confidence heuristics to retain HIGH and CRITICAL findings while adding default remediations. This ensures the security scan completes without silently dropping alerts.
How does SkillSpector prevent the LLM from dismissing real vulnerabilities?
The system implements a safety-gated filtering policy in the apply_filter method. Every HIGH or CRITICAL static finding is preserved unconditionally, even if the LLM disagrees with the assessment. The LLM only influences the filtering of MEDIUM and lower severity findings, and conservative tagging ("llm-unconfirmed") marks discrepancies for human review.
Can I use SkillSpector without an LLM for cost-sensitive environments?
Yes. Running the scan with the --no-llm flag disables the meta_analyzer LLM calls entirely. The pipeline executes only the static analysis phase and applies the heuristic fallback logic to generate findings with rule-based remediations, eliminating API costs while maintaining deterministic security coverage.
What is the structure of the MetaAnalyzerResult schema that the LLM must return?
The MetaAnalyzerResult Pydantic model requires two top-level fields: a list of findings containing enriched vulnerability data (is_vulnerability, confidence, intent, impact, explanation, remediation), and an overall_assessment string summarizing the file's aggregate risk level. This structured output ensures consistent downstream processing regardless of the underlying LLM provider configured in model_config.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →