How SkillSpector's Two-Stage Detection Pipeline Works: Static Analysis + LLM Validation

SkillSpector's two-stage detection pipeline combines high-recall static analyzers with an LLM-powered meta-analyzer to filter false positives and enrich confirmed vulnerabilities with contextual explanations and remediation steps.

NVIDIA's SkillSpector employs a hybrid architecture that balances speed with accuracy when auditing AI skill bundles. The pipeline first applies traditional static analysis to cast a wide net for potential security issues, then leverages a large language model to validate findings and add security context. This design minimizes expensive LLM calls while maintaining precision in vulnerability detection.

Stage 1: Static Analysis for High-Recall Detection

The first stage executes a collection of per-file analyzers that scan source code, configuration files, and other artifacts within the skill bundle. These analyzers are designed for maximum recall, capturing every potential vulnerability even if it generates false positives.

Per-File Analyzers and Finding Generation

In src/skillspector/nodes/analyzers/, the pipeline registers multiple detection modules—including static_yara.py, semantic_developer_intent.py, and semantic_security_discovery.py—within the ANALYZER_NODES dictionary. Each analyzer runs independently against the file cache built during the build_context node execution.

Every detector emits raw findings as structured objects containing rule IDs, file locations, severity levels, and confidence scores. These findings populate the SkillspectorState object defined in state.py, creating a coarse list of potential issues that require validation. The static analysis phase includes pattern matching, behavioral AST analysis, and taint tracking to identify suspicious code patterns without executing the skill.

Stage 2: LLM Meta-Analysis for Precision and Enrichment

The second stage introduces a precision layer through the meta_analyzer node, which processes the raw findings from stage one. This component acts as a filtering and enrichment engine, leveraging LLM reasoning to separate true vulnerabilities from false positives.

The Meta-Analyzer Node

Located in src/skillspector/nodes/meta_analyzer.py, the meta_analyzer function constructs a per-file prompt containing the file metadata, source content, and all static findings for that specific file. It invokes the LLMMetaAnalyzer class—utilizing the base functionality from llm_analyzer_base.py—to perform a single LLM call per file rather than per finding.

The LLM returns a structured MetaAnalyzerResult that indicates which findings represent genuine vulnerabilities. For confirmed issues, the result includes intent classification, impact assessment, confidence scores, and human-readable explanations with remediation steps. This approach keeps token consumption minimal by batching all findings for a given file into one comprehensive LLM request.

LangGraph Workflow Orchestration

The entire two-stage pipeline is wired together as a state graph in src/skillspector/graph.py using LangGraph's StateGraph API. The workflow orchestrates data flow between nodes while maintaining state in a SkillspectorState object.

The graph definition connects the components sequentially:

from langgraph.graph import StateGraph, START, END
from skillspector.state import SkillspectorState

workflow = StateGraph(SkillspectorState)

workflow.add_node("resolve_input", resolve_input)
workflow.add_node("build_context", build_context)
workflow.add_node("meta_analyzer", meta_analyzer)
workflow.add_node("report", report)

# Register all static analyzer nodes

for analyzer_id in ANALYZER_NODE_IDS:
    workflow.add_node(analyzer_id, ANALYZER_NODES[analyzer_id])

# Define edges: resolution → context building → parallel analyzers → meta-analysis → reporting

workflow.add_edge(START, "resolve_input")
workflow.add_edge("resolve_input", "build_context")
for analyzer_id in ANALYZER_NODE_IDS:
    workflow.add_edge("build_context", analyzer_id)
    workflow.add_edge(analyzer_id, "meta_analyzer")
workflow.add_edge("meta_analyzer", "report")
workflow.add_edge("report", END)

The resolve_input node parses command-line arguments or zip bundles, while build_context loads every file into the state cache. After static analysis completes, the report node in src/skillspector/nodes/report.py converts filtered findings into SARIF or JSON formats.

Executing the Pipeline

You can trigger the two-stage pipeline via CLI or programmatically. The command-line interface handles graph compilation internally:


# Run scan with specific LLM model for meta-analysis

skillspector scan ./my_skill.zip --model meta_analyzer=gemini-1.5-pro

For custom integrations, invoke the compiled graph directly:

from skillspector.graph import graph
from skillspector.state import SkillspectorState

state = SkillspectorState()
state["bundle_path"] = "./my_skill.zip"
state["use_llm"] = True
state["model_config"] = {"meta_analyzer": "gemini-1.5-pro"}

# Execute the full two-stage pipeline

final_state = graph.invoke(state)
print(final_state["report"])

Summary

  • SkillSpector's two-stage detection pipeline separates broad static analysis from precision LLM validation to optimize for both speed and accuracy.
  • Static analyzers in src/skillspector/nodes/analyzers/ generate high-recall raw findings using YARA rules, AST analysis, and taint tracking.
  • Meta-analyzer processes findings per-file in meta_analyzer.py, filtering false positives and enriching true vulnerabilities with confidence scores and remediation guidance.
  • LangGraph orchestration in graph.py wires components into a deterministic workflow that minimizes LLM token usage while maximizing security coverage.
  • Output formats include SARIF and JSON, produced by the report node after LLM validation completes.

Frequently Asked Questions

What types of static analyzers does SkillSpector use?

SkillSpector employs multiple static analysis techniques including YARA pattern matching (in static_yara.py), semantic developer intent analysis (in semantic_developer_intent.py), and security discovery heuristics. These analyzers examine source code, configuration files, and bundle manifests to identify potential vulnerabilities without executing the code.

How does the meta-analyzer reduce false positives?

The meta-analyzer in meta_analyzer.py sends static findings to an LLM with the full file context, allowing the model to determine whether flagged patterns represent actual vulnerabilities or benign code. It returns a structured MetaAnalyzerResult that filters out false positives while adding confidence scores and contextual explanations for confirmed issues.

Why does SkillSpector use a per-file rather than per-finding LLM approach?

Processing findings per-file rather than per-individual-finding minimizes API costs and latency while maintaining sufficient context for accurate validation. The llm_analyzer_base.py implements batch handling and token budget logic to optimize these calls, ensuring the LLM receives all findings for a given file in a single prompt.

What output formats does SkillSpector support?

The report node generates security reports in SARIF (Static Analysis Results Interchange Format) and JSON formats. These outputs include the filtered findings from the meta-analyzer, complete with vulnerability explanations, remediation steps, and intent classifications produced during the second stage of the pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →