# How SkillSpector's Two-Stage Static + LLM Analysis Pipeline Works

> Discover how SkillSpector's two-stage analysis pipeline combines static analyzers and LLMs to refine agent skills, filter false positives, and deliver secure structured assessments.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: internals
- Published: 2026-07-09

---

**SkillSpector processes agent skills through a LangGraph workflow that first runs deterministic static analyzers to generate raw findings, then employs a LLM meta-analyzer to filter false positives, enrich vulnerability context, and produce structured assessments with fail-closed safety guarantees.**

NVIDIA's SkillSpector implements a defense-in-depth approach to skill security scanning by combining fast static pattern matching with contextual large language model reasoning. This two-stage analysis pipeline ensures deterministic detection of known vulnerability signatures while leveraging LLMs to reduce noise and provide human-readable explanations. The entire workflow is orchestrated through a directed graph defined in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) that wires together context building, static analysis, and intelligent filtering nodes.

## Stage One: Static Analysis and Context Building

The pipeline begins with the **`build_context`** node located in [`src/skillspector/nodes/build_context.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/build_context.py). This node reads the skill manifest and materializes all skill files into a **file cache**, creating an expanded representation that includes surrounding code context for each file.

Once the context is built, the workflow dispatches files to independent static analyzer nodes. These analyzers include:

- **[`static_yara.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_yara.py)** – YARA signature matching for known malicious patterns
- **[`static_patterns_privilege_escalation.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_patterns_privilege_escalation.py)** – Pattern-based detection of privilege escalation vectors
- **[`osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/osv_client.py)** – Open Source Vulnerability database client for dependency scanning

Each analyzer returns a list of **`Finding`** objects defined in [`src/skillspector/models.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/models.py). These Pydantic models encapsulate:
- `rule_id` – The identifier of the triggered rule
- `message` – Human-readable description of the issue
- `severity` – Risk level (INFO, LOW, MEDIUM, HIGH, CRITICAL)
- `confidence` – Certainty score of the match
- `line_numbers` and `matched_text` – Precise location data
- Optional `remediation` – Suggested fixes when available

## Stage Two: LLM Filtering and Enrichment

After all static analyzers complete, the **`meta_analyzer`** node executes from [`src/skillspector/nodes/meta_analyzer.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/meta_analyzer.py). This stage introduces contextual reasoning through the **`LLMMetaAnalyzer`** class, which extends `LLMAnalyzerBase` from [`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py).

For every file containing at least one static finding, the meta-analyzer constructs a per-file LLM request using the **`PER_FILE_ANALYSIS_PROMPT`** template. The prompt injects:
- Skill metadata (name, description, intended purpose)
- Full file contents with syntax highlighting
- Raw static findings from stage one
- Security-focused instructions that explicitly forbid the model from trusting self-declared "safe" statements

The LLM response must conform to the **`MetaAnalyzerResult`** Pydantic schema, ensuring structured JSON output containing:
- **`findings`** – Enriched entries with boolean `is_vulnerability`, recalibrated `confidence`, `intent` classification, `impact` assessment, detailed `explanation`, and specific `remediation` steps
- **`overall_assessment`** – A file-level risk summary categorizing the aggregate threat level

The **`apply_filter`** routine then merges the LLM response with the original static findings using a safety-gated logic: it **preserves every HIGH or CRITICAL static finding** unconditionally, adding an `"llm-unconfirmed"` tag when the LLM disagrees with the severity. Lower-severity findings are discarded or downgraded based on the LLM verdict, significantly reducing false positives without risking silent suppression of serious vulnerabilities.

## Workflow Orchestration in LangGraph

The complete pipeline is assembled in **[`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py)** using LangGraph's state machine architecture. The workflow follows this directed acyclic graph:

1. **`resolve_input`** – Parses CLI arguments and manifests
2. **`build_context`** – Creates the file cache and analyzer inputs  
3. **Static analyzer nodes** – Parallel execution of all analyzers listed in `ANALYZER_NODE_IDS`
4. **`meta_analyzer`** – Sequential LLM processing of accumulated findings
5. **`report`** – Serialization to SARIF or JSON via [`src/skillspector/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/report.py)

The **`SkillspectorState`** typed dictionary (defined in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py)) carries immutable state through the graph, accumulating `Finding` objects at each stage until the final filtered results are produced.

## Fail-Closed Safety Mechanisms

SkillSpector implements **fail-closed semantics** to ensure pipeline reliability. If LLM calls are disabled via `use_llm=False` or if API calls fail, the system activates fallback routines:

- **`_fallback_filtered`** – Applies confidence-based heuristics to static findings and attaches default remediations when LLM enrichment is unavailable
- **`_passthrough_with_defaults`** – Ensures critical findings are never dropped by supplying conservative default values for missing LLM assessments

The credential resolution and model instantiation are handled by **[`src/skillspector/llm_utils.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_utils.py)**, which manages API keys and constructs LangChain `ChatModel` objects compatible with various provider endpoints.

## Running the Pipeline

**Command-line execution** (default configuration with LLM):

```bash
skillspector scan path/to/skill \
    --model meta_analyzer=gpt-4o-mini \
    --output results.sarif

```

**Programmatic invocation** using the compiled LangGraph:

```python
from skillspector.graph import graph
from skillspector.state import SkillspectorState

# Initialize state with manifest and configuration

state = SkillspectorState(
    manifest={"name": "demo-skill", "description": "example"},
    file_cache={},
    use_llm=True,
    model_config={"meta_analyzer": "gpt-4o-mini"},
)

# Execute the complete workflow

final_state = graph.invoke(state)

# Access filtered findings

for finding in final_state["filtered_findings"]:
    print(f"{finding.rule_id}: {finding.message} (confidence={finding.confidence:.2f})")

```

**CI/CD execution without LLM dependencies**:

```bash
skillspector scan path/to/skill --no-llm

```

Setting `--no-llm` triggers the heuristic fallback path, ensuring security scanning continues uninterrupted in environments without API access.

## Summary

- **Two-stage architecture** combines deterministic static analysis in `build_context` and analyzer nodes with contextual LLM filtering in `meta_analyzer`
- **Static phase** uses YARA signatures, pattern matchers, and OSV clients to generate structured `Finding` objects with precise line numbers and severity ratings
- **LLM phase** employs the `LLMMetaAnalyzer` class with the `PER_FILE_ANALYSIS_PROMPT` template to validate findings against the `MetaAnalyzerResult` schema, enriching true positives with impact assessments while filtering false positives
- **Fail-closed guarantees** ensure HIGH and CRITICAL static findings are always preserved regardless of LLM availability, with `_fallback_filtered` providing sensible defaults during outages
- **LangGraph orchestration** in [`graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/graph.py) manages the workflow through `SkillspectorState`, enabling both CLI usage via `skillspector scan` and programmatic integration through `graph.invoke()`

## Frequently Asked Questions

### What happens if the LLM service is unavailable or rate-limited?

The pipeline activates failure-resistant fallback mechanisms. The `meta_analyzer` node detects the failure and invokes **`_fallback_filtered`** or **`_passthrough_with_defaults`**, which apply confidence heuristics to retain HIGH and CRITICAL findings while adding default remediations. This ensures the security scan completes without silently dropping alerts.

### How does SkillSpector prevent the LLM from dismissing real vulnerabilities?

The system implements a safety-gated filtering policy in the `apply_filter` method. **Every HIGH or CRITICAL static finding is preserved unconditionally**, even if the LLM disagrees with the assessment. The LLM only influences the filtering of MEDIUM and lower severity findings, and conservative tagging (`"llm-unconfirmed"`) marks discrepancies for human review.

### Can I use SkillSpector without an LLM for cost-sensitive environments?

Yes. Running the scan with the **`--no-llm`** flag disables the `meta_analyzer` LLM calls entirely. The pipeline executes only the static analysis phase and applies the heuristic fallback logic to generate findings with rule-based remediations, eliminating API costs while maintaining deterministic security coverage.

### What is the structure of the MetaAnalyzerResult schema that the LLM must return?

The **`MetaAnalyzerResult`** Pydantic model requires two top-level fields: a list of `findings` containing enriched vulnerability data (`is_vulnerability`, `confidence`, `intent`, `impact`, `explanation`, `remediation`), and an `overall_assessment` string summarizing the file's aggregate risk level. This structured output ensures consistent downstream processing regardless of the underlying LLM provider configured in `model_config`.