# SkillSpector Architecture Explained: Inside NVIDIA’s LangGraph Security Pipeline

> Explore the SkillSpector architecture, a deterministic LangGraph pipeline by NVIDIA that orchestrates security analyzers and LLM meta-analysis for efficient skill processing from input to SARIF reporting.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: architecture
- Published: 2026-07-11

---

**NVIDIA SkillSpector uses a deterministic LangGraph pipeline to orchestrate static security analyzers and optional LLM meta-analysis, processing skills through discrete nodes from input resolution to SARIF reporting.**

The SkillSpector architecture centers on a **LangGraph-driven workflow** that treats security scanning as a deterministic state machine. Developed by NVIDIA, this open-source tool processes skills—whether directories, zip files, URLs, or single [`SKILL.md`](https://github.com/NVIDIA/SkillSpector/blob/main/SKILL.md) files—through a series of pure functions that mutate a shared state object. Understanding this architecture reveals how static pattern matching and semantic LLM analysis combine to produce actionable security findings.

## Core Components of the SkillSpector Architecture

### Entry Point and CLI Interface

The pipeline begins in [`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py), which parses command-line arguments and constructs the LangGraph workflow. This module handles user inputs such as `--format sarif` or `--no-llm`, validates the configuration, and launches the compiled graph with the initial `SkillspectorState`.

### StateGraph and Shared State Management

At the heart of the architecture lies [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py), which declares the **StateGraph** using the `SkillspectorState` type. This state object, defined in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) as a `TypedDict`, serves as the immutable shared context throughout execution. It tracks the `input_path`, `file_cache`, `ast_cache`, component metadata, raw findings, filtered findings, and final risk metrics.

### Input Normalization and Resolution

The `resolve_input` node in [`src/skillspector/nodes/resolve_input.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/resolve_input.py) normalizes diverse input sources. Whether users provide git repositories, zip archives, local directories, or single files, this node extracts and stores the canonical `skill_path` in the shared state, ensuring downstream analyzers receive a consistent filesystem view.

### Context Building and Caching

The `build_context` node ([`src/skillspector/nodes/build_context.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/build_context.py)) walks the skill tree to populate the state's `file_cache` and `ast_cache`. It extracts component metadata, parses abstract syntax trees for Python files, and builds a manifest that records component boundaries and dependencies required by downstream security checks.

### Static Analysis Engine

Located in `src/skillspector/nodes/analyzers/`, this layer contains 20+ pure-Python analyzers that execute in parallel after context building. Each analyzer receives the `SkillspectorState` and returns an `AnalyzerNodeResponse` containing `Finding` objects. These modules detect specific patterns including prompt injection vectors, supply-chain vulnerabilities, YARA rule matches, MCP-specific behaviors, and AST-level anti-patterns. The registry in [`src/skillspector/nodes/analyzers/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/__init__.py) manages the `ANALYZER_NODE_IDS` and `ANALYZER_NODES` mappings that wire these nodes into the graph.

### LLM Meta-Analysis Layer

The `meta_analyzer` node in [`src/skillspector/nodes/meta_analyzer.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/meta_analyzer.py) optionally invokes LLM providers (OpenAI, Anthropic, NVIDIA Build) through the abstraction layer in `src/skillspector/providers/`. This node filters false positives from static analyzers and generates human-readable explanations using common prompt-building logic from [`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py). Data structures like `Finding`, `Severity`, and `Location` are defined in [`src/skillspector/models.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/models.py), while SARIF-compatible types live in [`src/skillspector/sarif_models.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/sarif_models.py).

### Report Generation

The final `report` node ([`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py)) consumes the state's `risk_score`, `risk_severity`, `risk_recommendation`, and `sarif_report` fields to generate output. Supported formats include terminal (human-readable), JSON, Markdown, and SARIF for CI/CD integration.

## Execution Flow Through the LangGraph Pipeline

The SkillSpector architecture follows a strict six-step execution path defined in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py):

1. **START → resolve_input**: Ingest and normalize the skill source into a local path.
2. **resolve_input → build_context**: Build file caches, AST caches, and component metadata.
3. **build_context → analyzers**: Execute all 20+ static analyzers in parallel, each appending `Finding` objects to the state.
4. **analyzers → meta_analyzer**: Run optional LLM semantic analysis on aggregated findings to filter noise and add context.
5. **meta_analyzer → report**: Calculate final risk scores and generate remediation recommendations.
6. **report → END**: Format and write the final output, returning the terminal state to the CLI or API caller.

All nodes are pure functions accepting `SkillspectorState` and returning partial state updates, ensuring deterministic execution and straightforward unit testing.

## Practical Implementation Examples

### Programmatic API Usage

You can invoke the compiled graph directly from Python, bypassing the CLI:

```python
from skillspector.graph import graph

result = graph.invoke({
    "input_path": "/path/to/skill",
    "output_format": "json",
    "use_llm": True,
})

print(f"Risk Score: {result['risk_score']}")
print(f"Recommendation: {result['risk_recommendation']}")
for finding in result["filtered_findings"]:
    print(f"[{finding.severity}] {finding.id}: {finding.message}")

```

### Command-Line Interface

Typical CLI workflows leverage the same underlying graph:

```bash

# Full analysis with LLM semantic checks

skillspector scan ./my-skill/

# CI/CD integration with SARIF output

skillspector scan https://github.com/user/repo.git --format sarif --output report.sarif

# Fast static-only scan without LLM

skillspector scan ./my-skill/ --no-llm

```

### Extending with Custom Analyzers

Create a new analyzer at [`src/skillspector/nodes/analyzers/static_my_pattern.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/static_my_pattern.py):

```python
from skillspector.models import Finding
from skillspector.state import SkillspectorState, AnalyzerNodeResponse

def node(state: SkillspectorState) -> AnalyzerNodeResponse:
    findings = []
    for path, code in state["file_cache"].items():
        if "eval(" in code:
            findings.append(Finding(
                id="EVAL001",
                category="security",
                severity="HIGH",
                message="Dangerous eval() usage detected",
                location={"file": path, "start_line": 1}
            ))
    return {"findings": findings}

```

Register the analyzer in [`src/skillspector/nodes/analyzers/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/__init__.py) by adding it to `ANALYZER_NODE_IDS` and `ANALYZER_NODES`. The graph automatically includes your analyzer in the parallel execution phase after `build_context`.

## Summary

- **SkillSpector architecture** relies on a LangGraph `StateGraph` to orchestrate security scanning through deterministic, stateless nodes.
- The pipeline progresses through input resolution (`resolve_input`), context building (`build_context`), parallel static analysis (`analyzers/*`), optional LLM meta-analysis (`meta_analyzer`), and formatted reporting (`report`).
- State management uses a `TypedDict` defined in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py), shared immutably across all nodes via the `SkillspectorState` interface.
- Extensibility is built-in: new analyzers integrate by conforming to the `AnalyzerNodeResponse` protocol and registering in the analyzers module.
- Output formats include terminal, JSON, Markdown, and SARIF, making the tool suitable for both interactive development and CI/CD pipelines.

## Frequently Asked Questions

### What makes SkillSpector different from traditional static analysis tools?

Unlike standalone linters, the SkillSpector architecture uses a LangGraph workflow to chain deterministic static pattern matching with LLM semantic analysis. This dual-layer approach in [`src/skillspector/nodes/meta_analyzer.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/meta_analyzer.py) reduces false positives while maintaining the execution speed of traditional static checks, providing context-aware security findings.

### How does SkillSpector handle different input types?

The `resolve_input` node in [`src/skillspector/nodes/resolve_input.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/resolve_input.py) normalizes all inputs—git URLs, zip archives, local directories, or single files—into a canonical `skill_path` before processing begins. This ensures that downstream analyzers in `src/skillspector/nodes/analyzers/` receive a consistent filesystem abstraction regardless of the original source format.

### Can I run SkillSpector without LLM access?

Yes. Setting `use_llm: False` in the Python API or passing `--no-llm` in the CLI executes only the static analyzers, skipping the `meta_analyzer` node entirely. This enables fast, offline scans that rely solely on the 20+ pattern-based analyzers in `src/skillspector/nodes/analyzers/`.

### How do I add a new security check to SkillSpector?

Create a new analyzer module in `src/skillspector/nodes/analyzers/`, implement a `node(state: SkillspectorState)` function that returns an `AnalyzerNodeResponse`, and register it in [`src/skillspector/nodes/analyzers/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/__init__.py) by adding it to the `ANALYZER_NODE_IDS` and `ANALYZER_NODES` lists. The graph automatically discovers and executes your analyzer during the parallel analysis phase.