# How NVIDIA SkillSpector Implements Security Scanning with LangGraph: A Workflow Deep Dive

> Discover how NVIDIA SkillSpector uses LangGraph to orchestrate its security scanning workflow. Explore the mutable state and parallel execution of analyzers within this deep dive.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: deep-dive
- Published: 2026-07-11

---

**SkillSpector orchestrates its entire security scanning pipeline through LangGraph, where a mutable `SkillspectorState` object flows sequentially through specialized nodes while analyzers execute in parallel fan-out phases.**

The NVIDIA SkillSpector repository demonstrates production-grade application of LangGraph for automated security analysis. By treating each scanning stage as a node in a directed graph, the SkillSpector workflow LangGraph implementation eliminates manual orchestration code while enabling concurrent execution of static and LLM-backed analyzers. This architecture separates input handling, context building, analysis, and reporting into discrete, composable units defined in specific source files.

## Core Architecture Components

The workflow rests on four foundational pieces that define how data moves through the system.

### SkillspectorState: The Shared Data Contract

At [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py), the `SkillspectorState` typed dictionary serves as the immutable contract shared across all graph nodes. This state container holds the `input_path`, `file_cache` (a mapping of paths to contents), accumulated `findings`, LLM telemetry logs, and the final `sarif_report`.

Unlike traditional global state, this structure leverages LangGraph's reducer patterns (specifically `operator.add` for list fields) to safely merge partial state updates when parallel analyzers return results simultaneously.

### Graph Definition and Node Registry

The graph construction happens in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py), where the `create_graph()` function instantiates a `StateGraph` using `SkillspectorState` as its generic type. The system discovers available analyzers through two registry constants defined in [`src/skillspector/nodes/analyzers/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/__init__.py):

- **`ANALYZER_NODE_IDS`**: A list of string identifiers for every registered analyzer
- **`ANALYZER_NODES`**: A dictionary mapping those IDs to their callable implementations

This registry pattern allows the graph to dynamically wire edges from the context-building phase to each analyzer, then converge results into the meta-analyzer without hardcoding specific node names.

## Execution Flow: From Input to Report

The SkillSpector workflow LangGraph pipeline follows a linear "pre-process → fan-out → reduce → report" pattern with five distinct phases.

### Input Resolution and Context Building

The workflow begins with `resolve_input` ([`src/skillspector/nodes/resolve_input.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/resolve_input.py)), which normalizes user input—whether a local directory, remote URL, or zip file—into a local `skill_path` stored in state. This node handles the unpacking and temporary directory management automatically.

Next, `build_context` ([`src/skillspector/nodes/build_context.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/build_context.py)) traverses the skill directory to populate the `file_cache` and extract metadata from [`SKILL.md`](https://github.com/NVIDIA/SkillSpector/blob/main/SKILL.md) manifests. This node prepares the immutable context that all subsequent analyzers will reference.

### Parallel Analyzer Fan-Out

LangGraph's scheduling capability shines in the third phase. For each ID in `ANALYZER_NODE_IDS`, the graph creates an edge from `build_context` to that analyzer node, then another edge from the analyzer to `meta_analyzer`.

Because these analyzers operate on the same immutable context data stored in `SkillspectorState`, LangGraph executes them concurrently. Each analyzer appends its security findings to `state["findings"]` using the reducer pattern, preventing race conditions during parallel writes.

### Meta-Analysis and Reporting

The `meta_analyzer` node ([`src/skillspector/nodes/meta_analyzer.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/meta_analyzer.py)) receives the aggregated findings list and performs deduplication, false-positive filtering, and policy-level analysis. It stores the cleaned results in `state["filtered_findings"]`.

Finally, the `report` node ([`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py)) consumes the complete state to generate both a human-readable report (stored in `state["report_body"]`) and a machine-readable SARIF JSON document suitable for CI/CD integration.

## Extending the Pipeline

The modular design makes adding analyzers straightforward. Each analyzer is a pure function accepting `SkillspectorState` and returning an `AnalyzerNodeResponse` (a partial state update with new findings).

To register a new analyzer:

```python

# src/skillspector/nodes/analyzers/__init__.py

from .my_analyzer import my_analyzer

ANALYZER_NODE_IDS.append("my_analyzer")
ANALYZER_NODES["my_analyzer"] = my_analyzer

```

The next time `create_graph()` runs, LangGraph automatically includes the new node in the parallel execution fan-out without modifying [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py).

## Implementation Examples

### Running via CLI

The command-line interface ([`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py)) handles state initialization and graph invocation:

```bash

# Scan a local skill directory

skillspector scan /path/to/skill

# Scan a remote repository (URL resolved automatically)

skillspector scan https://github.com/example/agent-skill.git

```

Under the hood, the CLI constructs the initial `SkillspectorState` and calls `graph.invoke(state)`.

### Programmatic Execution

```python
from skillspector.graph import graph
from skillspector.state import SkillspectorState

# Initialize state

initial_state: SkillspectorState = {
    "input_path": "https://github.com/example/agent-skill.git",
    "mode": "scan",
    "use_llm": True,
    "show_suppressed": False,
    "model_config": {},
    "findings": [],
    "llm_call_log": [],
}

# Execute workflow

final_state = graph.invoke(initial_state)

# Access results

print(final_state["report_body"])

```

### Custom Analyzer Implementation

```python
from skillspector.state import SkillspectorState, AnalyzerNodeResponse

def env_file_scanner(state: SkillspectorState) -> AnalyzerNodeResponse:
    findings = []
    for path, content in state["file_cache"].items():
        if path.endswith(".env"):
            findings.append({
                "id": "ENV-001",
                "title": "Environment file detected",
                "severity": "high",
                "file_path": path,
            })
    return {"findings": findings}

```

## Summary

- **LangGraph orchestration** powers the entire SkillSpector pipeline, handling node scheduling and state management automatically.
- **`SkillspectorState`** acts as the typed, shared data contract that flows through all analysis phases, with reducer patterns ensuring safe parallel updates.
- **Parallel execution** of analyzers occurs naturally through LangGraph's fan-out edges from `build_context` to each registered analyzer node.
- **Dynamic registration** via `ANALYZER_NODE_IDS` and `ANALYZER_NODES` allows extension without core graph modifications.
- **Dual reporting** generates both human-readable text and SARIF JSON for CI integration in the final `report` node.

## Frequently Asked Questions

### What makes LangGraph suitable for SkillSpector's security scanning workflow?

LangGraph provides declarative wiring of directed-acyclic graphs with built-in support for parallel node execution and immutable state management. According to the NVIDIA SkillSpector source code, this allows the system to run multiple static and LLM-backed analyzers concurrently without manual thread management, while the reducer pattern safely merges findings from parallel branches into the shared `SkillspectorState`.

### How does SkillSpector handle parallel execution of analyzers?

The graph definition in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) creates independent edges from the `build_context` node to each analyzer listed in `ANALYZER_NODE_IDS`, then converges them to `meta_analyzer`. Because LangGraph recognizes these analyzer nodes have no interdependencies, it schedules them concurrently. Each analyzer appends to `state["findings"]` using the `operator.add` reducer, ensuring thread-safe aggregation without explicit synchronization code.

### Can I add custom analyzers to SkillSpector without modifying the core graph logic?

Yes. Create a node function that accepts `SkillspectorState` and returns `AnalyzerNodeResponse`, then register it in [`src/skillspector/nodes/analyzers/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/__init__.py) by appending its ID to `ANALYZER_NODE_IDS` and adding the callable to `ANALYZER_NODES`. The `create_graph()` function dynamically wires new analyzers into the parallel execution phase during graph compilation, requiring no changes to [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py).

### What output formats does SkillSpector generate?

The `report` node ([`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py)) produces two artifacts: a human-readable string stored in `state["report_body"]` containing formatted findings and risk assessments, and a machine-readable SARIF JSON document stored in `state["sarif_report"]` for integration with continuous integration pipelines and security dashboards.