# How SkillSpector's LangGraph Workflow Orchestrates Analyzer Nodes for Security Scanning

> Discover how SkillSpector leverages LangGraph StateGraph to orchestrate over 20 security analyzers in parallel, creating a scalable and dynamic analysis pipeline.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: architecture
- Published: 2026-06-25

---

**SkillSpector uses a LangGraph StateGraph to orchestrate over 20 security analyzers in parallel, dynamically registering nodes from a central registry to create a scalable analysis pipeline.**

NVIDIA's SkillSpector leverages **LangGraph** to coordinate a complex security analysis pipeline that examines AI skills for vulnerabilities. The **SkillSpector's LangGraph workflow orchestrates analyzer nodes** through a dynamic registry system, enabling concurrent execution of static pattern checks, semantic analysis, and policy validation. This architecture allows the tool to scale beyond twenty distinct analyzers without modifying the core graph topology in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py).

## The LangGraph Architecture in SkillSpector

SkillSpector builds its analysis pipeline on **LangGraph**, a lightweight graph-execution engine that manages state transitions between specialized nodes. The workflow defined in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) combines fixed infrastructure nodes with a dynamically expanding collection of analyzer implementations.

### Core Workflow Components

The graph consists of four permanent nodes plus the dynamic analyzer collection:

1. **resolve_input** – Normalizes user input (files, zip archives, URLs, or Git repositories) into a temporary working directory
2. **build_context** – Parses skill manifests, builds AST caches, and extracts component metadata
3. **meta_analyzer** – Performs deduplication, risk scoring, and filtering on aggregated findings
4. **report** – Formats final output as SARIF/JSON and writes to the state

The **analyzer nodes** execute between `build_context` and `meta_analyzer`, running in parallel branches to maximize throughput.

## State Management with SkillspectorState

All nodes share a typed dictionary defined in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) called `SkillspectorState`. This state object enforces type safety and maintains the contract between nodes throughout the pipeline execution.

The state contains:
- **input_path** – Original user-provided path
- **skill_path** – Resolved path to the skill directory
- **findings** – A mutable list that each analyzer appends to using safe list operations
- **components** – Parsed skill component metadata
- **file_cache** – AST and content caches for efficient analysis

Because analyzers only read from shared caches and append to the findings list, they can execute concurrently without race conditions or state corruption.

## Dynamic Node Registration via the Analyzer Registry

All analyzer implementations reside under `src/skillspector/nodes/analyzers/`. Each file exports a `node` object conforming to LangGraph's node signature, while the package-wide registry in [`src/skillspector/nodes/analyzers/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/__init__.py) aggregates these for dynamic loading.

```python

# src/skillspector/nodes/analyzers/__init__.py

ANALYZER_NODE_IDS = [
    "static_patterns_prompt_injection",
    "static_patterns_data_exfiltration",
    # … (20‑plus entries) …

    "semantic_quality_policy",
]

ANALYZER_NODES = {
    "static_patterns_prompt_injection": static_patterns_prompt_injection_node,
    # …

    "semantic_quality_policy": semantic_quality_policy_node,
}

```

This registry pattern decouples analyzer implementations from the graph construction logic. Adding a new analyzer requires only exporting the node and updating these two registry structures, with no changes needed to the core graph definition.

## Graph Construction and Parallel Execution Flow

The `create_graph()` function in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) wires the complete workflow by iterating over the analyzer registry:

```python

# src/skillspector/graph.py

workflow = StateGraph(SkillspectorState)

# Core nodes

workflow.add_node("resolve_input", resolve_input)
workflow.add_node("build_context", build_context)
workflow.add_node("meta_analyzer", meta_analyzer)
workflow.add_node("report", report)

# Dynamically add every analyzer from the registry

for analyzer_id in ANALYZER_NODE_IDS:
    workflow.add_node(analyzer_id, ANALYZER_NODES[analyzer_id])

# Edge layout

workflow.add_edge(START, "resolve_input")
workflow.add_edge("resolve_input", "build_context")

# Parallel branches: each analyzer receives the same context

for analyzer_id in ANALYZER_NODE_IDS:
    workflow.add_edge("build_context", analyzer_id)
    workflow.add_edge(analyzer_id, "meta_analyzer")

workflow.add_edge("meta_analyzer", "report")
workflow.add_edge("report", END)

```

**Parallelism** emerges from the edge topology: because all analyzers connect to the same predecessor (`build_context`) and each has its own outgoing edge to `meta_analyzer`, LangGraph schedules them concurrently (subject to the runtime's thread pool). The `meta_analyzer` acts as a synchronization point, collecting results from all parallel branches before generating the final report.

### Execution Guarantees

The **SkillSpector's LangGraph workflow orchestrates analyzer nodes** with specific guarantees:

- **Deterministic ordering**: `resolve_input` always precedes `build_context`, which always precedes the parallel analyzer phase
- **Graceful degradation**: Individual analyzers check `requires_api_key` or `is_available()` and silently skip if external services are unavailable
- **State-driven flow**: The typed state object prevents accidental mutation of unrelated fields across the 20+ node executions

## Extending the Pipeline: Adding a New Analyzer

Extending the analysis capability requires no modifications to [`graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/graph.py). Implement the node function, export it in the registry, and append the identifier to `ANALYZER_NODE_IDS`:

```python

# 1. Implement node in src/skillspector/nodes/analyzers/my_new_analyzer.py

# 2. Export in __init__.py

from .my_new_analyzer import node as my_new_analyzer_node

ANALYZER_NODE_IDS.append("my_new_analyzer")
ANALYZER_NODES["my_new_analyzer"] = my_new_analyzer_node

```

The next invocation of `create_graph()` automatically incorporates the new analyzer into the parallel execution flow.

## Summary

- **SkillSpector's LangGraph workflow orchestrates analyzer nodes** through a dynamic registry system defined in [`src/skillspector/nodes/analyzers/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/__init__.py), supporting 20+ analyzers without graph topology changes.
- The **StateGraph** in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) uses `SkillspectorState` to share typed data between nodes while enabling parallel execution of independent analyzers.
- **Parallel branches** connect `build_context` to all analyzers simultaneously, with `meta_analyzer` serving as a synchronization point for aggregation and risk scoring.
- **Deterministic execution** guarantees that input resolution and context building complete before analyzers run, while the registry pattern enables simple extension of analysis capabilities.

## Frequently Asked Questions

### How does SkillSpector handle failures in individual analyzer nodes?

Individual analyzers implement availability checks such as `is_available()` or `requires_api_key` validation before execution. If an analyzer encounters an error or lacks required credentials, it returns early without appending to the `findings` list, allowing the graph to continue execution. The **meta_analyzer** node downstream processes only the findings that were successfully generated, ensuring that a single analyzer failure does not crash the entire pipeline.

### Can the analyzer execution order be customized or prioritized?

The current implementation in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) treats all analyzers as parallel siblings connecting from `build_context` to `meta_analyzer`. LangGraph schedules them concurrently based on thread pool availability rather than a fixed sequence. To enforce prioritization, you would need to modify the edge construction logic to create sequential chains or priority groups within the `create_graph()` function, though the default registry pattern treats all analyzers as equal priority.

### What is the performance impact of adding more analyzers to the graph?

Because **SkillSpector's LangGraph workflow orchestrates analyzer nodes** in parallel branches, adding new analyzers increases total execution time only by the duration of the slowest analyzer in the set, not the sum of all durations. The `build_context` node performs heavy lifting (AST parsing, manifest extraction) once upfront, while individual analyzers read from shared caches. Network-bound analyzers (LLM-based semantic checks) may become the bottleneck, but CPU-bound static pattern matchers scale efficiently with the thread pool.

### How does the meta_analyzer node process findings from parallel analyzers?

The **meta_analyzer** node defined in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) receives the accumulated `findings` list from `SkillspectorState` after all parallel analyzer branches complete. It performs deduplication to remove overlapping detections, calculates aggregate risk scores based on severity weights, and applies user-specified filters (such as `--no-llm` flags) to produce `filtered_findings`. This aggregation happens synchronously before the **report** node formats the final SARIF output, ensuring consistent and deduplicated results regardless of how many analyzers contributed to the findings.