# How the LangGraph Workflow Orchestrates SkillSpector's Analysis Pipeline

> Discover how the LangGraph workflow stages security analysis in SkillSpector, from input resolution to LLM meta-analysis, optimizing performance with parallel execution and immutable state.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: how-to-guide
- Published: 2026-07-13

---

**SkillSpector uses a LangGraph StateGraph to orchestrate a multi-stage security analysis pipeline, where specialized nodes handle input resolution, static analysis, deduplication, LLM-powered meta-analysis, and final reporting, connected by deterministic edges that enable parallel execution and immutable state passing.**

NVIDIA's SkillSpector leverages the LangGraph library to define its core analysis engine as a deterministic state machine. The **LangGraph workflow** coordinates diverse static analysis tools and optional large language model (LLM) reasoning into a unified pipeline. This architecture, centered in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py), transforms raw skill packages into structured security reports through a series of modular, testable nodes.

## Graph Construction and State Management

The foundation of the pipeline is the `StateGraph` class imported from `langgraph.graph`. In [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py), the `create_graph()` function initializes this workflow engine and defines the strict state schema using Pydantic models from [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py).

The shared state dictionary carries fields like `skill_path`, `use_llm`, `output_format`, and `findings` throughout the execution. Type safety is enforced through models such as `AnalyzerUpdate` and `MetaAnalyzerUpdate`, which validate data as nodes read from and write to the shared state object.

### Node Registration

Individual processing stages are implemented as callable functions or classes within `src/skillspector/nodes/`. The graph registers each component using `graph.add_node("node_name", node_callable)`, mapping logical names to specific implementations like `resolve_input`, `build_context`, `deduplicate`, `meta_analyzer`, and `report`.

## Pipeline Execution Flow and Edge Wiring

The workflow follows a directed acyclic graph structure defined through explicit edge connections. The orchestration uses `START` and `END` constants from LangGraph to mark entry and exit points via `graph.set_entry_point(START)` and `graph.set_finish_point(END)`.

The execution sequence proceeds through these stages:

1. **Input Resolution**: `START → resolve_input` loads skill packages from directories, ZIP files, or URLs, normalizing the input into a consistent internal format.

2. **Context Building**: `resolve_input → build_context` extracts metadata from package manifests and dependency files, preparing the analysis context for downstream stages.

3. **Parallel Analysis**: `build_context → {static_analyzer_1, static_analyzer_2, …}` creates a fan-out pattern where multiple static analyzers (YARA scans, pattern matchers, binary filters) execute simultaneously. Each analyzer appends findings to the `state["findings"]` array.

4. **Deduplication**: The parallel branches converge to `deduplicate`, which collapses duplicate or overlapping findings before further processing.

5. **Meta-Analysis**: `deduplicate → meta_analyzer` optionally invokes an LLM when `use_llm=True` to synthesize higher-level security insights from the aggregated findings.

6. **Reporting**: `meta_analyzer → report` generates SARIF-compatible or JSON output formats, attaching temporary artifacts and formatting the final verdict.

7. **Termination**: `report → END` signals workflow completion.

## Execution Interfaces and Integration Patterns

The compiled graph object supports both synchronous and asynchronous invocation patterns, enabling integration across different runtime environments.

For command-line execution, [`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py) constructs the initial state dictionary from parsed arguments, then calls `graph.invoke(state)` to execute the pipeline synchronously. The function returns the enriched state containing the final report.

For serverless deployment, [`src/skillspector/mcp_server.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/mcp_server.py) wraps the graph in an async interface, calling `await graph.ainvoke(state, config={"run_name": "mcp_scan"})` to process HTTP requests without blocking the event loop.

Batch processing implementations in [`contrib/batch_scan/runner.py`](https://github.com/NVIDIA/SkillSpector/blob/main/contrib/batch_scan/runner.py) demonstrate how the same graph scales to process multiple skills sequentially, reusing the compiled workflow definition across invocations.

### Programmatic Invocation

```python
from skillspector.graph import graph

state = {
    "skill_path": "/path/to/skill",
    "use_llm": True,          # enable LLM-driven meta analysis

    "output_format": "json",  # or "sarif", "text"

}
result = graph.invoke(state)        # sync execution

# `result` contains the final report and the enriched state

print(result["report"])

```

### Command-Line Execution

```bash
skillsppector scan /tmp/my_skill.zip --use-llm --output-format sarif

```

### Asynchronous Server Integration

```python
from skillspector.graph import graph

async def scan_skill(payload: dict):
    result = await graph.ainvoke(payload, config={"run_name": "mcp_scan"})
    return result["verdict"]

```

## Summary

- **SkillSpector's LangGraph workflow** is defined in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) using `StateGraph` from the LangGraph library.
- The pipeline traverses nodes including `resolve_input`, `build_context`, parallel static analyzers, `deduplicate`, `meta_analyzer`, and `report`.
- **Parallel execution** is achieved through fan-out edges from `build_context` to multiple analyzer nodes, with results converging at the deduplication stage.
- State management uses Pydantic models (`AnalyzerUpdate`, `MetaAnalyzerUpdate`) in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) to maintain type safety across node boundaries.
- The graph supports both synchronous (`graph.invoke`) and asynchronous (`graph.ainvoke`) execution patterns, enabling CLI, MCP server, and batch scanner integrations.

## Frequently Asked Questions

### What is the entry point for SkillSpector's LangGraph workflow?

The workflow begins at the `START` node, which LangGraph automatically connects to the `resolve_input` node via `graph.set_entry_point(START)` in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py). This node handles initial loading and validation of the skill package from local paths, ZIP archives, or remote URLs.

### How does SkillSpector handle parallel execution of analyzers?

SkillSpector implements a fan-out pattern where the `build_context` node connects to multiple static analyzer nodes simultaneously using `graph.add_edge`. Each analyzer writes findings to the shared state under `state["findings"]`, and LangGraph handles the parallel execution. The branches converge at the `deduplicate` node, which receives control only after all analyzers complete.

### Can the LangGraph workflow run without LLM capabilities?

Yes. The `meta_analyzer` node checks the `use_llm` flag in the state dictionary. When set to `False`, the node bypasses LLM inference and relies on static rule sets for final analysis. The graph execution remains identical regardless of this flag, as the state schema accommodates both code paths through the `MetaAnalyzerUpdate` model.

### Where is the workflow state defined and validated?

The state structure is defined in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) using Pydantic models. Classes like `AnalyzerUpdate` and `MetaAnalyzerUpdate` enforce type constraints on what each node can write to the shared state, ensuring that `skill_path`, `output_format`, and `findings` maintain consistent types throughout the **LangGraph workflow** execution.