How SkillSpector's LangGraph Workflow Orchestrates Analyzer Nodes for Security Scanning

SkillSpector uses a LangGraph StateGraph to orchestrate over 20 security analyzers in parallel, dynamically registering nodes from a central registry to create a scalable analysis pipeline.

NVIDIA's SkillSpector leverages LangGraph to coordinate a complex security analysis pipeline that examines AI skills for vulnerabilities. The SkillSpector's LangGraph workflow orchestrates analyzer nodes through a dynamic registry system, enabling concurrent execution of static pattern checks, semantic analysis, and policy validation. This architecture allows the tool to scale beyond twenty distinct analyzers without modifying the core graph topology in src/skillspector/graph.py.

The LangGraph Architecture in SkillSpector

SkillSpector builds its analysis pipeline on LangGraph, a lightweight graph-execution engine that manages state transitions between specialized nodes. The workflow defined in src/skillspector/graph.py combines fixed infrastructure nodes with a dynamically expanding collection of analyzer implementations.

Core Workflow Components

The graph consists of four permanent nodes plus the dynamic analyzer collection:

  1. resolve_input – Normalizes user input (files, zip archives, URLs, or Git repositories) into a temporary working directory
  2. build_context – Parses skill manifests, builds AST caches, and extracts component metadata
  3. meta_analyzer – Performs deduplication, risk scoring, and filtering on aggregated findings
  4. report – Formats final output as SARIF/JSON and writes to the state

The analyzer nodes execute between build_context and meta_analyzer, running in parallel branches to maximize throughput.

State Management with SkillspectorState

All nodes share a typed dictionary defined in src/skillspector/state.py called SkillspectorState. This state object enforces type safety and maintains the contract between nodes throughout the pipeline execution.

The state contains:

  • input_path – Original user-provided path
  • skill_path – Resolved path to the skill directory
  • findings – A mutable list that each analyzer appends to using safe list operations
  • components – Parsed skill component metadata
  • file_cache – AST and content caches for efficient analysis

Because analyzers only read from shared caches and append to the findings list, they can execute concurrently without race conditions or state corruption.

Dynamic Node Registration via the Analyzer Registry

All analyzer implementations reside under src/skillspector/nodes/analyzers/. Each file exports a node object conforming to LangGraph's node signature, while the package-wide registry in src/skillspector/nodes/analyzers/__init__.py aggregates these for dynamic loading.


# src/skillspector/nodes/analyzers/__init__.py

ANALYZER_NODE_IDS = [
    "static_patterns_prompt_injection",
    "static_patterns_data_exfiltration",
    # … (20‑plus entries) …

    "semantic_quality_policy",
]

ANALYZER_NODES = {
    "static_patterns_prompt_injection": static_patterns_prompt_injection_node,
    # …

    "semantic_quality_policy": semantic_quality_policy_node,
}

This registry pattern decouples analyzer implementations from the graph construction logic. Adding a new analyzer requires only exporting the node and updating these two registry structures, with no changes needed to the core graph definition.

Graph Construction and Parallel Execution Flow

The create_graph() function in src/skillspector/graph.py wires the complete workflow by iterating over the analyzer registry:


# src/skillspector/graph.py

workflow = StateGraph(SkillspectorState)

# Core nodes

workflow.add_node("resolve_input", resolve_input)
workflow.add_node("build_context", build_context)
workflow.add_node("meta_analyzer", meta_analyzer)
workflow.add_node("report", report)

# Dynamically add every analyzer from the registry

for analyzer_id in ANALYZER_NODE_IDS:
    workflow.add_node(analyzer_id, ANALYZER_NODES[analyzer_id])

# Edge layout

workflow.add_edge(START, "resolve_input")
workflow.add_edge("resolve_input", "build_context")

# Parallel branches: each analyzer receives the same context

for analyzer_id in ANALYZER_NODE_IDS:
    workflow.add_edge("build_context", analyzer_id)
    workflow.add_edge(analyzer_id, "meta_analyzer")

workflow.add_edge("meta_analyzer", "report")
workflow.add_edge("report", END)

Parallelism emerges from the edge topology: because all analyzers connect to the same predecessor (build_context) and each has its own outgoing edge to meta_analyzer, LangGraph schedules them concurrently (subject to the runtime's thread pool). The meta_analyzer acts as a synchronization point, collecting results from all parallel branches before generating the final report.

Execution Guarantees

The SkillSpector's LangGraph workflow orchestrates analyzer nodes with specific guarantees:

  • Deterministic ordering: resolve_input always precedes build_context, which always precedes the parallel analyzer phase
  • Graceful degradation: Individual analyzers check requires_api_key or is_available() and silently skip if external services are unavailable
  • State-driven flow: The typed state object prevents accidental mutation of unrelated fields across the 20+ node executions

Extending the Pipeline: Adding a New Analyzer

Extending the analysis capability requires no modifications to graph.py. Implement the node function, export it in the registry, and append the identifier to ANALYZER_NODE_IDS:


# 1. Implement node in src/skillspector/nodes/analyzers/my_new_analyzer.py

# 2. Export in __init__.py

from .my_new_analyzer import node as my_new_analyzer_node

ANALYZER_NODE_IDS.append("my_new_analyzer")
ANALYZER_NODES["my_new_analyzer"] = my_new_analyzer_node

The next invocation of create_graph() automatically incorporates the new analyzer into the parallel execution flow.

Summary

  • SkillSpector's LangGraph workflow orchestrates analyzer nodes through a dynamic registry system defined in src/skillspector/nodes/analyzers/__init__.py, supporting 20+ analyzers without graph topology changes.
  • The StateGraph in src/skillspector/graph.py uses SkillspectorState to share typed data between nodes while enabling parallel execution of independent analyzers.
  • Parallel branches connect build_context to all analyzers simultaneously, with meta_analyzer serving as a synchronization point for aggregation and risk scoring.
  • Deterministic execution guarantees that input resolution and context building complete before analyzers run, while the registry pattern enables simple extension of analysis capabilities.

Frequently Asked Questions

How does SkillSpector handle failures in individual analyzer nodes?

Individual analyzers implement availability checks such as is_available() or requires_api_key validation before execution. If an analyzer encounters an error or lacks required credentials, it returns early without appending to the findings list, allowing the graph to continue execution. The meta_analyzer node downstream processes only the findings that were successfully generated, ensuring that a single analyzer failure does not crash the entire pipeline.

Can the analyzer execution order be customized or prioritized?

The current implementation in src/skillspector/graph.py treats all analyzers as parallel siblings connecting from build_context to meta_analyzer. LangGraph schedules them concurrently based on thread pool availability rather than a fixed sequence. To enforce prioritization, you would need to modify the edge construction logic to create sequential chains or priority groups within the create_graph() function, though the default registry pattern treats all analyzers as equal priority.

What is the performance impact of adding more analyzers to the graph?

Because SkillSpector's LangGraph workflow orchestrates analyzer nodes in parallel branches, adding new analyzers increases total execution time only by the duration of the slowest analyzer in the set, not the sum of all durations. The build_context node performs heavy lifting (AST parsing, manifest extraction) once upfront, while individual analyzers read from shared caches. Network-bound analyzers (LLM-based semantic checks) may become the bottleneck, but CPU-bound static pattern matchers scale efficiently with the thread pool.

How does the meta_analyzer node process findings from parallel analyzers?

The meta_analyzer node defined in src/skillspector/graph.py receives the accumulated findings list from SkillspectorState after all parallel analyzer branches complete. It performs deduplication to remove overlapping detections, calculates aggregate risk scores based on severity weights, and applies user-specified filters (such as --no-llm flags) to produce filtered_findings. This aggregation happens synchronously before the report node formats the final SARIF output, ensuring consistent and deduplicated results regardless of how many analyzers contributed to the findings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →