How the LangGraph Workflow Orchestrates SkillSpector's Analysis Pipeline
SkillSpector uses a LangGraph StateGraph to orchestrate a multi-stage security analysis pipeline, where specialized nodes handle input resolution, static analysis, deduplication, LLM-powered meta-analysis, and final reporting, connected by deterministic edges that enable parallel execution and immutable state passing.
NVIDIA's SkillSpector leverages the LangGraph library to define its core analysis engine as a deterministic state machine. The LangGraph workflow coordinates diverse static analysis tools and optional large language model (LLM) reasoning into a unified pipeline. This architecture, centered in src/skillspector/graph.py, transforms raw skill packages into structured security reports through a series of modular, testable nodes.
Graph Construction and State Management
The foundation of the pipeline is the StateGraph class imported from langgraph.graph. In src/skillspector/graph.py, the create_graph() function initializes this workflow engine and defines the strict state schema using Pydantic models from src/skillspector/state.py.
The shared state dictionary carries fields like skill_path, use_llm, output_format, and findings throughout the execution. Type safety is enforced through models such as AnalyzerUpdate and MetaAnalyzerUpdate, which validate data as nodes read from and write to the shared state object.
Node Registration
Individual processing stages are implemented as callable functions or classes within src/skillspector/nodes/. The graph registers each component using graph.add_node("node_name", node_callable), mapping logical names to specific implementations like resolve_input, build_context, deduplicate, meta_analyzer, and report.
Pipeline Execution Flow and Edge Wiring
The workflow follows a directed acyclic graph structure defined through explicit edge connections. The orchestration uses START and END constants from LangGraph to mark entry and exit points via graph.set_entry_point(START) and graph.set_finish_point(END).
The execution sequence proceeds through these stages:
-
Input Resolution:
START → resolve_inputloads skill packages from directories, ZIP files, or URLs, normalizing the input into a consistent internal format. -
Context Building:
resolve_input → build_contextextracts metadata from package manifests and dependency files, preparing the analysis context for downstream stages. -
Parallel Analysis:
build_context → {static_analyzer_1, static_analyzer_2, …}creates a fan-out pattern where multiple static analyzers (YARA scans, pattern matchers, binary filters) execute simultaneously. Each analyzer appends findings to thestate["findings"]array. -
Deduplication: The parallel branches converge to
deduplicate, which collapses duplicate or overlapping findings before further processing. -
Meta-Analysis:
deduplicate → meta_analyzeroptionally invokes an LLM whenuse_llm=Trueto synthesize higher-level security insights from the aggregated findings. -
Reporting:
meta_analyzer → reportgenerates SARIF-compatible or JSON output formats, attaching temporary artifacts and formatting the final verdict. -
Termination:
report → ENDsignals workflow completion.
Execution Interfaces and Integration Patterns
The compiled graph object supports both synchronous and asynchronous invocation patterns, enabling integration across different runtime environments.
For command-line execution, src/skillspector/cli.py constructs the initial state dictionary from parsed arguments, then calls graph.invoke(state) to execute the pipeline synchronously. The function returns the enriched state containing the final report.
For serverless deployment, src/skillspector/mcp_server.py wraps the graph in an async interface, calling await graph.ainvoke(state, config={"run_name": "mcp_scan"}) to process HTTP requests without blocking the event loop.
Batch processing implementations in contrib/batch_scan/runner.py demonstrate how the same graph scales to process multiple skills sequentially, reusing the compiled workflow definition across invocations.
Programmatic Invocation
from skillspector.graph import graph
state = {
"skill_path": "/path/to/skill",
"use_llm": True, # enable LLM-driven meta analysis
"output_format": "json", # or "sarif", "text"
}
result = graph.invoke(state) # sync execution
# `result` contains the final report and the enriched state
print(result["report"])
Command-Line Execution
skillsppector scan /tmp/my_skill.zip --use-llm --output-format sarif
Asynchronous Server Integration
from skillspector.graph import graph
async def scan_skill(payload: dict):
result = await graph.ainvoke(payload, config={"run_name": "mcp_scan"})
return result["verdict"]
Summary
- SkillSpector's LangGraph workflow is defined in
src/skillspector/graph.pyusingStateGraphfrom the LangGraph library. - The pipeline traverses nodes including
resolve_input,build_context, parallel static analyzers,deduplicate,meta_analyzer, andreport. - Parallel execution is achieved through fan-out edges from
build_contextto multiple analyzer nodes, with results converging at the deduplication stage. - State management uses Pydantic models (
AnalyzerUpdate,MetaAnalyzerUpdate) insrc/skillspector/state.pyto maintain type safety across node boundaries. - The graph supports both synchronous (
graph.invoke) and asynchronous (graph.ainvoke) execution patterns, enabling CLI, MCP server, and batch scanner integrations.
Frequently Asked Questions
What is the entry point for SkillSpector's LangGraph workflow?
The workflow begins at the START node, which LangGraph automatically connects to the resolve_input node via graph.set_entry_point(START) in src/skillspector/graph.py. This node handles initial loading and validation of the skill package from local paths, ZIP archives, or remote URLs.
How does SkillSpector handle parallel execution of analyzers?
SkillSpector implements a fan-out pattern where the build_context node connects to multiple static analyzer nodes simultaneously using graph.add_edge. Each analyzer writes findings to the shared state under state["findings"], and LangGraph handles the parallel execution. The branches converge at the deduplicate node, which receives control only after all analyzers complete.
Can the LangGraph workflow run without LLM capabilities?
Yes. The meta_analyzer node checks the use_llm flag in the state dictionary. When set to False, the node bypasses LLM inference and relies on static rule sets for final analysis. The graph execution remains identical regardless of this flag, as the state schema accommodates both code paths through the MetaAnalyzerUpdate model.
Where is the workflow state defined and validated?
The state structure is defined in src/skillspector/state.py using Pydantic models. Classes like AnalyzerUpdate and MetaAnalyzerUpdate enforce type constraints on what each node can write to the shared state, ensuring that skill_path, output_format, and findings maintain consistent types throughout the LangGraph workflow execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →