SkillSpector Architecture Explained: Inside NVIDIA’s LangGraph Security Pipeline
NVIDIA SkillSpector uses a deterministic LangGraph pipeline to orchestrate static security analyzers and optional LLM meta-analysis, processing skills through discrete nodes from input resolution to SARIF reporting.
The SkillSpector architecture centers on a LangGraph-driven workflow that treats security scanning as a deterministic state machine. Developed by NVIDIA, this open-source tool processes skills—whether directories, zip files, URLs, or single SKILL.md files—through a series of pure functions that mutate a shared state object. Understanding this architecture reveals how static pattern matching and semantic LLM analysis combine to produce actionable security findings.
Core Components of the SkillSpector Architecture
Entry Point and CLI Interface
The pipeline begins in src/skillspector/cli.py, which parses command-line arguments and constructs the LangGraph workflow. This module handles user inputs such as --format sarif or --no-llm, validates the configuration, and launches the compiled graph with the initial SkillspectorState.
StateGraph and Shared State Management
At the heart of the architecture lies src/skillspector/graph.py, which declares the StateGraph using the SkillspectorState type. This state object, defined in src/skillspector/state.py as a TypedDict, serves as the immutable shared context throughout execution. It tracks the input_path, file_cache, ast_cache, component metadata, raw findings, filtered findings, and final risk metrics.
Input Normalization and Resolution
The resolve_input node in src/skillspector/nodes/resolve_input.py normalizes diverse input sources. Whether users provide git repositories, zip archives, local directories, or single files, this node extracts and stores the canonical skill_path in the shared state, ensuring downstream analyzers receive a consistent filesystem view.
Context Building and Caching
The build_context node (src/skillspector/nodes/build_context.py) walks the skill tree to populate the state's file_cache and ast_cache. It extracts component metadata, parses abstract syntax trees for Python files, and builds a manifest that records component boundaries and dependencies required by downstream security checks.
Static Analysis Engine
Located in src/skillspector/nodes/analyzers/, this layer contains 20+ pure-Python analyzers that execute in parallel after context building. Each analyzer receives the SkillspectorState and returns an AnalyzerNodeResponse containing Finding objects. These modules detect specific patterns including prompt injection vectors, supply-chain vulnerabilities, YARA rule matches, MCP-specific behaviors, and AST-level anti-patterns. The registry in src/skillspector/nodes/analyzers/__init__.py manages the ANALYZER_NODE_IDS and ANALYZER_NODES mappings that wire these nodes into the graph.
LLM Meta-Analysis Layer
The meta_analyzer node in src/skillspector/nodes/meta_analyzer.py optionally invokes LLM providers (OpenAI, Anthropic, NVIDIA Build) through the abstraction layer in src/skillspector/providers/. This node filters false positives from static analyzers and generates human-readable explanations using common prompt-building logic from src/skillspector/llm_analyzer_base.py. Data structures like Finding, Severity, and Location are defined in src/skillspector/models.py, while SARIF-compatible types live in src/skillspector/sarif_models.py.
Report Generation
The final report node (src/skillspector/nodes/report.py) consumes the state's risk_score, risk_severity, risk_recommendation, and sarif_report fields to generate output. Supported formats include terminal (human-readable), JSON, Markdown, and SARIF for CI/CD integration.
Execution Flow Through the LangGraph Pipeline
The SkillSpector architecture follows a strict six-step execution path defined in src/skillspector/graph.py:
- START → resolve_input: Ingest and normalize the skill source into a local path.
- resolve_input → build_context: Build file caches, AST caches, and component metadata.
- build_context → analyzers: Execute all 20+ static analyzers in parallel, each appending
Findingobjects to the state. - analyzers → meta_analyzer: Run optional LLM semantic analysis on aggregated findings to filter noise and add context.
- meta_analyzer → report: Calculate final risk scores and generate remediation recommendations.
- report → END: Format and write the final output, returning the terminal state to the CLI or API caller.
All nodes are pure functions accepting SkillspectorState and returning partial state updates, ensuring deterministic execution and straightforward unit testing.
Practical Implementation Examples
Programmatic API Usage
You can invoke the compiled graph directly from Python, bypassing the CLI:
from skillspector.graph import graph
result = graph.invoke({
"input_path": "/path/to/skill",
"output_format": "json",
"use_llm": True,
})
print(f"Risk Score: {result['risk_score']}")
print(f"Recommendation: {result['risk_recommendation']}")
for finding in result["filtered_findings"]:
print(f"[{finding.severity}] {finding.id}: {finding.message}")
Command-Line Interface
Typical CLI workflows leverage the same underlying graph:
# Full analysis with LLM semantic checks
skillspector scan ./my-skill/
# CI/CD integration with SARIF output
skillspector scan https://github.com/user/repo.git --format sarif --output report.sarif
# Fast static-only scan without LLM
skillspector scan ./my-skill/ --no-llm
Extending with Custom Analyzers
Create a new analyzer at src/skillspector/nodes/analyzers/static_my_pattern.py:
from skillspector.models import Finding
from skillspector.state import SkillspectorState, AnalyzerNodeResponse
def node(state: SkillspectorState) -> AnalyzerNodeResponse:
findings = []
for path, code in state["file_cache"].items():
if "eval(" in code:
findings.append(Finding(
id="EVAL001",
category="security",
severity="HIGH",
message="Dangerous eval() usage detected",
location={"file": path, "start_line": 1}
))
return {"findings": findings}
Register the analyzer in src/skillspector/nodes/analyzers/__init__.py by adding it to ANALYZER_NODE_IDS and ANALYZER_NODES. The graph automatically includes your analyzer in the parallel execution phase after build_context.
Summary
- SkillSpector architecture relies on a LangGraph
StateGraphto orchestrate security scanning through deterministic, stateless nodes. - The pipeline progresses through input resolution (
resolve_input), context building (build_context), parallel static analysis (analyzers/*), optional LLM meta-analysis (meta_analyzer), and formatted reporting (report). - State management uses a
TypedDictdefined insrc/skillspector/state.py, shared immutably across all nodes via theSkillspectorStateinterface. - Extensibility is built-in: new analyzers integrate by conforming to the
AnalyzerNodeResponseprotocol and registering in the analyzers module. - Output formats include terminal, JSON, Markdown, and SARIF, making the tool suitable for both interactive development and CI/CD pipelines.
Frequently Asked Questions
What makes SkillSpector different from traditional static analysis tools?
Unlike standalone linters, the SkillSpector architecture uses a LangGraph workflow to chain deterministic static pattern matching with LLM semantic analysis. This dual-layer approach in src/skillspector/nodes/meta_analyzer.py reduces false positives while maintaining the execution speed of traditional static checks, providing context-aware security findings.
How does SkillSpector handle different input types?
The resolve_input node in src/skillspector/nodes/resolve_input.py normalizes all inputs—git URLs, zip archives, local directories, or single files—into a canonical skill_path before processing begins. This ensures that downstream analyzers in src/skillspector/nodes/analyzers/ receive a consistent filesystem abstraction regardless of the original source format.
Can I run SkillSpector without LLM access?
Yes. Setting use_llm: False in the Python API or passing --no-llm in the CLI executes only the static analyzers, skipping the meta_analyzer node entirely. This enables fast, offline scans that rely solely on the 20+ pattern-based analyzers in src/skillspector/nodes/analyzers/.
How do I add a new security check to SkillSpector?
Create a new analyzer module in src/skillspector/nodes/analyzers/, implement a node(state: SkillspectorState) function that returns an AnalyzerNodeResponse, and register it in src/skillspector/nodes/analyzers/__init__.py by adding it to the ANALYZER_NODE_IDS and ANALYZER_NODES lists. The graph automatically discovers and executes your analyzer during the parallel analysis phase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →