How Does SkillSpector Analyze Code? Inside NVIDIA's LangGraph Security Pipeline
SkillSpector analyzes code through a LangGraph-driven pipeline that transforms raw snippets into structured security findings via six sequential stages: input resolution, context building, static pattern matching, LLM-augmented analysis, meta-analysis aggregation, and SARIF reporting.
NVIDIA's SkillSpector is an open-source security scanner that analyzes code for policy violations and security risks. Understanding how SkillSpector analyzes code reveals a sophisticated LangGraph architecture that orchestrates multiple static and AI-driven analyzers into a unified detection pipeline.
The LangGraph Analysis Pipeline Architecture
At the core of SkillSpector's code analysis is a StateGraph compiled in src/skillspector/graph.py. The create_graph() function wires nodes together to process code through distinct stages, passing a shared SkillspectorState object between each transformation.
Stage 1: Entry Point and Input Resolution
Analysis begins when the CLI (src/skillspector/cli.py) or HTTP API creates a SkillspectorState instance containing the raw input and configuration. The resolve_input node (src/skillspector/nodes/resolve_input.py) normalizes file paths, detects programming languages, and handles optional Git checkouts for repository-wide scans.
Stage 2: Context Building
The build_context node (src/skillspector/nodes/build_context.py) gathers surrounding files and extracts code blocks to prepare a context dictionary. This ensures downstream analyzers evaluate security patterns within their proper codebase surroundings rather than in isolation.
Stage 3: Static Analysis Execution
Multiple analyzer nodes execute in parallel after context building. Each analyzer defined in src/skillspector/nodes/analyzers/__init__.py (such as static_patterns_anti_refusal, static_yara, and semantic_security_discovery) implements an analyze function that scans content for known-bad patterns.
The static runner (src/skillspector/nodes/analyzers/static_runner.py) centralizes pattern execution. It iterates over compiled regular expressions or YARA rules, computes confidence scores, applies penalties for matches inside code examples, and deduplicates overlapping findings.
Stage 4: Meta-Analysis and Aggregation
The meta_analyzer node (src/skillspector/nodes/meta_analyzer.py) collects results from all static and LLM-driven analyzers. It resolves conflicts between findings, performs deduplication, and optionally enriches results with LLM-generated explanations via src/skillspector/llm_analyzer_base.py.
Stage 5: Report Generation
The report node (src/skillspector/nodes/report.py) serializes finalized findings into SARIF (Static Analysis Results Interchange Format) or plain-text output using definitions from src/skillspector/sarif_models.py, creating standardized reports ready for CI pipeline integration.
Graph Execution Flow
The following graph definition in src/skillspector/graph.py illustrates how SkillSpector orchestrates the analysis pipeline:
# src/skillspector/graph.py
def create_graph():
workflow = StateGraph(SkillspectorState)
workflow.add_node("resolve_input", resolve_input)
workflow.add_node("build_context", build_context)
workflow.add_node("meta_analyzer", meta_analyzer)
workflow.add_node("report", report)
for analyzer_id in ANALYZER_NODE_IDS:
workflow.add_node(analyzer_id, ANALYZER_NODES[analyzer_id])
workflow.add_edge(START, "resolve_input")
workflow.add_edge("resolve_input", "build_context")
for analyzer_id in ANALYZER_NODE_IDS:
workflow.add_edge("build_context", analyzer_id)
workflow.add_edge(analyzer_id, "meta_analyzer")
workflow.add_edge("meta_analyzer", "report")
workflow.add_edge("report", END)
return workflow.compile()
The graph is compiled once (graph = create_graph()) and invoked for each scan request, ensuring efficient reuse of the workflow structure.
Core Components and Data Structures
Understanding SkillSpector's code analysis requires familiarity with three foundational components:
- SkillspectorState (
src/skillspector/state.py): The shared data model holding original input, derived context, and accumulating findings throughout the pipeline. - AnalyzerFinding: The standardized result object capturing rule IDs, severity levels, confidence scores, and code locations.
- Static Runner (
src/skillspector/nodes/analyzers/static_runner.py): Centralized utility that executes all static pattern modules and returnsAnalyzerFindingobjects.
Static Pattern Matching Implementation
Static analysis in SkillSpector relies on regex and YARA-based detectors. Each analyzer module (such as src/skillspector/nodes/analyzers/static_patterns_anti_refusal.py) defines pattern tuples with associated confidence scores.
When analyzing code, the static runner calculates line numbers, extracts surrounding context, and applies penalty scoring to matches found within documentation examples. This prevents false positives from tutorial code while maintaining detection of actual vulnerabilities.
Running SkillSpector Scans
You can invoke SkillSpector's analysis pipeline through the command line or programmatically via Python.
Command Line Execution
skillspector scan path/to/project
The CLI (src/skillspector/cli.py) initializes a SkillspectorState, invokes the compiled graph, and prints the SARIF report.
Programmatic Execution
from skillspector.graph import graph
from skillspector.state import SkillspectorState
state = SkillspectorState(
input_path="example.py",
config={}, # optional runtime flags
)
results = graph.invoke(state)
print(results["report"]) # contains the final SARIF JSON
Extending Analysis with Custom Patterns
You can extend how SkillSpector analyzes code by adding custom static pattern analyzers. Create a new module using the static_runner utilities:
# my_custom_analyzer.py
from skillspector.models import AnalyzerFinding, Location, Severity
from skillspector.nodes.analyzers import static_runner
ANALYZER_ID = "static_patterns_my_custom"
PATTERNS = [
(r"\bdangerous_function\(\)", 0.9), # simple regex example
]
def analyze(content, file_path, file_type):
findings = []
for match in static_runner.find_matches(content, PATTERNS):
findings.append(
AnalyzerFinding(
rule_id="MY01",
message="Dangerous function call",
severity=Severity.HIGH,
location=Location(file=file_path,
start_line=static_runner.get_line_number(content, match.start())),
confidence=0.9,
tags=["custom"],
context=static_runner.get_context(content, match.start()),
matched_text=match.group(0),
)
)
return findings
def node(state):
return {"findings": static_runner.run_static_patterns(state, [sys.modules[__name__]])}
Register ANALYZER_ID in src/skillspector/nodes/analyzers/__init__.py to automatically insert it into the compiled graph.
Summary
- SkillSpector analyzes code through a LangGraph pipeline with distinct nodes for input resolution, context building, static analysis, meta-analysis, and reporting.
- The static runner (
src/skillspector/nodes/analyzers/static_runner.py) executes regex and YARA patterns with confidence scoring and deduplication. - SkillspectorState (
src/skillspector/state.py) maintains shared state across all analysis stages, from raw input to final findings. - Results output follows the SARIF standard (
src/skillspector/sarif_models.py) for seamless CI/CD integration. - Custom analyzers integrate by implementing the
analyzefunction and registering insrc/skillspector/nodes/analyzers/__init__.py.
Frequently Asked Questions
What analysis engines does SkillSpector use?
SkillSpector employs a hybrid approach combining static pattern matching (regex and YARA rules) with LLM-augmented semantic analysis. The static runner executes compiled patterns from modules like static_patterns_anti_refusal.py, while optional LLM analyzers provide deeper semantic understanding through llm_analyzer_base.py.
How does SkillSpector reduce false positives in code examples?
The static runner applies penalty scoring to matches found within documentation strings or example blocks. When patterns match inside commented sections or docstrings, the confidence score decreases, preventing false positives from tutorial code while maintaining detection of actual vulnerabilities.
Can SkillSpector analyze entire repositories or just single files?
SkillSpector handles both single files and complete repositories. The resolve_input node (src/skillspector/nodes/resolve_input.py) normalizes input paths and performs Git checkouts when necessary, while build_context aggregates surrounding files to provide comprehensive repository-wide analysis context.
What output formats does SkillSpector support?
SkillSpector outputs findings in SARIF (Static Analysis Results Interchange Format) as defined in src/skillspector/sarif_models.py, enabling integration with GitHub Advanced Security, Azure DevOps, and other CI platforms. Plain-text human-readable reports are also available through the report node.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →