SkillSpectorState Fields: Complete TypedDict Reference for NVIDIA SkillSpector
SkillSpectorState is a TypedDict defining 23 fields that serve as the shared schema for LangGraph workflows, enabling type-safe data passing between input resolution, context building, security analysis, and reporting nodes.
The SkillSpectorState class acts as the immutable contract for state management in the NVIDIA/SkillSpector repository. Defined in [src/skillspector/state.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) (lines 28-75), these SkillSpectorState fields allow analyzer nodes to communicate without side effects while maintaining strict type checking across the security scanning pipeline.
Input Resolution and Path Fields
These six fields handle ingestion, normalization, and temporary resource management during the initial skill loading phase.
input_path:str | None— The raw user-supplied path pointing to a file, URL, or zip archive. Theresolve_inputnode processes this to locate the skill definition.skill_path:str | None— Normalized absolute location of the skill after resolution and potential extraction.zip_bytes:bytes | None— Raw binary content of the skill when provided as a zip archive.temp_dir_for_cleanup:str | None— Temporary directory created for extracted archives; the caller must delete this after processing.mode:str— Execution mode discriminator, typically"local"or"remote".yara_rules_dir:str | None— Optional directory path containing additional YARA rules for static analysis, supplied via--yara-rules-dir.
Context Building and Caching Fields
Populated primarily by the build_context node, these seven fields optimize I/O operations and store component metadata for downstream analysis.
components:list[str]— List of component identifiers discovered within the skill package.file_cache:dict[str, str]— Mapping of file paths to content strings, preventing redundant disk reads.ast_cache:dict[str, str]— Cached abstract syntax tree representations for source files.manifest:dict[str, object]— Parsedskill.yamlcontents describing skill metadata and structure.previous_manifest:dict[str, object] | None— Manifest from prior scans, enabling differential analysis when rerunning checks.component_metadata:list[dict[str, object]]— Extended metadata per component used for risk evaluation and reporting.has_executable_scripts:bool— Security flag indicating whether any component contains executable scripts.
Security Analysis and LLM Configuration Fields
These five fields control the scanning behavior, LLM integration, and collection of security findings.
use_llm:bool— Global toggle that enables or disables LLM-based analysis across all nodes.model_config:dict[str, str]— Node-to-model mapping (e.g.,{"default": "gpt-4"}) specifying which LLM handles each analysis stage.findings:list[Finding]— Central collection of security issues. This field usesoperator.addaggregation, allowing parallel analyzer nodes to safely append results via list concatenation.filtered_findings:list[Finding]— Deduplicated and risk-scored findings that survive post-processing filters.risk_score:int— Numeric risk calculation derived from finding severity and component metadata.
Risk Assessment and Reporting Fields
Generated in final pipeline stages, these five fields produce consumable security reports and remediation guidance.
risk_severity:str— Human-readable classification (e.g.,"high","medium","low") derived from the numeric risk score.risk_recommendation:str— Actionable remediation guidance based on the computed risk profile.output_format:str— Target report format, supporting"markdown"or"sarif"for CI/CD integration.report_body:str— Final rendered report content in the requested format.sarif_report:dict[str, object]— Complete SARIF document structure for integration with security scanning tools.
Working with SkillSpectorState in LangGraph Nodes
Initializing the State
from skillspector.state import SkillspectorState
# Minimal state required to start the graph
state: SkillspectorState = {
"input_path": "/path/to/skill.yaml",
"mode": "local",
"findings": [], # Typed as List[Finding] via operator.add
"use_llm": True,
"output_format": "markdown",
}
Updating Context in build_context
def build_context(state: SkillspectorState) -> SkillspectorState:
"""Populate caches and component lists from source files."""
state["components"] = ["component_a", "component_b"]
state["file_cache"] = {"component_a/main.py": "...source..."}
state["ast_cache"] = {"component_a/main.py": "...AST..."}
state["manifest"] = {"name": "my-skill", "version": "1.0"}
return state
Appending Findings from Analyzers
from skillspector.models import Finding
def static_yara_analyzer(state: SkillspectorState) -> SkillspectorState:
"""Append new security findings using operator.add aggregation."""
new_findings = detect_yara_issues(state["file_cache"])
# findings field supports += operation via operator.add
state["findings"] += new_findings
return state
Generating Final Reports
def render_report(state: SkillspectorState) -> SkillspectorState:
"""Produce SARIF and markdown outputs for CI/CD consumption."""
sarif = generate_sarif(state["findings"], state["risk_score"])
state["sarif_report"] = sarif
state["report_body"] = format_markdown(sarif)
return state
State Flow Through Key Source Files
- [
src/skillspector/state.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py): Contains theSkillSpectorStateTypedDict declaration (lines 28-75). - [
src/skillspector/graph.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py): Constructs the LangGraph workflow and wires nodes to the shared state schema. - [
src/skillspector/nodes/resolve_input.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/resolve_input.py): Populatesinput_path,skill_path, andtemp_dir_for_cleanup. - [
src/skillspector/nodes/build_context.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/build_context.py): Fillscomponents,file_cache,ast_cache, andmanifestfields. - [
src/skillspector/nodes/report.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py): Consumesoutput_formatto producereport_bodyandsarif_report.
Summary
SkillSpectorStateis defined as a TypedDict insrc/skillspector/state.pywith 23 strictly-typed fields spanning input handling, caching, analysis, and reporting.- The
findingsfield usesoperator.addaggregation to safely collect results from parallel analyzer nodes without race conditions. - Caching fields (
file_cache,ast_cache) prevent redundant I/O during multi-node analysis. - State mutations are isolated to specific nodes:
resolve_inputhandles paths,build_contexthandles metadata, analyzers handle security findings, andreporthandles output generation.
Frequently Asked Questions
What is SkillSpectorState in NVIDIA SkillSpector?
SkillSpectorState is a TypedDict that defines the shared schema for LangGraph workflows in the SkillSpector security scanner. It enables type-safe communication between nodes by establishing which fields each pipeline stage can read and write.
How does the findings field handle concurrent updates?
The findings field is configured with operator.add aggregation in LangGraph, allowing multiple analyzer nodes running in parallel to append Finding objects simultaneously. The framework automatically concatenates lists from different branches using the operator.add reducer.
What is the difference between findings and filtered_findings?
findings contains all raw security issues discovered by static analyzers and LLM checks, while filtered_findings stores the subset that remains after post-processing steps like deduplication, false-positive removal, and risk score thresholding.
How do I add custom fields to SkillSpectorState?
Modify the SkillSpectorState definition in src/skillspector/nodes/build_context.py and other dependent nodes must be updated to handle the new key, typically by providing sensible defaults in the initial state construction to avoid KeyError exceptions during graph execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →