SkillSpectorState Fields: Complete TypedDict Reference for NVIDIA SkillSpector

SkillSpectorState is a TypedDict defining 23 fields that serve as the shared schema for LangGraph workflows, enabling type-safe data passing between input resolution, context building, security analysis, and reporting nodes.

The SkillSpectorState class acts as the immutable contract for state management in the NVIDIA/SkillSpector repository. Defined in [src/skillspector/state.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) (lines 28-75), these SkillSpectorState fields allow analyzer nodes to communicate without side effects while maintaining strict type checking across the security scanning pipeline.

Input Resolution and Path Fields

These six fields handle ingestion, normalization, and temporary resource management during the initial skill loading phase.

  • input_path: str | None — The raw user-supplied path pointing to a file, URL, or zip archive. The resolve_input node processes this to locate the skill definition.
  • skill_path: str | None — Normalized absolute location of the skill after resolution and potential extraction.
  • zip_bytes: bytes | None — Raw binary content of the skill when provided as a zip archive.
  • temp_dir_for_cleanup: str | None — Temporary directory created for extracted archives; the caller must delete this after processing.
  • mode: str — Execution mode discriminator, typically "local" or "remote".
  • yara_rules_dir: str | None — Optional directory path containing additional YARA rules for static analysis, supplied via --yara-rules-dir.

Context Building and Caching Fields

Populated primarily by the build_context node, these seven fields optimize I/O operations and store component metadata for downstream analysis.

  • components: list[str] — List of component identifiers discovered within the skill package.
  • file_cache: dict[str, str] — Mapping of file paths to content strings, preventing redundant disk reads.
  • ast_cache: dict[str, str] — Cached abstract syntax tree representations for source files.
  • manifest: dict[str, object] — Parsed skill.yaml contents describing skill metadata and structure.
  • previous_manifest: dict[str, object] | None — Manifest from prior scans, enabling differential analysis when rerunning checks.
  • component_metadata: list[dict[str, object]] — Extended metadata per component used for risk evaluation and reporting.
  • has_executable_scripts: bool — Security flag indicating whether any component contains executable scripts.

Security Analysis and LLM Configuration Fields

These five fields control the scanning behavior, LLM integration, and collection of security findings.

  • use_llm: bool — Global toggle that enables or disables LLM-based analysis across all nodes.
  • model_config: dict[str, str] — Node-to-model mapping (e.g., {"default": "gpt-4"}) specifying which LLM handles each analysis stage.
  • findings: list[Finding] — Central collection of security issues. This field uses operator.add aggregation, allowing parallel analyzer nodes to safely append results via list concatenation.
  • filtered_findings: list[Finding] — Deduplicated and risk-scored findings that survive post-processing filters.
  • risk_score: int — Numeric risk calculation derived from finding severity and component metadata.

Risk Assessment and Reporting Fields

Generated in final pipeline stages, these five fields produce consumable security reports and remediation guidance.

  • risk_severity: str — Human-readable classification (e.g., "high", "medium", "low") derived from the numeric risk score.
  • risk_recommendation: str — Actionable remediation guidance based on the computed risk profile.
  • output_format: str — Target report format, supporting "markdown" or "sarif" for CI/CD integration.
  • report_body: str — Final rendered report content in the requested format.
  • sarif_report: dict[str, object] — Complete SARIF document structure for integration with security scanning tools.

Working with SkillSpectorState in LangGraph Nodes

Initializing the State

from skillspector.state import SkillspectorState

# Minimal state required to start the graph

state: SkillspectorState = {
    "input_path": "/path/to/skill.yaml",
    "mode": "local",
    "findings": [],  # Typed as List[Finding] via operator.add

    "use_llm": True,
    "output_format": "markdown",
}

Updating Context in build_context

def build_context(state: SkillspectorState) -> SkillspectorState:
    """Populate caches and component lists from source files."""
    state["components"] = ["component_a", "component_b"]
    state["file_cache"] = {"component_a/main.py": "...source..."}
    state["ast_cache"] = {"component_a/main.py": "...AST..."}
    state["manifest"] = {"name": "my-skill", "version": "1.0"}
    return state

Appending Findings from Analyzers

from skillspector.models import Finding

def static_yara_analyzer(state: SkillspectorState) -> SkillspectorState:
    """Append new security findings using operator.add aggregation."""
    new_findings = detect_yara_issues(state["file_cache"])
    # findings field supports += operation via operator.add

    state["findings"] += new_findings
    return state

Generating Final Reports

def render_report(state: SkillspectorState) -> SkillspectorState:
    """Produce SARIF and markdown outputs for CI/CD consumption."""
    sarif = generate_sarif(state["findings"], state["risk_score"])
    state["sarif_report"] = sarif
    state["report_body"] = format_markdown(sarif)
    return state

State Flow Through Key Source Files

Summary

  • SkillSpectorState is defined as a TypedDict in src/skillspector/state.py with 23 strictly-typed fields spanning input handling, caching, analysis, and reporting.
  • The findings field uses operator.add aggregation to safely collect results from parallel analyzer nodes without race conditions.
  • Caching fields (file_cache, ast_cache) prevent redundant I/O during multi-node analysis.
  • State mutations are isolated to specific nodes: resolve_input handles paths, build_context handles metadata, analyzers handle security findings, and report handles output generation.

Frequently Asked Questions

What is SkillSpectorState in NVIDIA SkillSpector?

SkillSpectorState is a TypedDict that defines the shared schema for LangGraph workflows in the SkillSpector security scanner. It enables type-safe communication between nodes by establishing which fields each pipeline stage can read and write.

How does the findings field handle concurrent updates?

The findings field is configured with operator.add aggregation in LangGraph, allowing multiple analyzer nodes running in parallel to append Finding objects simultaneously. The framework automatically concatenates lists from different branches using the operator.add reducer.

What is the difference between findings and filtered_findings?

findings contains all raw security issues discovered by static analyzers and LLM checks, while filtered_findings stores the subset that remains after post-processing steps like deduplication, false-positive removal, and risk score thresholding.

How do I add custom fields to SkillSpectorState?

Modify the SkillSpectorState definition in src/skillspector/nodes/build_context.py and other dependent nodes must be updated to handle the new key, typically by providing sensible defaults in the initial state construction to avoid KeyError exceptions during graph execution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →