# SkillSpectorState Fields: Complete TypedDict Reference for NVIDIA SkillSpector

> Explore the SkillSpectorState TypedDict reference for NVIDIA SkillSpector. Understand its 23 fields, the shared schema for LangGraph workflows, and type-safe data passing.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: api-reference
- Published: 2026-06-23

---

**SkillSpectorState is a TypedDict defining 23 fields that serve as the shared schema for LangGraph workflows, enabling type-safe data passing between input resolution, context building, security analysis, and reporting nodes.**

The `SkillSpectorState` class acts as the immutable contract for state management in the **NVIDIA/SkillSpector** repository. Defined in [[`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py)](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) (lines 28-75), these **SkillSpectorState fields** allow analyzer nodes to communicate without side effects while maintaining strict type checking across the security scanning pipeline.

## Input Resolution and Path Fields

These six fields handle ingestion, normalization, and temporary resource management during the initial skill loading phase.

- **`input_path`**: `str | None` — The raw user-supplied path pointing to a file, URL, or zip archive. The `resolve_input` node processes this to locate the skill definition.
- **`skill_path`**: `str | None` — Normalized absolute location of the skill after resolution and potential extraction.
- **`zip_bytes`**: `bytes | None` — Raw binary content of the skill when provided as a zip archive.
- **`temp_dir_for_cleanup`**: `str | None` — Temporary directory created for extracted archives; the caller must delete this after processing.
- **`mode`**: `str` — Execution mode discriminator, typically `"local"` or `"remote"`.
- **`yara_rules_dir`**: `str | None` — Optional directory path containing additional YARA rules for static analysis, supplied via `--yara-rules-dir`.

## Context Building and Caching Fields

Populated primarily by the `build_context` node, these seven fields optimize I/O operations and store component metadata for downstream analysis.

- **`components`**: `list[str]` — List of component identifiers discovered within the skill package.
- **`file_cache`**: `dict[str, str]` — Mapping of file paths to content strings, preventing redundant disk reads.
- **`ast_cache`**: `dict[str, str]` — Cached abstract syntax tree representations for source files.
- **`manifest`**: `dict[str, object]` — Parsed [`skill.yaml`](https://github.com/NVIDIA/SkillSpector/blob/main/skill.yaml) contents describing skill metadata and structure.
- **`previous_manifest`**: `dict[str, object] | None` — Manifest from prior scans, enabling differential analysis when rerunning checks.
- **`component_metadata`**: `list[dict[str, object]]` — Extended metadata per component used for risk evaluation and reporting.
- **`has_executable_scripts`**: `bool` — Security flag indicating whether any component contains executable scripts.

## Security Analysis and LLM Configuration Fields

These five fields control the scanning behavior, LLM integration, and collection of security findings.

- **`use_llm`**: `bool` — Global toggle that enables or disables LLM-based analysis across all nodes.
- **`model_config`**: `dict[str, str]` — Node-to-model mapping (e.g., `{"default": "gpt-4"}`) specifying which LLM handles each analysis stage.
- **`findings`**: `list[Finding]` — Central collection of security issues. This field uses `operator.add` aggregation, allowing parallel analyzer nodes to safely append results via list concatenation.
- **`filtered_findings`**: `list[Finding]` — Deduplicated and risk-scored findings that survive post-processing filters.
- **`risk_score`**: `int` — Numeric risk calculation derived from finding severity and component metadata.

## Risk Assessment and Reporting Fields

Generated in final pipeline stages, these five fields produce consumable security reports and remediation guidance.

- **`risk_severity`**: `str` — Human-readable classification (e.g., `"high"`, `"medium"`, `"low"`) derived from the numeric risk score.
- **`risk_recommendation`**: `str` — Actionable remediation guidance based on the computed risk profile.
- **`output_format`**: `str` — Target report format, supporting `"markdown"` or `"sarif"` for CI/CD integration.
- **`report_body`**: `str` — Final rendered report content in the requested format.
- **`sarif_report`**: `dict[str, object]` — Complete SARIF document structure for integration with security scanning tools.

## Working with SkillSpectorState in LangGraph Nodes

### Initializing the State

```python
from skillspector.state import SkillspectorState

# Minimal state required to start the graph

state: SkillspectorState = {
    "input_path": "/path/to/skill.yaml",
    "mode": "local",
    "findings": [],  # Typed as List[Finding] via operator.add

    "use_llm": True,
    "output_format": "markdown",
}

```

### Updating Context in build_context

```python
def build_context(state: SkillspectorState) -> SkillspectorState:
    """Populate caches and component lists from source files."""
    state["components"] = ["component_a", "component_b"]
    state["file_cache"] = {"component_a/main.py": "...source..."}
    state["ast_cache"] = {"component_a/main.py": "...AST..."}
    state["manifest"] = {"name": "my-skill", "version": "1.0"}
    return state

```

### Appending Findings from Analyzers

```python
from skillspector.models import Finding

def static_yara_analyzer(state: SkillspectorState) -> SkillspectorState:
    """Append new security findings using operator.add aggregation."""
    new_findings = detect_yara_issues(state["file_cache"])
    # findings field supports += operation via operator.add

    state["findings"] += new_findings
    return state

```

### Generating Final Reports

```python
def render_report(state: SkillspectorState) -> SkillspectorState:
    """Produce SARIF and markdown outputs for CI/CD consumption."""
    sarif = generate_sarif(state["findings"], state["risk_score"])
    state["sarif_report"] = sarif
    state["report_body"] = format_markdown(sarif)
    return state

```

## State Flow Through Key Source Files

- **[[`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py)](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py)**: Contains the `SkillSpectorState` TypedDict declaration (lines 28-75).
- **[[`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py)](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py)**: Constructs the LangGraph workflow and wires nodes to the shared state schema.
- **[[`src/skillspector/nodes/resolve_input.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/resolve_input.py)](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/resolve_input.py)**: Populates `input_path`, `skill_path`, and `temp_dir_for_cleanup`.
- **[[`src/skillspector/nodes/build_context.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/build_context.py)](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/build_context.py)**: Fills `components`, `file_cache`, `ast_cache`, and `manifest` fields.
- **[[`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py)](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py)**: Consumes `output_format` to produce `report_body` and `sarif_report`.

## Summary

- `SkillSpectorState` is defined as a TypedDict in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) with 23 strictly-typed fields spanning input handling, caching, analysis, and reporting.
- The `findings` field uses `operator.add` aggregation to safely collect results from parallel analyzer nodes without race conditions.
- Caching fields (`file_cache`, `ast_cache`) prevent redundant I/O during multi-node analysis.
- State mutations are isolated to specific nodes: `resolve_input` handles paths, `build_context` handles metadata, analyzers handle security findings, and `report` handles output generation.

## Frequently Asked Questions

### What is SkillSpectorState in NVIDIA SkillSpector?

`SkillSpectorState` is a `TypedDict` that defines the shared schema for LangGraph workflows in the SkillSpector security scanner. It enables type-safe communication between nodes by establishing which fields each pipeline stage can read and write.

### How does the findings field handle concurrent updates?

The `findings` field is configured with `operator.add` aggregation in LangGraph, allowing multiple analyzer nodes running in parallel to append `Finding` objects simultaneously. The framework automatically concatenates lists from different branches using the `operator.add` reducer.

### What is the difference between findings and filtered_findings?

`findings` contains all raw security issues discovered by static analyzers and LLM checks, while `filtered_findings` stores the subset that remains after post-processing steps like deduplication, false-positive removal, and risk score thresholding.

### How do I add custom fields to SkillSpectorState?

Modify the `SkillSpectorState` definition in [`src/skillspector/nodes/build_context.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/build_context.py) and other dependent nodes must be updated to handle the new key, typically by providing sensible defaults in the initial state construction to avoid `KeyError` exceptions during graph execution.