# How to Contribute to SkillSpector Development: A Complete Guide

> Contribute to SkillSpector development by cloning the NVIDIA repository, installing dev dependencies, and adding new analyzers to the LangGraph security pipeline. Follow our guide to contribute now.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: how-to-guide
- Published: 2026-07-10

---

**To contribute to SkillSpector development, clone the NVIDIA repository, install dependencies with `make install-dev`, and extend the LangGraph security pipeline by adding new analyzers in `src/skillspector/nodes/analyzers/` and registering them in the `ANALYZER_NODES` registry.**

SkillSpector is NVIDIA's open-source **LangGraph-based security scanner** for AI-agent skills that uses a directed acyclic graph to process code through 22 parallel analyzers. Contributing to SkillSpector development requires understanding its state-driven architecture defined in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) and its pure functional node pattern. This guide provides the exact file paths, code patterns, and commands you need to add features, run tests, and submit pull requests.

## Understand the LangGraph Architecture

Before you contribute to SkillSpector development, familiarize yourself with its core data flow. The pipeline processes a **`SkillspectorState`** TypedDict (defined in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py)) through five high-level stages:

1. **Input resolution** – The `resolve_input` node in [`src/skillspector/input_handler.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/input_handler.py) converts URLs, Git repos, or ZIP files into local directories.
2. **Context building** – The `build_context` node walks the directory, caches file contents, builds ASTs, and extracts component metadata.
3. **Parallel analysis** – 22 analyzer nodes (static regex, YARA, AST-based detectors, and optional LLM checks) emit `Finding` objects defined in [`src/skillspector/models.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/models.py).
4. **Meta-analysis** – The `meta_analyzer` node optionally filters findings using `LLMMetaAnalyzer`.
5. **Reporting** – The `report` node applies baseline suppression, computes risk scores, and renders SARIF/JSON/Markdown output.

All state mutations use LangGraph’s reducer pattern, making the pipeline side-effect free and easy to unit test.

## Set Up Your Development Environment

Start by cloning the repository and creating an isolated Python environment. The project prefers **`uv`** for fast dependency resolution, though standard `venv` works.

```bash
git clone https://github.com/NVIDIA/SkillSpector.git
cd SkillSpector

# Create virtual environment (uv preferred)

uv venv .venv && source .venv/bin/activate

# Install development dependencies

make install-dev

```

The `Makefile` targets assume your virtual environment is active. See [`docs/DEVELOPMENT.md`](https://github.com/NVIDIA/SkillSpector/blob/main/docs/DEVELOPMENT.md) for platform-specific setup details.

## Run the Workflow Locally

Validate your environment by scanning a sample skill. You can invoke SkillSpector via CLI or launch the interactive LangGraph Studio.

**CLI usage:**

```bash

# Default scan (static + LLM analysis)

skillspector scan ./examples/my-skill/

# Fast iteration without LLM calls

skillspector scan ./examples/my-skill/ --no-llm

# Generate structured JSON output

skillspector scan ./examples/my-skill/ --format json -o report.json

```

**Interactive development server:**

```bash
make langgraph-dev

```

This command starts a local HTTP server and opens **LangGraph Studio**, where you can step through each node, inspect the `SkillspectorState` at every stage, and debug analyzer logic visually.

## Add a New Analyzer

Extending the detection engine is the most common way to contribute to SkillSpector development. The analyzer registry in [`src/skillspector/nodes/analyzers/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/__init__.py) dynamically loads analysis modules.

### Create the Analyzer Module

Create a new Python file under `src/skillspector/nodes/analyzers/`. A minimal static analyzer implements an `analyze()` function that returns a list of `AnalyzerFinding` objects:

```python

# src/skillspector/nodes/analyzers/static_patterns_secret_leak.py

from .common import AnalyzerFinding, Location, Severity
from .pattern_defaults import CATEGORY, EXPLANATION, REMEDIATION

def analyze(content: str, file_path: str, file_type: str) -> list[AnalyzerFinding]:
    findings = []
    # Detect hard-coded AWS access key patterns

    if "AKIA" in content:
        line_no = next(
            i for i, line in enumerate(content.splitlines(), 1) if "AKIA" in line
        )
        findings.append(
            AnalyzerFinding(
                location=Location(file=file_path, line=line_no),
                severity=Severity.HIGH,
                rule_id="SEC1",
                message="Hard-coded AWS access key detected",
                category=CATEGORY,
                explanation=EXPLANATION,
                remediation=REMEDIATION,
            )
        )
    return findings

```

### Register in the Analyzer Registry

Import your module and append to the registry lists in [`src/skillspector/nodes/analyzers/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/__init__.py):

```python
from . import static_patterns_secret_leak

ANALYZER_NODE_IDS.append("static_secret_leak")
ANALYZER_NODES["static_secret_leak"] = static_runner.run_static_patterns(
    static_patterns_secret_leak.analyze
)

```

The `ANALYZER_NODE_IDS` list controls execution order, while `ANALYZER_NODES` maps IDs to the runnable node functions.

### Write Unit Tests

Create a test file under `tests/nodes/analyzers/` to ensure your analyzer correctly identifies issues without false positives:

```python

# tests/nodes/analyzers/test_static_secret_leak.py

from skillspector.nodes.analyzers.static_patterns_secret_leak import analyze

def test_detects_aws_key():
    content = "aws_access_key_id = AKIAIOSFODNN7EXAMPLE"
    findings = analyze(content, "config.py", "python")
    assert len(findings) == 1
    assert findings[0].rule_id == "SEC1"
    assert findings[0].severity.value == "high"

```

## Test and Validate Your Changes

Before submitting, run the full quality assurance pipeline to ensure no regressions:

```bash

# Run all unit tests

make test

# Check code style and typing

make lint

# Auto-format code

make format

```

All three commands must pass to satisfy CI requirements. The repository uses a **risk-score gating** model, so reviewers will verify that new rules correctly categorize findings and do not inflate false positive rates.

## Submit Your Contribution

Once your analyzer passes local validation, submit a pull request following the repository guidelines in [`CONTRIBUTING.md`](https://github.com/NVIDIA/SkillSpector/blob/main/CONTRIBUTING.md):

1. Fork the repository and create a feature branch (`git checkout -b feature/my-analyzer`).
2. Update [`docs/DEVELOPMENT.md`](https://github.com/NVIDIA/SkillSpector/blob/main/docs/DEVELOPMENT.md) if your analyzer introduces new configuration flags or behavioral changes.
3. Ensure all CI checks pass (`make lint && make test && make format`).
4. Open a PR against `main` with a clear description linking to any related issues.
5. Respond to reviewer feedback regarding rule severity classification and performance impact.

## Summary

- **SkillSpector** uses a LangGraph DAG defined in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) to process security scans through 22 parallel analyzers.
- Set up your environment with `make install-dev` and use `make langgraph-dev` to visualize the execution graph.
- Add analyzers by creating modules in `src/skillspector/nodes/analyzers/` and registering them in [`__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/__init__.py) via `ANALYZER_NODES`.
- Write tests in `tests/nodes/analyzers/` and validate with `make test` before submitting.
- Follow the risk-score gating model and update documentation when contributing to SkillSpector development.

## Frequently Asked Questions

### What is the SkillspectorState TypedDict?

**`SkillspectorState`** is the central data structure defined in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) that carries input paths, cached file contents, AST representations, and finding lists through the LangGraph workflow. Every node in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py) receives and returns partial updates to this state using LangGraph’s reducer pattern.

### How do I debug my analyzer without running the full LLM pipeline?

Run the scan with the `--no-llm` flag or set `"use_llm": false` in the state when invoking programmatically. This executes only the 14 static pattern analyzers, YARA rules, and AST-based detectors, completing in seconds rather than minutes. You can also use `make langgraph-dev` to step through individual nodes in LangGraph Studio.

### Can I add LLM-based analyzers instead of static pattern matching?

Yes. While most analyzers use static patterns in `src/skillspector/nodes/analyzers/`, you can implement custom logic that calls `chat_completion` from [`src/skillspector/llm_utils.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_utils.py). However, static analyzers are preferred for deterministic, fast feedback. LLM-based filtering should typically reside in the `meta_analyzer` node or use the `LLMMetaAnalyzer` class to avoid excessive API costs.

### Where are the severity and risk score calculations defined?

The **`report`** node in [`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py) aggregates all findings, applies baseline suppression rules, and computes the numeric risk score mapped to severity levels (Low, Medium, High, Critical). When adding analyzers, ensure your `Severity` enum values align with the scoring thresholds defined in this file to maintain consistent risk assessment.