# How to Integrate SkillSpector into a Python Project Programmatically

> Integrate SkillSpector into your Python project programmatically. Import the graph, create a state dictionary, and invoke the analysis pipeline for seamless security scanning without the CLI.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: how-to-guide
- Published: 2026-07-11

---

**You can integrate SkillSpector programmatically by importing the compiled `graph` singleton from `skillspector.graph`, constructing a state dictionary using the `_scan_state` helper from [`cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/cli.py), and invoking `graph.invoke(state)` to execute the full security analysis pipeline without using the CLI.**

NVIDIA/SkillSpector is a LangGraph-based security scanner that exposes its core workflow as a ready-made graph. Instead of shelling out to the command line, you can import the `skillspector` package directly into your Python application to analyze skills, archives, or directories while maintaining full control over input, configuration, and result handling.

## Import the Compiled Graph

The entry point for programmatic integration lives in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py). This module exposes a **pre-compiled graph singleton** named `graph` that wires together all analyzer nodes—including input resolution, context building, static analysis, meta-analysis, and reporting.

You have two options for initialization:

- **`from skillspector.graph import graph`** – Imports the singleton compiled at import time (recommended for most use cases).
- **`from skillspector.graph import create_graph`** – Call this function to build a fresh `StateGraph` instance if you need custom node wiring.

The graph expects a **state dictionary** matching the `SkillspectorState` model defined in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py). This Pydantic-style object tracks the workflow's data through keys like `input_path`, `output_format`, and `use_llm`.

## Build the Initial State

To ensure your state dictionary contains all required keys, mirror the CLI's internal logic by using the `_scan_state()` helper from [`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py). This function maps high-level parameters to the exact structure the graph consumes.

Required arguments include:

- **`input_path`** – Path to a skill directory, zip archive, or remote URL.
- **`format`** – Output format (`"terminal"`, `"json"`, `"markdown"`, or `"sarif"`).
- **`no_llm`** – Boolean to disable LLM-based analyzers (set `True` to skip LLM analysis).

Optional keys you can inject directly into the state dict include `yara_rules_dir`, `baseline`, and provider-specific configurations.

## Invoke the Analysis Pipeline

Execute the workflow by calling `graph.invoke()` with your constructed state. Optionally pass a LangSmith trace configuration via the `config` parameter to enable observability.

```python
from skillspector.graph import graph

result = graph.invoke(state, config={"run_name": "my-scan", "tags": ["api"]})

```

The returned dictionary contains:

- **`risk_score`** – Integer from 0 to 100.
- **`risk_severity`** – String value (`"LOW"`, `"MEDIUM"`, or `"HIGH"`).
- **`findings`** / **`filtered_findings`** – Lists of vulnerability objects.
- **`report_body`** – Pre-rendered output matching your chosen format.

## Handle Multi-Skill Directories

For repositories containing multiple independent skills, use the `detect_skills()` utility from [`src/skillspector/multi_skill.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/multi_skill.py) before invoking the graph. This returns a `MultiSkillDetectionResult` indicating whether the path represents a single skill or a multi-skill layout, allowing you to loop over sub-skills individually.

```python
from skillspector.multi_skill import detect_skills

detection = detect_skills("/path/to/repo")
if detection.is_multi_skill:
    for skill in detection.skills:
        # Invoke graph for each skill.path

        pass

```

## Complete Integration Example

Below is a self-contained script demonstrating the full integration pattern. It builds the state using the CLI helper, invokes the graph, and handles both single and multi-skill layouts.

```python

# example_programmatic_scan.py

from pathlib import Path
from skillspector.graph import graph
from skillspector.cli import _scan_state
from skillspector.multi_skill import detect_skills

def build_trace_config() -> dict:
    """Optional LangSmith tracing configuration."""
    return {
        "run_name": "programmatic-scan",
        "tags": ["skillspector", "example"],
        "metadata": {"caller": "my-script"},
    }

def run_scan(input_path: str, use_llm: bool = True) -> dict:
    """
    Execute SkillSpector on *input_path* and return the result dictionary.
    """
    # Build state dict matching SkillspectorState expectations

    state = _scan_state(
        input_path=input_path,
        format="json",
        no_llm=not use_llm,
    )
    
    # Invoke the compiled LangGraph workflow

    result = graph.invoke(state, config=build_trace_config())
    return result

def main():
    skill_dir = Path("./my-skill/").resolve()
    if not skill_dir.exists():
        raise FileNotFoundError(f"Skill directory not found: {skill_dir}")

    # Detect multi-skill layouts

    detection = detect_skills(skill_dir)
    if detection.is_multi_skill:
        print(f"Detected {len(detection.skills)} independent skills:")
        for sub in detection.skills:
            print(f"  Scanning {sub.name}...")
            sub_result = run_scan(str(sub.path))
            print(f"  Score: {sub_result.get('risk_score')} ({sub_result.get('risk_severity')})")
        return

    # Single skill scan

    result = run_scan(str(skill_dir))
    print(result.get("report_body"))

if __name__ == "__main__":
    # Optional: Configure LLM provider via environment variables

    # import os

    # os.environ["SKILLSPECTOR_PROVIDER"] = "openai"

    # os.environ["OPENAI_API_KEY"] = "sk-..."

    main()

```

## Summary

- **Import the graph** from [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py)—use the pre-compiled `graph` singleton for immediate execution or `create_graph()` for customization.
- **Build state correctly** by leveraging `_scan_state()` in [`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py) to generate the dictionary keys the `SkillspectorState` model requires.
- **Invoke with `graph.invoke(state)`** to run the security pipeline and receive structured results including `risk_score`, `findings`, and `report_body`.
- **Support multi-skill repos** using `detect_skills()` from [`src/skillspector/multi_skill.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/multi_skill.py) to iterate over individual skills within a larger directory.

## Frequently Asked Questions

### What is the difference between `create_graph()` and the `graph` import?

`create_graph()` constructs and returns a new `StateGraph` instance that you can modify before calling `.compile()`. The `graph` import is a module-level singleton that is already compiled and ready to use, which is the recommended entry point for standard programmatic integrations as implemented in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py).

### Can I run SkillSpector without installing the CLI dependencies?

Yes. The core analysis logic resides in the `skillspector` package and depends only on LangGraph and the analyzer libraries. You can import `skillspector.graph` and `skillspector.state` in a minimal environment without the CLI-specific dependencies, provided you handle environment variables for LLM providers manually.

### How do I disable LLM-based analysis when calling the graph programmatically?

Set `no_llm=True` when calling `_scan_state()` from [`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py) (or set `use_llm: false` directly in your state dictionary). This omits the LLM provider nodes from the execution graph, running only static analyzers like YARA and pattern matching.

### What keys are guaranteed to exist in the result dictionary returned by `graph.invoke()`?

According to the `SkillspectorState` definition in [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) and the terminal nodes in [`src/skillspector/graph.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/graph.py), the result always contains `risk_score` (int), `risk_severity` (str), `findings` (list), and `report_body` (str). Optional keys like `sarif_report` appear only when the output format is set to `"sarif"`.