How to Use SkillSpector as a Python Library or API

SkillSpector exposes a compiled LangGraph workflow via skillspector.graph that you can invoke programmatically with graph.invoke(state) to run static and LLM-powered security analysis on AI skills without touching the CLI.

SkillSpector is NVIDIA's open-source security scanner for AI skills packaged as a standard Python library named skillspector. Instead of relying solely on command-line interactions, you can import the analysis engine directly into your Python applications to programmatically scan skill directories, zip files, or git repositories. This guide demonstrates how to use the SkillSpector Python library to embed security checks into your automation workflows.

Core Architecture and Entry Points

The library centers on a LangGraph workflow defined in src/skillspector/graph.py that orchestrates analysis nodes sequentially. Understanding this architecture helps you configure the state correctly when calling the API programmatically.

The Compiled Graph Object

The primary entry point is the compiled graph object available at skillspector.graph. According to the source code in src/skillspector/__init__.py, this is a re-export of the compiled StateGraph built by create_graph(). When you call graph.invoke(initial_state), the workflow executes six phases: input resolution, context building, static analysis, semantic analysis (optional), report generation, and cleanup.

State Management with SkillspectorState

All data flows between nodes via a single mutable object defined in src/skillspector/state.py as SkillspectorState, a TypedDict that carries the skill path, file caches, AST representations, findings, and configuration flags. When invoking the graph programmatically, you provide a dictionary matching this schema—missing fields automatically populate with defaults.

Step-by-Step Implementation Guide

To run SkillSpector programmatically, import the compiled graph and construct an initial state dictionary with your analysis parameters.

from skillspector import graph

# Configure the initial state

initial_state = {
    "input_path": "/path/to/your/skill",   # Directory, zip, git URL, or SKILL.md file

    "output_format": "json",               # Options: terminal | json | markdown | sarif

    "use_llm": True,                        # Set False for static-only analysis

}

# Execute the full analysis pipeline

result = graph.invoke(initial_state)

# Access structured results

print(f"Risk score: {result['risk_score']}/100")
print(f"Severity: {result['risk_severity']}")
print(f"Recommendation: {result['risk_recommendation']}")

# Iterate through findings

for finding in result["filtered_findings"]:
    print(f"[{finding['severity']}] {finding['rule_id']}: {finding['message']}")

# Access the formatted report body

print(result["report_body"])

The result object contains the complete SkillspectorState plus computed fields including risk_score, filtered_findings, and the rendered report_body formatted according to your output_format specification.

Understanding the Analysis Pipeline

When you invoke graph.invoke(), the library executes a fixed sequence of nodes defined in src/skillspector/graph.py:

  1. resolve_input (src/skillspector/nodes/resolve_input.py): Normalizes the user-provided path and handles temporary directory creation for remote URLs or zip archives
  2. build_context (src/skillspector/nodes/build_context.py): Traverses the skill directory, populates components, file_cache, and ast_cache, extracts the skill manifest from SKILL.md, and detects executable scripts
  3. Static analyzers (src/skillspector/nodes/analyzers/*): Executes 64 pattern detectors including regex matchers, AST analyzers, YARA rules, and OSV lookups that append Finding objects to the state
  4. meta_analyzer (src/skillspector/nodes/meta_analyzer.py): Optionally invokes an LLM (OpenAI, Anthropic, or NVIDIA nv_build) to filter false positives and enrich explanations when use_llm is True
  5. report (src/skillspector/nodes/report.py): Computes the final risk score and generates the formatted output

Temporary directories created for git URLs or zip files are automatically cleaned up after the graph completes execution.

Configuration and Environment Variables

When enabling LLM analysis by setting "use_llm": True in your state, you must configure provider credentials via environment variables. The meta_analyzer node reads the SKILLSPECTOR_PROVIDER variable to determine which client to initialize from src/skillspector/providers/*. Set OPENAI_API_KEY, ANTHROPIC_API_KEY, or NVIDIA Build credentials as required by your selected provider.

Practical Code Examples

Basic Static-Only Scan

For CI environments where you want fast analysis without LLM latency, disable semantic analysis:

from skillspector import graph

state = {
    "input_path": "https://github.com/example/my-skill",
    "output_format": "terminal",
    "use_llm": False,
}

result = graph.invoke(state)
print(result["report_body"])

Loading Credentials from Environment Files

When embedding SkillSpector in existing applications, load API keys securely using python-dotenv:

import os
from pathlib import Path
from dotenv import load_dotenv
from skillspector import graph

load_dotenv(Path(__file__).with_name(".env"))

state = {
    "input_path": "./my-skill/",
    "output_format": "markdown",
    "use_llm": True,
}

result = graph.invoke(state)
Path("skill_report.md").write_text(result["report_body"])

Embedding in CI/CD Pipelines

Create a reusable wrapper function to integrate SkillSpector into larger automation systems:

from skillspector import graph
import json
import logging

def scan_skill(path: str) -> dict:
    """Run SkillSpector analysis and return standardized risk metrics."""
    state = {
        "input_path": path,
        "output_format": "json",
        "use_llm": True,
    }
    result = graph.invoke(state)
    
    return {
        "score": result["risk_score"],
        "severity": result["risk_severity"],
        "recommendation": result["risk_recommendation"],
        "findings": result["filtered_findings"],
    }

if __name__ == "__main__":
    report = scan_skill("./candidate-skill/")
    logging.info("Skill risk: %s (%d)", report["severity"], report["score"])
    print(json.dumps(report, indent=2))

Summary

  • Import the compiled graph from skillspector (defined in src/skillspector/__init__.py) to access the full LangGraph workflow
  • Construct a state dictionary matching SkillspectorState with at minimum input_path, output_format, and use_llm keys
  • Call graph.invoke(state) to execute the pipeline through resolve_input, build_context, static analyzers, meta_analyzer, and report nodes
  • Access results via result["risk_score"], result["filtered_findings"], and result["report_body"] for programmatic consumption
  • Configure LLM providers using environment variables when use_llm is True, or set to False for static-only analysis

Frequently Asked Questions

How do I import the SkillSpector graph in my Python script?

Import the compiled graph directly from the package namespace: from skillspector import graph. This object is re-exported from src/skillspector/__init__.py and represents the compiled LangGraph workflow ready for immediate invocation.

What state parameters are required to invoke SkillSpector programmatically?

You must provide at least input_path (string pointing to a directory, zip, git URL, or SKILL.md file) and output_format (one of terminal, json, markdown, or sarif). The use_llm boolean is optional and defaults to False. All other fields in the SkillspectorState TypedDict from src/skillspector/state.py are populated with sensible defaults.

Can I run SkillSpector analysis without using an LLM?

Yes. Set "use_llm": False in your initial state dictionary. This skips the meta_analyzer node entirely and runs only the static analyzers (regex, AST, YARA, and OSV lookups) from src/skillspector/nodes/analyzers/*, producing faster results without requiring API keys.

How does SkillSpector handle temporary files from git URLs or zip archives?

The resolve_input node in src/skillspector/nodes/resolve_input.py automatically clones repositories or extracts archives to temporary directories, storing the path in temp_dir_for_cleanup. The graph automatically removes these temporary resources after execution completes, requiring no manual cleanup from your application code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →