How to Integrate SkillSpector into a Python Project Programmatically
You can integrate SkillSpector programmatically by importing the compiled graph singleton from skillspector.graph, constructing a state dictionary using the _scan_state helper from cli.py, and invoking graph.invoke(state) to execute the full security analysis pipeline without using the CLI.
NVIDIA/SkillSpector is a LangGraph-based security scanner that exposes its core workflow as a ready-made graph. Instead of shelling out to the command line, you can import the skillspector package directly into your Python application to analyze skills, archives, or directories while maintaining full control over input, configuration, and result handling.
Import the Compiled Graph
The entry point for programmatic integration lives in src/skillspector/graph.py. This module exposes a pre-compiled graph singleton named graph that wires together all analyzer nodes—including input resolution, context building, static analysis, meta-analysis, and reporting.
You have two options for initialization:
from skillspector.graph import graph– Imports the singleton compiled at import time (recommended for most use cases).from skillspector.graph import create_graph– Call this function to build a freshStateGraphinstance if you need custom node wiring.
The graph expects a state dictionary matching the SkillspectorState model defined in src/skillspector/state.py. This Pydantic-style object tracks the workflow's data through keys like input_path, output_format, and use_llm.
Build the Initial State
To ensure your state dictionary contains all required keys, mirror the CLI's internal logic by using the _scan_state() helper from src/skillspector/cli.py. This function maps high-level parameters to the exact structure the graph consumes.
Required arguments include:
input_path– Path to a skill directory, zip archive, or remote URL.format– Output format ("terminal","json","markdown", or"sarif").no_llm– Boolean to disable LLM-based analyzers (setTrueto skip LLM analysis).
Optional keys you can inject directly into the state dict include yara_rules_dir, baseline, and provider-specific configurations.
Invoke the Analysis Pipeline
Execute the workflow by calling graph.invoke() with your constructed state. Optionally pass a LangSmith trace configuration via the config parameter to enable observability.
from skillspector.graph import graph
result = graph.invoke(state, config={"run_name": "my-scan", "tags": ["api"]})
The returned dictionary contains:
risk_score– Integer from 0 to 100.risk_severity– String value ("LOW","MEDIUM", or"HIGH").findings/filtered_findings– Lists of vulnerability objects.report_body– Pre-rendered output matching your chosen format.
Handle Multi-Skill Directories
For repositories containing multiple independent skills, use the detect_skills() utility from src/skillspector/multi_skill.py before invoking the graph. This returns a MultiSkillDetectionResult indicating whether the path represents a single skill or a multi-skill layout, allowing you to loop over sub-skills individually.
from skillspector.multi_skill import detect_skills
detection = detect_skills("/path/to/repo")
if detection.is_multi_skill:
for skill in detection.skills:
# Invoke graph for each skill.path
pass
Complete Integration Example
Below is a self-contained script demonstrating the full integration pattern. It builds the state using the CLI helper, invokes the graph, and handles both single and multi-skill layouts.
# example_programmatic_scan.py
from pathlib import Path
from skillspector.graph import graph
from skillspector.cli import _scan_state
from skillspector.multi_skill import detect_skills
def build_trace_config() -> dict:
"""Optional LangSmith tracing configuration."""
return {
"run_name": "programmatic-scan",
"tags": ["skillspector", "example"],
"metadata": {"caller": "my-script"},
}
def run_scan(input_path: str, use_llm: bool = True) -> dict:
"""
Execute SkillSpector on *input_path* and return the result dictionary.
"""
# Build state dict matching SkillspectorState expectations
state = _scan_state(
input_path=input_path,
format="json",
no_llm=not use_llm,
)
# Invoke the compiled LangGraph workflow
result = graph.invoke(state, config=build_trace_config())
return result
def main():
skill_dir = Path("./my-skill/").resolve()
if not skill_dir.exists():
raise FileNotFoundError(f"Skill directory not found: {skill_dir}")
# Detect multi-skill layouts
detection = detect_skills(skill_dir)
if detection.is_multi_skill:
print(f"Detected {len(detection.skills)} independent skills:")
for sub in detection.skills:
print(f" Scanning {sub.name}...")
sub_result = run_scan(str(sub.path))
print(f" Score: {sub_result.get('risk_score')} ({sub_result.get('risk_severity')})")
return
# Single skill scan
result = run_scan(str(skill_dir))
print(result.get("report_body"))
if __name__ == "__main__":
# Optional: Configure LLM provider via environment variables
# import os
# os.environ["SKILLSPECTOR_PROVIDER"] = "openai"
# os.environ["OPENAI_API_KEY"] = "sk-..."
main()
Summary
- Import the graph from
src/skillspector/graph.py—use the pre-compiledgraphsingleton for immediate execution orcreate_graph()for customization. - Build state correctly by leveraging
_scan_state()insrc/skillspector/cli.pyto generate the dictionary keys theSkillspectorStatemodel requires. - Invoke with
graph.invoke(state)to run the security pipeline and receive structured results includingrisk_score,findings, andreport_body. - Support multi-skill repos using
detect_skills()fromsrc/skillspector/multi_skill.pyto iterate over individual skills within a larger directory.
Frequently Asked Questions
What is the difference between create_graph() and the graph import?
create_graph() constructs and returns a new StateGraph instance that you can modify before calling .compile(). The graph import is a module-level singleton that is already compiled and ready to use, which is the recommended entry point for standard programmatic integrations as implemented in src/skillspector/graph.py.
Can I run SkillSpector without installing the CLI dependencies?
Yes. The core analysis logic resides in the skillspector package and depends only on LangGraph and the analyzer libraries. You can import skillspector.graph and skillspector.state in a minimal environment without the CLI-specific dependencies, provided you handle environment variables for LLM providers manually.
How do I disable LLM-based analysis when calling the graph programmatically?
Set no_llm=True when calling _scan_state() from src/skillspector/cli.py (or set use_llm: false directly in your state dictionary). This omits the LLM provider nodes from the execution graph, running only static analyzers like YARA and pattern matching.
What keys are guaranteed to exist in the result dictionary returned by graph.invoke()?
According to the SkillspectorState definition in src/skillspector/state.py and the terminal nodes in src/skillspector/graph.py, the result always contains risk_score (int), risk_severity (str), findings (list), and report_body (str). Optional keys like sarif_report appear only when the output format is set to "sarif".
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →