How to Invoke SkillSpector's LangGraph Workflow Programmatically Using the Python API

To invoke SkillSpector's LangGraph workflow programmatically, import the compiled graph object from skillspector.graph, construct a state dictionary matching the SkillspectorState schema, and call graph.invoke(state) for synchronous execution or await graph.ainvoke(state) for asynchronous execution.

The NVIDIA/SkillSpector repository packages its core security analysis engine as a LangGraph state machine. While the tool ships with a command-line interface, you can embed the scanner directly into Python applications, CI pipelines, or notebooks by interacting with the underlying graph API. This approach gives you fine-grained control over execution, output formats, and cleanup.

Importing and Configuring the Graph

Locating the Compiled Graph

The entry point for programmatic access is the module-level variable graph defined in src/skillspector/graph.py. This object is created by the create_graph() function, which constructs a StateGraph[SkillspectorState], wires the analyzer nodes, and compiles the workflow. Import it directly without needing to recompile:

from skillspector.graph import graph

The graph is already compiled and ready for invocation, so you avoid the overhead of rebuilding the state machine on every run.

Constructing the Initial State

Before invocation, you must populate a dictionary that conforms to the SkillspectorState schema defined in src/skillspector/state.py. The most critical keys include:

  • skill_path – Absolute path to the skill bundle (directory, zip file, or URL).
  • output_format – Choose "terminal", "json", "markdown", or "sarif" (defaults to "terminal").
  • use_llm – Boolean flag; set True to enable LLM-backed analyzers or False for static-only scans.
  • Optional settings – yara_rules_dir, baseline, show_suppressed, and other configuration parameters.

The state dictionary flows through the graph's nodes—including resolve_input, build_context, the analyzer fan-out, meta_analyzer, and finally the report node—accumulating findings and metadata at each step.

Synchronous and Asynchronous Invocation

Blocking Execution with graph.invoke

For straightforward scripts or single-threaded applications, use graph.invoke() to run the entire workflow blocking:

from pathlib import Path
from skillspector.graph import graph

state = {
    "skill_path": str(Path("/path/to/your/skill_bundle")),
    "output_format": "json",
    "use_llm": False,
}

result = graph.invoke(state)

print("Risk score:", result["risk_score"])
print("Total findings:", len(result["findings"]))
print("Report:", result["report_body"])

The method returns a plain Python dictionary containing all final state fields populated by the graph nodes.

Async Execution with graph.ainvoke

For applications requiring non-blocking execution, such as web servers or async data pipelines, use graph.ainvoke():

import asyncio
from pathlib import Path
from skillspector.graph import graph

async def scan_skill():
    state = {
        "skill_path": "./my_skill.zip",
        "output_format": "terminal",
        "use_llm": True,
    }
    
    result = await graph.ainvoke(state)
    
    print(result["report_body"])
    
    for entry in result["llm_call_log"]:
        print(entry)

asyncio.run(scan_skill())

This coroutine-based approach prevents I/O blocking when the graph downloads remote skill bundles or performs LLM inference. Both invocation methods return an identical result structure.

Understanding the Result Dictionary

Core Output Fields

After graph.invoke() or await graph.ainvoke() completes, the returned dictionary contains the complete SkillspectorState plus generated reports. Key fields include:

  • findings – List of Finding objects (see src/skillspector/models.py) containing vulnerability details.
  • filtered_findings – Findings remaining after baseline suppression and filtering.
  • report_body – Human-readable report formatted according to your output_format selection.
  • sarif_report – Full SARIF JSON payload for integration with security platforms.
  • risk_score, risk_severity, risk_recommendation – Aggregated risk assessment computed by the meta_analyzer node.
  • llm_call_log – Telemetry data for each LLM invocation, including latency and token usage, present when use_llm=True.

The graph's reducer (using operator.add) automatically merges list-type fields like findings and llm_call_log as nodes execute in parallel.

Handling Temporary Directories

When the input path is a URL or zip file, the resolve_input node creates a temporary directory for extraction. The result dictionary includes a temp_dir_for_cleanup key pointing to this directory. You must delete it manually or use the helper function from src/skillspector/cleanup.py:

from skillspector.cleanup import cleanup_result
from skillspector.graph import graph

result = graph.invoke(state)

# Automatic cleanup

cleanup_result(result)

Failing to clean up leaves temporary clones or extracted archives in your filesystem.

Complete Code Examples

Static Analysis Only (Synchronous)

from pathlib import Path
from skillspector.graph import graph

state = {
    "skill_path": str(Path("/path/to/your/skill_bundle")),
    "output_format": "json",
    "use_llm": False,
}

result = graph.invoke(state)

print("Risk score:", result["risk_score"])
print("Findings:", len(result["findings"]))
print("Report (JSON):", result["report_body"])

LLM-Enabled Async Scan

import asyncio
from pathlib import Path
from skillspector.graph import graph

async def async_scan():
    state = {
        "skill_path": str(Path("./my_skill")),
        "output_format": "terminal",
        "use_llm": True,
    }

    result = await graph.ainvoke(state)

    print(result["report_body"])

    for entry in result["llm_call_log"]:
        print(entry)

asyncio.run(async_scan())

Remote URL with Cleanup

import shutil
from pathlib import Path
from skillspector.graph import graph

state = {
    "skill_path": "https://github.com/example/skill.git",
    "output_format": "markdown",
    "use_llm": True,
}

result = graph.invoke(state)

temp_dir = result.get("temp_dir_for_cleanup")
if temp_dir:
    shutil.rmtree(temp_dir)

Path("skill_report.md").write_text(result["report_body"], encoding="utf-8")
print("Report saved to skill_report.md")

Summary

  • Import the compiled graph from src/skillspector/graph.py to access the graph variable.
  • Construct a valid state dictionary using keys defined in src/skillspector/state.py, specifying skill_path, output_format, and use_llm.
  • Invoke synchronously with graph.invoke(state) or asynchronously with await graph.ainvoke(state).
  • Process the result to extract findings, report_body, sarif_report, and risk metrics.
  • Clean up temporary directories referenced by temp_dir_for_cleanup in the result, or use skillspector.cleanup.cleanup_result.

Frequently Asked Questions

What is the difference between graph.invoke and graph.ainvoke?

graph.invoke() blocks the current thread until the LangGraph workflow completes, making it suitable for scripts and batch processing. graph.ainvoke() returns a coroutine that you must await, allowing concurrent execution in async applications like FastAPI servers or Jupyter notebooks. Both methods accept the same state dictionary and return identical result structures.

Can I modify the graph before invocation?

Yes. Because graph is a standard LangGraph CompiledStateGraph object, you can access its underlying structure or create a custom graph using create_graph() from src/skillspector/graph.py. This allows you to replace specific nodes, add tracing hooks, or insert middleware between the analyzer fan-out and the meta-analyzer step.

What output formats does the programmatic API support?

The output_format state key accepts "terminal", "json", "markdown", or "sarif". The report_body field in the result dictionary contains the human-readable formatted report, while sarif_report contains the structured SARIF JSON payload regardless of the chosen format.

How do I handle authentication for private skill repositories?

When specifying a URL in skill_path, include authentication tokens in the URL string (e.g., https://token@github.com/user/repo.git) or configure Git credentials in the environment before invocation. The resolve_input node in src/skillspector/nodes/ handles the cloning process using standard Git commands, so it respects system-level credential helpers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →