How to Invoke SkillSpector's LangGraph Workflow Programmatically Using the Python API
To invoke SkillSpector's LangGraph workflow programmatically, import the compiled graph object from skillspector.graph, construct a state dictionary matching the SkillspectorState schema, and call graph.invoke(state) for synchronous execution or await graph.ainvoke(state) for asynchronous execution.
The NVIDIA/SkillSpector repository packages its core security analysis engine as a LangGraph state machine. While the tool ships with a command-line interface, you can embed the scanner directly into Python applications, CI pipelines, or notebooks by interacting with the underlying graph API. This approach gives you fine-grained control over execution, output formats, and cleanup.
Importing and Configuring the Graph
Locating the Compiled Graph
The entry point for programmatic access is the module-level variable graph defined in src/skillspector/graph.py. This object is created by the create_graph() function, which constructs a StateGraph[SkillspectorState], wires the analyzer nodes, and compiles the workflow. Import it directly without needing to recompile:
from skillspector.graph import graph
The graph is already compiled and ready for invocation, so you avoid the overhead of rebuilding the state machine on every run.
Constructing the Initial State
Before invocation, you must populate a dictionary that conforms to the SkillspectorState schema defined in src/skillspector/state.py. The most critical keys include:
skill_path– Absolute path to the skill bundle (directory, zip file, or URL).output_format– Choose"terminal","json","markdown", or"sarif"(defaults to"terminal").use_llm– Boolean flag; setTrueto enable LLM-backed analyzers orFalsefor static-only scans.- Optional settings –
yara_rules_dir,baseline,show_suppressed, and other configuration parameters.
The state dictionary flows through the graph's nodes—including resolve_input, build_context, the analyzer fan-out, meta_analyzer, and finally the report node—accumulating findings and metadata at each step.
Synchronous and Asynchronous Invocation
Blocking Execution with graph.invoke
For straightforward scripts or single-threaded applications, use graph.invoke() to run the entire workflow blocking:
from pathlib import Path
from skillspector.graph import graph
state = {
"skill_path": str(Path("/path/to/your/skill_bundle")),
"output_format": "json",
"use_llm": False,
}
result = graph.invoke(state)
print("Risk score:", result["risk_score"])
print("Total findings:", len(result["findings"]))
print("Report:", result["report_body"])
The method returns a plain Python dictionary containing all final state fields populated by the graph nodes.
Async Execution with graph.ainvoke
For applications requiring non-blocking execution, such as web servers or async data pipelines, use graph.ainvoke():
import asyncio
from pathlib import Path
from skillspector.graph import graph
async def scan_skill():
state = {
"skill_path": "./my_skill.zip",
"output_format": "terminal",
"use_llm": True,
}
result = await graph.ainvoke(state)
print(result["report_body"])
for entry in result["llm_call_log"]:
print(entry)
asyncio.run(scan_skill())
This coroutine-based approach prevents I/O blocking when the graph downloads remote skill bundles or performs LLM inference. Both invocation methods return an identical result structure.
Understanding the Result Dictionary
Core Output Fields
After graph.invoke() or await graph.ainvoke() completes, the returned dictionary contains the complete SkillspectorState plus generated reports. Key fields include:
findings– List ofFindingobjects (seesrc/skillspector/models.py) containing vulnerability details.filtered_findings– Findings remaining after baseline suppression and filtering.report_body– Human-readable report formatted according to youroutput_formatselection.sarif_report– Full SARIF JSON payload for integration with security platforms.risk_score,risk_severity,risk_recommendation– Aggregated risk assessment computed by themeta_analyzernode.llm_call_log– Telemetry data for each LLM invocation, including latency and token usage, present whenuse_llm=True.
The graph's reducer (using operator.add) automatically merges list-type fields like findings and llm_call_log as nodes execute in parallel.
Handling Temporary Directories
When the input path is a URL or zip file, the resolve_input node creates a temporary directory for extraction. The result dictionary includes a temp_dir_for_cleanup key pointing to this directory. You must delete it manually or use the helper function from src/skillspector/cleanup.py:
from skillspector.cleanup import cleanup_result
from skillspector.graph import graph
result = graph.invoke(state)
# Automatic cleanup
cleanup_result(result)
Failing to clean up leaves temporary clones or extracted archives in your filesystem.
Complete Code Examples
Static Analysis Only (Synchronous)
from pathlib import Path
from skillspector.graph import graph
state = {
"skill_path": str(Path("/path/to/your/skill_bundle")),
"output_format": "json",
"use_llm": False,
}
result = graph.invoke(state)
print("Risk score:", result["risk_score"])
print("Findings:", len(result["findings"]))
print("Report (JSON):", result["report_body"])
LLM-Enabled Async Scan
import asyncio
from pathlib import Path
from skillspector.graph import graph
async def async_scan():
state = {
"skill_path": str(Path("./my_skill")),
"output_format": "terminal",
"use_llm": True,
}
result = await graph.ainvoke(state)
print(result["report_body"])
for entry in result["llm_call_log"]:
print(entry)
asyncio.run(async_scan())
Remote URL with Cleanup
import shutil
from pathlib import Path
from skillspector.graph import graph
state = {
"skill_path": "https://github.com/example/skill.git",
"output_format": "markdown",
"use_llm": True,
}
result = graph.invoke(state)
temp_dir = result.get("temp_dir_for_cleanup")
if temp_dir:
shutil.rmtree(temp_dir)
Path("skill_report.md").write_text(result["report_body"], encoding="utf-8")
print("Report saved to skill_report.md")
Summary
- Import the compiled graph from
src/skillspector/graph.pyto access thegraphvariable. - Construct a valid state dictionary using keys defined in
src/skillspector/state.py, specifyingskill_path,output_format, anduse_llm. - Invoke synchronously with
graph.invoke(state)or asynchronously withawait graph.ainvoke(state). - Process the result to extract
findings,report_body,sarif_report, and risk metrics. - Clean up temporary directories referenced by
temp_dir_for_cleanupin the result, or useskillspector.cleanup.cleanup_result.
Frequently Asked Questions
What is the difference between graph.invoke and graph.ainvoke?
graph.invoke() blocks the current thread until the LangGraph workflow completes, making it suitable for scripts and batch processing. graph.ainvoke() returns a coroutine that you must await, allowing concurrent execution in async applications like FastAPI servers or Jupyter notebooks. Both methods accept the same state dictionary and return identical result structures.
Can I modify the graph before invocation?
Yes. Because graph is a standard LangGraph CompiledStateGraph object, you can access its underlying structure or create a custom graph using create_graph() from src/skillspector/graph.py. This allows you to replace specific nodes, add tracing hooks, or insert middleware between the analyzer fan-out and the meta-analyzer step.
What output formats does the programmatic API support?
The output_format state key accepts "terminal", "json", "markdown", or "sarif". The report_body field in the result dictionary contains the human-readable formatted report, while sarif_report contains the structured SARIF JSON payload regardless of the chosen format.
How do I handle authentication for private skill repositories?
When specifying a URL in skill_path, include authentication tokens in the URL string (e.g., https://token@github.com/user/repo.git) or configure Git credentials in the environment before invocation. The resolve_input node in src/skillspector/nodes/ handles the cloning process using standard Git commands, so it respects system-level credential helpers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →