How SkillSpector Handles Recursive Directory Scanning for Skill Collections
SkillSpector performs recursive directory scanning by detecting immediate subdirectories containing SKILL.md files and processing each as an independent skill through a reusable LangGraph workflow, triggered via the --recursive CLI flag.
SkillSpector, NVIDIA's security analysis tool for AI skills, supports recursive directory scanning to audit multiple independent skills in a single command. When you invoke the skillspector scan command with the --recursive flag, the tool treats each immediate subdirectory containing a SKILL.md file as a separate skill package, running the full analysis pipeline on each while preserving per-skill reporting.
How Recursive Scanning Works
CLI Entry Point and Flag Detection
In src/skillspector/cli.py, the scan function checks for the --recursive (-r) flag and validates that the supplied path is a directory:
if recursive and resolved_path.is_dir():
detection = detect_skills(resolved_path)
if detection.is_multi_skill:
_scan_multi_skill(detection, format, output, no_llm, yara_rules_dir, verbose)
return
(source: cli.py lines 277-284)
This logic branches execution to the multi-skill handler when the flag is present and the target is a directory.
Multi-Skill Detection Logic
The detect_skills function in src/skillspector/multi_skill.py implements the discovery logic. It walks only the immediate subdirectories of the root path, skipping hidden directories (those starting with .):
for child in sorted(directory.iterdir()):
if not child.is_dir() or child.name.startswith("."):
continue
if _has_skill_md(child):
name = _extract_skill_name(child)
skills.append(SkillDirectory(path=child, name=name,
relative_path=child.name))
is_multi = len(skills) >= 2
(source: multi_skill.py lines 70-88)
A directory qualifies as a skill if it contains a SKILL.md or skill.md file. The function returns a MultiSkillDetectionResult where is_multi_skill is True only when two or more valid skill directories are found and no root-level SKILL.md exists. Notably, this design does not recurse deeper than one level; nested collections require explicit re-invocation.
Per-Skill Analysis Execution
Once detected, the _scan_multi_skill function (in src/skillspector/cli.py) iterates over each SkillDirectory object, constructs a fresh graph state via _scan_state, and invokes the analysis pipeline:
# From cli.py lines 55-84 and 90-114
for skill in detection.skills:
state = _scan_state(str(skill.path), format, no_llm, yara_rules_dir, verbose)
result = graph.invoke(state, config=trace_config)
The workflow graph itself is defined in src/skillspector/graph.py (lines 34-53) and is compiled once, then reused for every skill invocation to ensure consistent analysis across the collection. Results are aggregated and output as combined JSON or concatenated reports.
Practical Usage Examples
Command Line Scanning
Scan a single skill directory without recursion:
skillspector scan ./my_skill/
Scan a collection where each immediate subdirectory is an independent skill:
skillspector scan ./skill_collection/ --recursive
Generate a combined JSON report for the entire collection:
skillspector scan ./skill_collection/ -r -f json -o combined_report.json
Programmatic Access
You can also trigger recursive scanning programmatically:
from pathlib import Path
from skillspector.cli import _scan_state, _build_trace_config
from skillspector.graph import graph
from skillspector.multi_skill import detect_skills
root = Path("./skill_collection")
det = detect_skills(root)
if det.is_multi_skill:
for skill in det.skills:
skill_state = _scan_state(str(skill.path), format="json", no_llm=False)
result = graph.invoke(
skill_state,
config=_build_trace_config(str(skill.path), "json", False)
)
print(f"{skill.name}: {result['risk_score']}/100")
else:
# Handle single-skill path
state = _scan_state(str(root), format="json", no_llm=False)
result = graph.invoke(state, config=_build_trace_config(str(root), "json", False))
Key Implementation Files
src/skillspector/cli.py: Contains thescanfunction that parses--recursiveand orchestrates multi-skill scanning via_scan_multi_skill.src/skillspector/multi_skill.py: Implementsdetect_skillsto identify valid skill directories based onSKILL.mdpresence.src/skillspector/graph.py: Defines the LangGraph workflow that processes each skill independently.src/skillspector/state.py: Holds the mutable state passed to the graph, includinginput_pathanduse_llmflags.
Summary
- Recursive scanning is triggered by the
--recursive(-r) flag on theskillspector scancommand. - The scanner walks only immediate subdirectories, treating each containing a
SKILL.mdfile as an independent skill. - Detection requires two or more valid skill directories and the absence of a root-level
SKILL.md. - Each skill is processed through the same LangGraph workflow defined in
graph.py, ensuring consistent analysis. - Results can be output as individual reports or merged into combined JSON/concatenated formats.
Frequently Asked Questions
Does SkillSpector scan nested subdirectories recursively?
No. The detect_skills function in src/skillspector/multi_skill.py only iterates over immediate children of the target directory. It skips hidden directories and does not descend into sub-subdirectories. To scan nested collections, you must invoke the scan command separately on each nested level.
What file markers identify a valid skill directory?
SkillSpector recognizes a directory as a skill if it contains either SKILL.md or skill.md (case-insensitive). The detection logic in _has_skill_md checks for these files within each immediate subdirectory.
How does SkillSpector handle single-skill directories when using the recursive flag?
If the target directory contains fewer than two valid skill subdirectories, or if it contains a SKILL.md at the root level, the is_multi_skill flag returns False. The tool then falls back to standard single-skill scanning mode, treating the entire directory as one skill package.
Is the analysis pipeline identical for recursive and single-skill scans?
Yes. The compiled LangGraph workflow defined in src/skillspector/graph.py is reused for every skill in a recursive scan. Each invocation receives a fresh state object constructed by _scan_state, ensuring isolated analysis while maintaining identical pipeline behavior across all skills.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →