How to Scan Multi-Skill Directories with the --recursive Flag in SkillSpector
SkillSpector's --recursive flag enables batch scanning of multiple independent skills within a single directory by automatically detecting sub-directories containing individual SKILL.md files and executing the analysis pipeline for each one separately.
NVIDIA's SkillSpector provides a recursive scanning mode that treats directories containing several independent skills as discrete scan targets. When you use the --recursive (or -r) flag with the scan command, the tool automatically identifies multi-skill repositories and executes the analysis pipeline for each sub-skill independently. This capability streamlines security auditing across large skill collections without requiring manual invocation for each individual skill directory.
Understanding Multi-Skill Directory Detection
According to the detection logic in src/skillspector/multi_skill.py, SkillSpector distinguishes between single-skill and multi-skill directories based on the presence and location of SKILL.md files.
The Detection Algorithm
The detect_skills() function (lines 51-61 in src/skillspector/multi_skill.py) determines whether a target directory qualifies as multi-skill. A directory is classified as multi-skill only when it satisfies two conditions simultaneously: it lacks a top-level SKILL.md file, and it contains at least two immediate sub-directories that each contain their own SKILL.md file.
Single-Skill Override Behavior
If SkillSpector finds a SKILL.md file in the root directory, it treats the entire structure as a single skill regardless of nested skill definitions. This design ensures that composite skills or monorepos with a primary skill definition aren't incorrectly fragmented during analysis.
CLI Implementation and Scanning Logic
Flag Definition and Entry Point
The --recursive option is defined in the CLI at lines 217-224 of src/skillspector/cli.py. When supplied alongside a directory path, the entry point at lines 80-86 invokes detect_skills() to evaluate whether the target requires multi-skill handling.
Multi-Skill Orchestration
Upon detecting a multi-skill structure (is_multi_skill=True), the CLI delegates to _scan_multi_skill() (lines 59-66 in src/skillspector/cli.py). This orchestrator iterates over every detected SkillDirectory object and invokes the standard analysis pipeline—_scan_state followed by graph.invoke—for each sub-skill independently.
Results are collected and optionally consolidated into a combined JSON report when using --format json with an --output path (lines 94-112).
Command-Line and Programmatic Usage
SkillSpector supports recursive scanning through both CLI invocation and direct Python API calls.
Command-Line Examples
Scan a directory containing multiple independent skills:
skillspector scan ./skill-collection/ --recursive
Generate a combined JSON report for all discovered skills:
skillspector scan ./skill-collection/ --recursive --format json --output multi_report.json
Python API Implementation
from pathlib import Path
from skillspector.multi_skill import detect_skills
from skillspector.cli import _scan_multi_skill, FormatChoice, _scan_state, _build_trace_config
from skillspector.graph import graph # the LangGraph pipeline
# Path to a directory that may hold many skills
root = Path("./skill-collection/").resolve()
# Detect sub-skills
detection = detect_skills(root)
if detection.is_multi_skill:
# Run the same logic the CLI uses
_scan_multi_skill(
detection=detection,
format=FormatChoice.json,
output=Path("multi_report.json"),
no_llm=False,
yara_rules_dir=None,
verbose=True,
)
else:
# Fallback to a normal single-skill scan
state = _scan_state(str(root), FormatChoice.json, no_llm=False)
result = graph.invoke(state, config=_build_trace_config(str(root), FormatChoice.json, False))
# …process `result` as needed…
Exit Codes and Risk Aggregation
After completing a multi-skill scan, SkillSpector calculates the process exit code based on the highest risk score discovered across all sub-skills. As implemented in lines 39-41 of src/skillspector/cli.py, the process returns a failure exit code if any individual skill produces a risk score greater than 50, ensuring that high-severity findings in any sub-skill trigger appropriate CI/CD pipeline failures.
Summary
- Use
--recursiveto scan directories containing multiple independent skills, each defined by their ownSKILL.mdfile. - The
detect_skills()function insrc/skillspector/multi_skill.pyidentifies multi-skill structures by checking for the absence of root-levelSKILL.mdand presence of skill definitions in sub-directories. - The
_scan_multi_skill()orchestrator processes each sub-skill independently through the standardgraph.invokepipeline. - Combine
--recursivewith--format jsonand--outputto generate consolidated reports for entire skill collections. - Exit codes reflect the highest risk score across all scanned skills, failing if any score exceeds 50.
Frequently Asked Questions
What happens if I use --recursive on a single-skill directory?
SkillSpector automatically falls back to standard single-skill scanning mode if the target directory contains a top-level SKILL.md file. According to the logic in src/skillspector/multi_skill.py, the recursive detection only activates when the directory structure matches the multi-skill criteria of having no root skill file but multiple sub-directory skill definitions.
Can I customize the risk threshold for exit codes?
The exit code logic in src/skillspector/cli.py (lines 39-41) uses a fixed threshold of 50 for determining scan failure. Any individual skill scoring above this threshold causes the entire multi-skill scan to return a non-zero exit code, regardless of other skills' scores.
Does the --recursive flag support deeply nested skill discovery?
The current implementation detects immediate sub-directories only. A directory qualifies as multi-skill when at least two direct sub-directories contain SKILL.md files. While the scanning logic processes each detected skill independently, it does not recursively traverse deeper than the immediate sub-directory level for initial discovery purposes.
How are results combined when using JSON output?
When you specify --format json with --output, the _scan_multi_skill() function aggregates individual scan results into a single JSON file. This combined report contains the analysis data from all discovered skills while maintaining separation between individual skill results for traceability and debugging purposes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →