How to Scan Multi-Skill Directories Recursively with SkillSpector
SkillSpector's --recursive flag detects and scans independent skills within a parent directory, aggregating results into a single report or individual outputs.
The NVIDIA SkillSpector repository provides a specialized command-line interface for analyzing AI skills defined in SKILL.md files. When working with repositories that contain multiple independent skills organized in subdirectories, you need a way to scan them all at once without invoking the tool separately for each one. The recursive scanning feature treats a parent directory as a collection of separate skills, automatically discovering and analyzing each one according to the detection logic implemented in the source code.
How Recursive Scanning Works in SkillSpector
Recursive scanning follows a detect-then-scan workflow that first identifies whether a directory qualifies as a multi-skill repository, then processes each skill independently through the standard analysis pipeline.
Detection Logic in multi_skill.py
The core detection algorithm resides in src/skillspector/multi_skill.py. According to the source code, a directory qualifies as a multi-skill directory only when two conditions are met:
- The root directory does not contain a top-level
SKILL.mdfile. - At least two immediate subdirectories each contain their own
SKILL.mdfile.
If a root SKILL.md exists, SkillSpector treats the entire directory as a single skill regardless of nested skill files. This logic is implemented in the detect_skills() function at lines 51-61 of multi_skill.py.
CLI Flag Definition
The --recursive (or -r) flag is defined in src/skillspector/cli.py at lines 217-224. When present, the CLI invokes detect_skills() at lines 80-86 to determine whether the target path represents a multi-skill structure before proceeding with the appropriate scan strategy.
Scanning Workflow for Multi-Skill Repositories
Once the detection phase confirms a multi-skill structure, SkillSpector orchestrates the scan through specialized functions that handle parallel execution and result aggregation.
The detect_skills() Function
This helper function analyzes the directory structure and returns a DetectionResult indicating whether the target is a multi-skill repository. It identifies valid SkillDirectory objects by checking for the presence of SKILL.md files in immediate subdirectories using the internal _has_skill_md and _extract_skill_name helpers.
The _scan_multi_skill() Implementation
When detect_skills() returns is_multi_skill=True, the CLI forwards the result to _scan_multi_skill() at lines 59-66 in cli.py. This routine:
- Iterates over every detected
SkillDirectoryobject. - Invokes the normal scan pipeline (
_scan_state→graph.invoke) for each sub-skill independently. - Collects per-skill results and prints a summary to the console.
- Writes a combined JSON report when
--format jsonand--outputare specified (lines 70-88 and 94-112).
The function leverages the LangGraph pipeline defined in src/skillspector/graph.py to perform the actual analysis for each skill.
Command-Line Usage Examples
Use the --recursive flag to scan directories containing multiple independent skills. The following commands demonstrate common usage patterns:
# Scan a directory containing several independent skills
skillspector scan ./skill-collection/ --recursive
# Generate a combined JSON report for all discovered skills
skillspector scan ./skill-collection/ --recursive --format json --output multi_report.json
# Use the short flag form
skillspector scan ./skill-collection/ -r
Programmatic Usage with Python
You can replicate the CLI's recursive scanning behavior in Python by importing the detection and scanning functions directly:
from pathlib import Path
from skillspector.multi_skill import detect_skills
from skillspector.cli import _scan_multi_skill, FormatChoice, _scan_state, _build_trace_config
from skillspector.graph import graph
# Define the target directory
root = Path("./skill-collection/").resolve()
# Detect sub-skills recursively
detection = detect_skills(root)
if detection.is_multi_skill:
# Execute multi-skill scan with JSON output
_scan_multi_skill(
detection=detection,
format=FormatChoice.json,
output=Path("multi_report.json"),
no_llm=False,
yara_rules_dir=None,
verbose=True,
)
else:
# Fallback to single-skill scan
state = _scan_state(str(root), FormatChoice.json, no_llm=False)
result = graph.invoke(
state,
config=_build_trace_config(str(root), FormatChoice.json, False)
)
# Process result as needed
Understanding Exit Codes and Risk Aggregation
After completing a multi-skill scan, SkillSpector determines the process exit code based on the highest risk score detected among all sub-skills. According to lines 39-41 in src/skillspector/cli.py, the process exits with a failure code if any skill reports a risk score greater than 50. This ensures that CI/CD pipelines can fail builds when any skill in the repository exhibits high-risk characteristics.
Summary
- Use
--recursiveor-rto enable multi-skill directory scanning in the SkillSpector CLI. - Detection requires no root-level
SKILL.mdand at least two subdirectories containingSKILL.mdfiles. - Implementation spans
src/skillspector/multi_skill.py(detection) andsrc/skillspector/cli.py(orchestration). - Results aggregation combines individual skill outputs into a single report when using
--format json --output. - Exit codes reflect the highest risk score across all scanned skills, failing if any exceed 50.
Frequently Asked Questions
What defines a multi-skill directory in SkillSpector?
A directory qualifies as a multi-skill repository when it lacks a SKILL.md file at its root level while containing at least two immediate subdirectories that each have their own SKILL.md. If a root-level skill file exists, SkillSpector treats the entire directory as a single skill regardless of nested structures.
How does SkillSpector handle nested skill directories?
SkillSpector only examines immediate subdirectories when detecting multi-skill structures. The detect_skills() function in src/skillspector/multi_skill.py checks for SKILL.md files one level deep, meaning deeply nested skill directories beyond the first level are not automatically discovered in recursive scans.
Can I export recursive scan results as a single JSON file?
Yes. When using --recursive combined with --format json and --output, the _scan_multi_skill() function aggregates all individual skill results into a single combined JSON report. This output contains the analysis results for every discovered skill in the directory structure.
What exit code does SkillSpector return when scanning multiple skills?
SkillSpector returns a failure exit code if any skill in the multi-skill scan reports a risk score greater than 50. The exit code logic at lines 39-41 of src/skillspector/cli.py uses the maximum risk score across all scanned skills to determine the final process status, ensuring high-risk skills in any subdirectory trigger a pipeline failure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →