How SkillSpector Handles Multiple SKILL.md Files for Multi-Skill Detection
SkillSpector detects multiple SKILL.md files by scanning subdirectories for individual skill manifests while checking for a root-level SKILL.md, treating directories as multi-skill only when two or more independent skill directories are found without a root manifest.
NVIDIA/SkillSpector provides a recursive scanning mode that automatically discovers and processes directories containing multiple independent skills. When analyzing a codebase with several SKILL.md files spread across subdirectories, the tool uses a specific detection algorithm defined in src/skillspector/multi_skill.py to determine whether to treat the path as a single skill or as a multi-skill collection.
The Multi-Skill Detection Algorithm
The core logic resides in src/skillspector/multi_skill.py, specifically within the detect_skills() function. This implementation follows a six-step validation process to categorize directories correctly and returns a MultiSkillDetectionResult dataclass.
Step 1: Validate the Target Path
First, the function verifies the supplied path is actually a directory. If the path is not a folder, the function immediately returns a negative result.
# src/skillspector/multi_skill.py (lines 63-65)
if not directory.is_dir():
# Early return for non-directory paths
return MultiSkillDetectionResult(...)
Step 2: Check for Root-Level SKILL.md
The algorithm checks for a SKILL.md or skill.md at the root level using _has_skill_md(). If present, the entire directory is treated as a single skill regardless of nested files, setting has_root_skill=True and returning early.
# src/skillspector/multi_skill.py (lines 94-97)
has_root = _has_skill_md(directory)
if has_root:
return MultiSkillDetectionResult(has_root_skill=True, ...)
Step 3: Scan Immediate Subdirectories
If no root manifest exists, the code iterates through direct children using sorted(directory.iterdir()). It skips non-directories and hidden folders, checking each child for its own SKILL.md via _has_skill_md(child).
# src/skillspector/multi_skill.py (lines 70-78)
for child in sorted(directory.iterdir()):
if not child.is_dir() or child.name.startswith("."):
continue
if _has_skill_md(child):
# Process as independent skill directory
Step 4: Extract Skill Names
For each valid skill directory found, _extract_skill_name() parses the frontmatter using yaml.safe_load. The function uses the name key from the YAML if available, otherwise falling back to the directory name.
# src/skillspector/multi_skill.py (lines 99-129)
def _extract_skill_name(skill_dir: Path) -> str:
skill_md = skill_dir / "SKILL.md"
content = skill_md.read_text()
# Parse YAML frontmatter and return name or directory name
Step 5: Determine Multi-Skill Status
The function counts discovered skills and sets is_multi_skill=True only when len(skills) >= 2. This threshold ensures that single-skill directories are not processed as multi-skill collections.
# src/skillspector/multi_skill.py (lines 86-91)
is_multi = len(skills) >= 2
return MultiSkillDetectionResult(
is_multi_skill=is_multi,
skills=skills,
has_root_skill=False
)
Step 6: Return Structured Results
The function returns a MultiSkillDetectionResult dataclass containing is_multi_skill, the list of SkillDirectory objects, and the has_root_skill boolean flag.
CLI Integration with Recursive Scanning
The command-line interface in src/skillspector/cli.py leverages this detection logic when the --recursive flag is passed. At lines 73-78, the CLI resolves the path and calls detect_skills().
# src/skillspector/cli.py (lines 73-78)
if recursive and resolved_path.is_dir():
detection = detect_skills(resolved_path)
if detection.is_multi_skill:
_scan_multi_skill(detection, format, output, no_llm, yara_rules_dir, verbose)
return
If detection.is_multi_skill is true, execution delegates to _scan_multi_skill(), which walks each discovered skill directory separately and produces independent reports. If no root skill exists and no sub-skills are found, the tool falls back to single-skill mode with a warning.
Practical Implementation Examples
Direct Python API Usage
You can programmatically detect multiple skills without invoking the CLI:
from pathlib import Path
from skillspector.multi_skill import detect_skills
# Point to the directory you want to scan
directory = Path("/path/to/skill-collection")
result = detect_skills(directory)
if result.is_multi_skill:
print(f"Detected {len(result.skills)} independent skills:")
for skill in result.skills:
print(f" • {skill.name} (at {skill.relative_path})")
elif result.has_root_skill:
print("Single-skill directory (root SKILL.md present).")
else:
print("No SKILL.md files found – nothing to scan.")
Command-Line Execution
Scan a collection of skills and treat each sub-skill independently:
skillspector scan ./my-skill-collection --recursive
If the folder contains skill-a/SKILL.md and skill-b/SKILL.md but no top-level SKILL.md, the CLI automatically runs two separate scans and emits two distinct reports.
Unit Test Validation
The repository's test suite in tests/test_multi_skill.py exercises the detection logic:
# tests/test_multi_skill.py
result = detect_skills(multi_skill_dir)
assert result.is_multi_skill
assert len(result.skills) == 3
assert result.skills[0].name == "first"
Summary
- Root-level precedence: A
SKILL.mdat the directory root causes the entire folder to be treated as a single skill, overriding any nested manifests. - Two-skill threshold: The algorithm requires at least two independent skill directories to set
is_multi_skill=True, preventing single-skill directories from unnecessary multi-skill processing. - Frontmatter parsing: Skill names are extracted from YAML frontmatter via
_extract_skill_name()or derived from directory names when thenamekey is absent. - Recursive CLI flag: The
--recursiveflag triggersdetect_skills()insrc/skillspector/cli.pyand delegates to_scan_multi_skill()for batch processing of independent reports. - Structured results: The
MultiSkillDetectionResultdataclass provides boolean flagsis_multi_skillandhas_root_skill, plus a list ofSkillDirectoryobjects for downstream processing.
Frequently Asked Questions
What happens if a directory contains both a root SKILL.md and subdirectories with their own SKILL.md files?
When a root-level SKILL.md is present, SkillSpector treats the entire directory as a single skill. The _has_skill_md() helper detects the root manifest at lines 94-97 in src/skillspector/multi_skill.py, causing an early return with has_root_skill=True. Nested SKILL.md files are ignored in this scenario, ensuring the root manifest takes precedence.
How does SkillSpector determine the name of each detected skill?
The _extract_skill_name() function in src/skillspector/multi_skill.py (lines 99-129) parses the frontmatter of each SKILL.md using yaml.safe_load. If a name key exists in the YAML header, it is used as the skill name; otherwise, the function falls back to the directory name. This ensures consistent identification even when metadata is missing.
Why does the detection require at least two skills to be considered multi-skill?
The algorithm specifically checks is_multi = len(skills) >= 2 at lines 86-91 in src/skillspector/multi_skill.py. This threshold prevents single-skill directories from being processed through the multi-skill pipeline, ensuring appropriate handling of isolated skill manifests without unnecessary overhead or report fragmentation.
Can I use the multi-skill detection programmatically without the CLI?
Yes, you can import detect_skills directly from skillspector.multi_skill and pass a pathlib.Path object to analyze any directory. The function returns a MultiSkillDetectionResult object containing boolean flags and skill metadata, allowing you to build custom workflows around the detection logic without invoking the command-line interface.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →