How to Handle Large Multi-Skill Repositories with SkillSpector: Complete Detection and Processing Guide

SkillSpector handles large multi-skill repositories by detecting individual skills within sub-directories and processing each through isolated LangGraph workflows, preventing memory bloat and cross-contamination of findings.

Managing repositories that contain dozens of independent LLM-enabled skills presents unique scalability challenges. When you need to handle large multi-skill repositories with SkillSpector, the tool's Multi-Skill Detector automatically segments the codebase into discrete analysis units. This architecture ensures that each skill generates separate reports while maintaining bounded memory usage, as implemented in the NVIDIA/SkillSpector open-source project.

Multi-Skill Detection Logic in multi_skill.py

The detection system resides in src/skillspector/multi_skill.py and follows a deterministic process to identify whether a repository should be treated as a single skill or a collection of independent ones.

Root-Level SKILL.md Verification

First, the _has_skill_md(directory) function (lines 94-96) checks for a top-level SKILL.md or skill.md file. When present, the entire repository is processed as a single skill, bypassing sub-directory scanning entirely. This design decision prevents unnecessary filesystem traversal when the repository explicitly declares itself as one cohesive unit.

Sub-Directory Discovery and Filtering

If no root skill file exists, the detector iterates through immediate sub-directories while filtering out hidden folders (line 74). For each candidate directory, _has_skill_md(child) is called again (line 76) to identify potential skill locations. This selective scanning reduces I/O overhead by ignoring non-skill folders and system directories.

Skill Name Extraction and Fallback Handling

When a skill file is found, _extract_skill_name(child) parses the front-matter (lines 99-129) to extract the human-readable name field. If the front-matter is malformed or lacks a name, the detector falls back to the directory name (line 29), ensuring the pipeline never crashes due to missing identifiers.

Multi-Skill Threshold

The repository is only classified as multi-skill when two or more valid skill directories are discovered (line 86). The result is encapsulated in a MultiSkillDetectionResult object consumed by the main workflow.

Isolated Analysis Architecture

Once detection completes, SkillSpector leverages LangGraph workflows to maintain strict isolation between skills.

Per-Skill State Management

Each skill executes within its own SkillspectorState instance (defined in src/skillspector/state.py). This isolation prevents cross-contamination of mutable caches—including file_cache and ast_cache—ensuring that findings from one skill do not influence another. When processing a multi-skill repository, the CLI iterates over result.skills and invokes the analysis graph once per skill.

The LangGraph Workflow Pipeline

The compiled workflow in src/skillspector/graph.py (lines 34-53) orchestrates analyzers registered in src/skillspector/nodes/analyzers/__init__.py. For each skill, the graph resolves the input path, builds context, runs static and LLM-backed analyzers, and feeds findings into a meta-analyzer. This per-skill execution generates isolated SARIF reports rather than a single monolithic output.

Parallelism and Resource Control

The architecture supports concurrent execution within each skill via LangGraph's node parallelism. Additionally, the outer loop in src/skillspector/cli.py can process multiple skills simultaneously using thread pools, while selective scanning skips non-skill folders to reduce I/O overhead.

Practical Implementation Example

Below is a complete implementation demonstrating how to programmatically handle large multi-skill repositories with SkillSpector's detection API:

from pathlib import Path
from skillspector.multi_skill import detect_skills
from skillspector.graph import graph  # compiled LangGraph workflow

from skillspector.logging_config import get_logger

logger = get_logger(__name__)

def scan_repository(root: Path):
    """Run SkillSpector on a repository that may contain many skills."""
    detection = detect_skills(root)

    # Single-skill case – run the graph once

    if not detection.is_multi_skill:
        logger.info("Scanning as a single skill")
        state = {"input_path": str(root)}
        result = graph.invoke(state)          # returns the final SkillspectorState

        print(result["report_body"])
        return

    # Multi-skill case – iterate over each discovered skill

    logger.info(f"Detected {len(detection.skills)} sub-skills")
    for skill in detection.skills:
        logger.info(f"Scanning sub-skill: {skill.name}")
        state = {"input_path": str(skill.path)}
        result = graph.invoke(state)
        # Each skill gets its own SARIF fragment; write to disk:

        sarif_path = root / f"{skill.name}.sarif.json"
        sarif_path.write_text(result["sarif_report"], encoding="utf-8")
        print(f"Report for {skill.name} written to {sarif_path}")

Execute the scanner on your local repository:

python - <<'PY'
from pathlib import Path
from my_script import scan_repository

scan_repository(Path("/path/to/large-multi-skill-repo"))
PY

Summary

  • Root Detection: The presence of a top-level SKILL.md triggers single-skill mode, while its absence activates multi-skill scanning in src/skillspector/multi_skill.py.
  • Sub-Directory Filtering: Hidden directories are ignored, and skill candidates are validated via _has_skill_md() checks on immediate children.
  • State Isolation: Each skill runs in its own SkillspectorState instance, preventing memory leaks and cross-contamination across the analysis graph.
  • Scalable Execution: The LangGraph pipeline processes skills independently, generating separate SARIF reports per component while supporting internal node parallelism.
  • Robust Fallbacks: Malformed skill front-matter defaults to directory names, ensuring deterministic execution even with imperfect metadata.

Frequently Asked Questions

How does SkillSpector determine if a repository contains multiple skills?

SkillSpector checks for a root-level SKILL.md file first; if absent, it scans immediate sub-directories for their own SKILL.md files. When two or more valid skill directories are found (line 86 of multi_skill.py), the repository is classified as multi-skill and processed as a collection.

What happens if a skill's SKILL.md file is missing the name field?

The _extract_skill_name() function extracts the name from front-matter (lines 99-129), but if this field is missing or malformed, it falls back to the directory name (line 29). This ensures the analysis pipeline continues without interruption regardless of metadata quality.

How does SkillSpector prevent memory issues when scanning dozens of skills?

Each skill executes within an isolated SkillspectorState object with separate file_cache and ast_cache instances. The CLI processes skills sequentially or via thread pools, but never shares mutable state between them, keeping the memory footprint bounded per skill rather than cumulative across the repository.

Can I customize which analyzers run for each skill in a multi-skill repository?

The analyzer registry in src/skillspector/nodes/analyzers/__init__.py defines the complete set of static and LLM-backed analyzers used by the graph. While the current implementation runs the full suite for each skill, the modular LangGraph architecture in src/skillspector/graph.py supports custom node composition for specialized workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →