SkillSpector Input Types for Scanning: Git, ZIP, URLs, and Local Files Explained

SkillSpector supports six distinct input types—Git repository URLs, direct file URLs, ZIP archives, Markdown files, local directories, and other local files—normalizing all of them into a temporary directory structure via the InputHandler.resolve() method.

NVIDIA's SkillSpector is an open-source scanning engine that analyzes code repositories and files to extract skill-related metadata. Understanding the supported input types is essential for effectively using the tool, as the InputHandler class in src/skillspector/input_handler.py automatically detects and transforms various input formats into a standardized local directory for downstream processing.

How SkillSpector Normalizes Input

The scanning engine requires a concrete filesystem path to operate, but users can supply inputs in multiple formats. The InputHandler class bridges this gap through its resolve() method (lines 93-102), which analyzes the input string and routes it to specialized helper methods. Each helper creates a temporary directory via _get_temp_dir(), ensuring the scanner always receives an absolute path to a local directory regardless of the original source.

Supported SkillSpector Input Types

Git Repository URLs

When the resolve() method detects a Git URL pattern, it invokes _clone_git() to clone the repository into a temporary folder. This allows SkillSpector to scan remote repositories without requiring manual download steps.

Direct File URLs

Raw file URLs from supported hosts (including raw GitHub and HuggingFace links) trigger the _download_file() helper. The file is downloaded to a temporary location, with security checks against allowed hosts defined in src/skillspector/constants.py to prevent SSRF attacks.

ZIP Archives

Both local and remote ZIP archives are handled by _extract_zip(). The archive is extracted into a temporary directory, and the resulting folder path is passed to the scanner. This supports compressed project bundles without requiring manual extraction.

Markdown Files

Individual .md files are processed by _wrap_single_file(), which copies the file into a temporary directory structure. This ensures the scanner receives a consistent directory-based input even when analyzing single documentation files.

Local Directories

If the input resolves to an existing local directory, resolve() returns the absolute path unchanged. This bypasses temporary directory creation and uses the existing filesystem structure directly.

Other Local Files

Any non-Markdown local file is treated similarly to Markdown files—copied into a temporary directory via _wrap_single_file(). This catch-all behavior ensures single-file inputs are always presented to the scanner within a directory context.

Practical Implementation Examples

The following examples demonstrate how to use the InputHandler class to process different SkillSpector input types:

from skillspector.input_handler import InputHandler

handler = InputHandler()

# 1️⃣ Scan a public Git repo

repo_path, src_type = handler.resolve("https://github.com/NVIDIA/SkillSpector.git")
print(repo_path, src_type)      # → /tmp/skillspector_abc123/repo  git

# 2️⃣ Scan a raw file URL

file_path, src_type = handler.resolve(
    "https://raw.githubusercontent.com/NVIDIA/SkillSpector/main/README.md"
)
print(file_path, src_type)      # → /tmp/skillspector_abc123  url

# 3️⃣ Scan a local zip archive

zip_path, src_type = handler.resolve("/home/user/project.zip")
print(zip_path, src_type)       # → /tmp/skillspector_abc123/extracted  zip

# 4️⃣ Scan a single markdown file

md_path, src_type = handler.resolve("/home/user/notes.md")
print(md_path, src_type)        # → /tmp/skillspector_abc123  file

# 5️⃣ Scan an entire directory

dir_path, src_type = handler.resolve("/home/user/project/")
print(dir_path, src_type)       # → /home/user/project  directory

After scanning, clean up temporary resources:

handler.cleanup()   # removes the temporary folders created above

Security and Architecture Notes

The input resolution process includes several security considerations. The InputHandler validates URL inputs against allowed host lists defined in src/skillspector/constants.py, preventing Server-Side Request Forgery (SSRF) attacks against internal network resources. All temporary directories are created with unique identifiers to prevent collisions between concurrent scanning operations.

The CLI entry point in src/skillspector/cli.py instantiates InputHandler and manages the lifecycle of the resolved path, ensuring temporary directories are cleaned up after scanning completes or when errors occur.

Summary

  • Git URLs: Cloned to temporary directories via _clone_git()
  • Direct file URLs: Downloaded securely to temp folders via _download_file()
  • ZIP archives: Extracted automatically via _extract_zip()
  • Markdown files: Wrapped in temporary directories via _wrap_single_file()
  • Local directories: Returned as absolute paths without modification
  • Other local files: Treated as single-file inputs and wrapped in temp directories

Key implementation files include src/skillspector/input_handler.py (core logic), src/skillspector/constants.py (security constants), and src/skillspector/cli.py (entry point).

Frequently Asked Questions

Does SkillSpector support private Git repositories?

The InputHandler class uses standard Git clone operations via _clone_git(), so it supports any URL accessible to the Git binary in the environment. However, authentication must be configured at the system level (SSH keys or Git credentials), as the resolve() method does not accept authentication parameters directly.

How does SkillSpector handle large ZIP archives?

Large archives are extracted to temporary directories created by _get_temp_dir() without size validation checks in the current implementation. Ensure your system has sufficient temporary disk space, as the extraction process requires space for both the compressed and uncompressed contents.

Can I scan multiple inputs in a single command?

The current InputHandler implementation processes one input string per resolve() call. To scan multiple inputs, instantiate the handler and call resolve() for each target, or invoke the CLI separately for each input. The cleanup() method removes all temporary directories created by that handler instance.

Where are temporary files stored during scanning?

Temporary directories are created via _get_temp_dir() in the system's default temporary location (typically /tmp on Linux systems). These directories follow the pattern skillspector_abc123 and are automatically removed when cleanup() is called or when the handler instance is destroyed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →