SkillSpector Input Types for Scanning: Git, ZIP, URLs, and Local Files Explained
SkillSpector supports six distinct input types—Git repository URLs, direct file URLs, ZIP archives, Markdown files, local directories, and other local files—normalizing all of them into a temporary directory structure via the InputHandler.resolve() method.
NVIDIA's SkillSpector is an open-source scanning engine that analyzes code repositories and files to extract skill-related metadata. Understanding the supported input types is essential for effectively using the tool, as the InputHandler class in src/skillspector/input_handler.py automatically detects and transforms various input formats into a standardized local directory for downstream processing.
How SkillSpector Normalizes Input
The scanning engine requires a concrete filesystem path to operate, but users can supply inputs in multiple formats. The InputHandler class bridges this gap through its resolve() method (lines 93-102), which analyzes the input string and routes it to specialized helper methods. Each helper creates a temporary directory via _get_temp_dir(), ensuring the scanner always receives an absolute path to a local directory regardless of the original source.
Supported SkillSpector Input Types
Git Repository URLs
When the resolve() method detects a Git URL pattern, it invokes _clone_git() to clone the repository into a temporary folder. This allows SkillSpector to scan remote repositories without requiring manual download steps.
Direct File URLs
Raw file URLs from supported hosts (including raw GitHub and HuggingFace links) trigger the _download_file() helper. The file is downloaded to a temporary location, with security checks against allowed hosts defined in src/skillspector/constants.py to prevent SSRF attacks.
ZIP Archives
Both local and remote ZIP archives are handled by _extract_zip(). The archive is extracted into a temporary directory, and the resulting folder path is passed to the scanner. This supports compressed project bundles without requiring manual extraction.
Markdown Files
Individual .md files are processed by _wrap_single_file(), which copies the file into a temporary directory structure. This ensures the scanner receives a consistent directory-based input even when analyzing single documentation files.
Local Directories
If the input resolves to an existing local directory, resolve() returns the absolute path unchanged. This bypasses temporary directory creation and uses the existing filesystem structure directly.
Other Local Files
Any non-Markdown local file is treated similarly to Markdown files—copied into a temporary directory via _wrap_single_file(). This catch-all behavior ensures single-file inputs are always presented to the scanner within a directory context.
Practical Implementation Examples
The following examples demonstrate how to use the InputHandler class to process different SkillSpector input types:
from skillspector.input_handler import InputHandler
handler = InputHandler()
# 1️⃣ Scan a public Git repo
repo_path, src_type = handler.resolve("https://github.com/NVIDIA/SkillSpector.git")
print(repo_path, src_type) # → /tmp/skillspector_abc123/repo git
# 2️⃣ Scan a raw file URL
file_path, src_type = handler.resolve(
"https://raw.githubusercontent.com/NVIDIA/SkillSpector/main/README.md"
)
print(file_path, src_type) # → /tmp/skillspector_abc123 url
# 3️⃣ Scan a local zip archive
zip_path, src_type = handler.resolve("/home/user/project.zip")
print(zip_path, src_type) # → /tmp/skillspector_abc123/extracted zip
# 4️⃣ Scan a single markdown file
md_path, src_type = handler.resolve("/home/user/notes.md")
print(md_path, src_type) # → /tmp/skillspector_abc123 file
# 5️⃣ Scan an entire directory
dir_path, src_type = handler.resolve("/home/user/project/")
print(dir_path, src_type) # → /home/user/project directory
After scanning, clean up temporary resources:
handler.cleanup() # removes the temporary folders created above
Security and Architecture Notes
The input resolution process includes several security considerations. The InputHandler validates URL inputs against allowed host lists defined in src/skillspector/constants.py, preventing Server-Side Request Forgery (SSRF) attacks against internal network resources. All temporary directories are created with unique identifiers to prevent collisions between concurrent scanning operations.
The CLI entry point in src/skillspector/cli.py instantiates InputHandler and manages the lifecycle of the resolved path, ensuring temporary directories are cleaned up after scanning completes or when errors occur.
Summary
- Git URLs: Cloned to temporary directories via
_clone_git() - Direct file URLs: Downloaded securely to temp folders via
_download_file() - ZIP archives: Extracted automatically via
_extract_zip() - Markdown files: Wrapped in temporary directories via
_wrap_single_file() - Local directories: Returned as absolute paths without modification
- Other local files: Treated as single-file inputs and wrapped in temp directories
Key implementation files include src/skillspector/input_handler.py (core logic), src/skillspector/constants.py (security constants), and src/skillspector/cli.py (entry point).
Frequently Asked Questions
Does SkillSpector support private Git repositories?
The InputHandler class uses standard Git clone operations via _clone_git(), so it supports any URL accessible to the Git binary in the environment. However, authentication must be configured at the system level (SSH keys or Git credentials), as the resolve() method does not accept authentication parameters directly.
How does SkillSpector handle large ZIP archives?
Large archives are extracted to temporary directories created by _get_temp_dir() without size validation checks in the current implementation. Ensure your system has sufficient temporary disk space, as the extraction process requires space for both the compressed and uncompressed contents.
Can I scan multiple inputs in a single command?
The current InputHandler implementation processes one input string per resolve() call. To scan multiple inputs, instantiate the handler and call resolve() for each target, or invoke the CLI separately for each input. The cleanup() method removes all temporary directories created by that handler instance.
Where are temporary files stored during scanning?
Temporary directories are created via _get_temp_dir() in the system's default temporary location (typically /tmp on Linux systems). These directories follow the pattern skillspector_abc123 and are automatically removed when cleanup() is called or when the handler instance is destroyed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →