# SkillSpector Input Types for Scanning: Git, ZIP, URLs, and Local Files Explained

> Explore SkillSpector input types: Git, ZIP, URLs, and local files. Learn how NVIDIA/SkillSpector normalizes diverse inputs for effective code scanning.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: how-to-guide
- Published: 2026-06-25

---

**SkillSpector supports six distinct input types—Git repository URLs, direct file URLs, ZIP archives, Markdown files, local directories, and other local files—normalizing all of them into a temporary directory structure via the `InputHandler.resolve()` method.**

NVIDIA's SkillSpector is an open-source scanning engine that analyzes code repositories and files to extract skill-related metadata. Understanding the supported input types is essential for effectively using the tool, as the `InputHandler` class in [`src/skillspector/input_handler.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/input_handler.py) automatically detects and transforms various input formats into a standardized local directory for downstream processing.

## How SkillSpector Normalizes Input

The scanning engine requires a concrete filesystem path to operate, but users can supply inputs in multiple formats. The `InputHandler` class bridges this gap through its `resolve()` method (lines 93-102), which analyzes the input string and routes it to specialized helper methods. Each helper creates a temporary directory via `_get_temp_dir()`, ensuring the scanner always receives an absolute path to a local directory regardless of the original source.

## Supported SkillSpector Input Types

### Git Repository URLs

When the `resolve()` method detects a Git URL pattern, it invokes `_clone_git()` to clone the repository into a temporary folder. This allows SkillSpector to scan remote repositories without requiring manual download steps.

### Direct File URLs

Raw file URLs from supported hosts (including raw GitHub and HuggingFace links) trigger the `_download_file()` helper. The file is downloaded to a temporary location, with security checks against allowed hosts defined in [`src/skillspector/constants.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/constants.py) to prevent SSRF attacks.

### ZIP Archives

Both local and remote ZIP archives are handled by `_extract_zip()`. The archive is extracted into a temporary directory, and the resulting folder path is passed to the scanner. This supports compressed project bundles without requiring manual extraction.

### Markdown Files

Individual `.md` files are processed by `_wrap_single_file()`, which copies the file into a temporary directory structure. This ensures the scanner receives a consistent directory-based input even when analyzing single documentation files.

### Local Directories

If the input resolves to an existing local directory, `resolve()` returns the absolute path unchanged. This bypasses temporary directory creation and uses the existing filesystem structure directly.

### Other Local Files

Any non-Markdown local file is treated similarly to Markdown files—copied into a temporary directory via `_wrap_single_file()`. This catch-all behavior ensures single-file inputs are always presented to the scanner within a directory context.

## Practical Implementation Examples

The following examples demonstrate how to use the `InputHandler` class to process different SkillSpector input types:

```python
from skillspector.input_handler import InputHandler

handler = InputHandler()

# 1️⃣ Scan a public Git repo

repo_path, src_type = handler.resolve("https://github.com/NVIDIA/SkillSpector.git")
print(repo_path, src_type)      # → /tmp/skillspector_abc123/repo  git

# 2️⃣ Scan a raw file URL

file_path, src_type = handler.resolve(
    "https://raw.githubusercontent.com/NVIDIA/SkillSpector/main/README.md"
)
print(file_path, src_type)      # → /tmp/skillspector_abc123  url

# 3️⃣ Scan a local zip archive

zip_path, src_type = handler.resolve("/home/user/project.zip")
print(zip_path, src_type)       # → /tmp/skillspector_abc123/extracted  zip

# 4️⃣ Scan a single markdown file

md_path, src_type = handler.resolve("/home/user/notes.md")
print(md_path, src_type)        # → /tmp/skillspector_abc123  file

# 5️⃣ Scan an entire directory

dir_path, src_type = handler.resolve("/home/user/project/")
print(dir_path, src_type)       # → /home/user/project  directory

```

After scanning, clean up temporary resources:

```python
handler.cleanup()   # removes the temporary folders created above

```

## Security and Architecture Notes

The input resolution process includes several security considerations. The `InputHandler` validates URL inputs against allowed host lists defined in [`src/skillspector/constants.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/constants.py), preventing Server-Side Request Forgery (SSRF) attacks against internal network resources. All temporary directories are created with unique identifiers to prevent collisions between concurrent scanning operations.

The CLI entry point in [`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py) instantiates `InputHandler` and manages the lifecycle of the resolved path, ensuring temporary directories are cleaned up after scanning completes or when errors occur.

## Summary

- **Git URLs**: Cloned to temporary directories via `_clone_git()`
- **Direct file URLs**: Downloaded securely to temp folders via `_download_file()`
- **ZIP archives**: Extracted automatically via `_extract_zip()`
- **Markdown files**: Wrapped in temporary directories via `_wrap_single_file()`
- **Local directories**: Returned as absolute paths without modification
- **Other local files**: Treated as single-file inputs and wrapped in temp directories

Key implementation files include [`src/skillspector/input_handler.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/input_handler.py) (core logic), [`src/skillspector/constants.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/constants.py) (security constants), and [`src/skillspector/cli.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/cli.py) (entry point).

## Frequently Asked Questions

### Does SkillSpector support private Git repositories?

The `InputHandler` class uses standard Git clone operations via `_clone_git()`, so it supports any URL accessible to the Git binary in the environment. However, authentication must be configured at the system level (SSH keys or Git credentials), as the `resolve()` method does not accept authentication parameters directly.

### How does SkillSpector handle large ZIP archives?

Large archives are extracted to temporary directories created by `_get_temp_dir()` without size validation checks in the current implementation. Ensure your system has sufficient temporary disk space, as the extraction process requires space for both the compressed and uncompressed contents.

### Can I scan multiple inputs in a single command?

The current `InputHandler` implementation processes one input string per `resolve()` call. To scan multiple inputs, instantiate the handler and call `resolve()` for each target, or invoke the CLI separately for each input. The `cleanup()` method removes all temporary directories created by that handler instance.

### Where are temporary files stored during scanning?

Temporary directories are created via `_get_temp_dir()` in the system's default temporary location (typically `/tmp` on Linux systems). These directories follow the pattern `skillspector_abc123` and are automatically removed when `cleanup()` is called or when the handler instance is destroyed.