What Are the Main Components of SkillSpector? A Complete Technical Breakdown
SkillSpector's main components include a multi-format input parser, a two-stage analysis engine combining static analysis with optional LLM evaluation, a risk scoring system, baseline suppression module, MCP server interface, and pluggable LLM provider registry, all exposed through both CLI and Python API.
SkillSpector is an open-source security scanner for AI-agent skills developed by NVIDIA. Understanding the main components of SkillSpector reveals how it answers the critical question "Is this skill safe to install?" by analyzing code, metadata, and dependencies before execution.
Input Handling and Ingestion Layer
The input processing component supports multiple ingestion formats to accommodate diverse workflows. According to the NVIDIA/SkillSpector source code, the scanner accepts Git repositories, URLs, zip files, local directories, or single SKILL.md files.
This flexibility allows integration into various stages of the development lifecycle, from local developer testing to automated CI pipeline scanning. The input layer normalizes these disparate formats into a consistent internal representation before passing them to the analysis engine.
Analysis Engine: Static and Semantic Components
The core of SkillSpector is its two-stage analysis pipeline, implemented primarily in src/skillspector/graph.py using LangGraph workflow orchestration.
Static Analysis Core
The first stage performs fast static analysis using multiple techniques:
- Regex pattern matching against 68 vulnerability patterns across 17 categories
- AST-based execution analysis to detect code flow anomalies
- YARA signature matching for known malicious indicators
- Live OSV lookups querying OSV.dev for known CVEs in dependencies (with offline fallback)
These checks detect prompt injection vectors, data exfiltration attempts, privilege escalation paths, supply-chain vulnerabilities, and MCP-specific security issues.
LLM-Based Semantic Layer
The optional second stage uses an LLM evaluation component to reduce false positives and provide semantic context. This layer is controlled via the --no-llm flag in src/skillspector/cli.py and supports multiple providers including OpenAI, Anthropic, Bedrock, NVIDIA Build, Claude CLI, and Codex CLI.
The provider registry in src/skillspector/providers/registry.py handles automatic model discovery and authentication, while src/skillspector/llm_utils.py manages prompt generation and response parsing.
Vulnerability Detection Capabilities
The detection engine identifies 68 distinct vulnerability patterns organized into 17 categories. These include prompt injection, data exfiltration, privilege escalation, supply-chain attacks, AST-based execution risks, and MCP-specific security checks.
The system performs live vulnerability lookups (SC4) by querying OSV.dev for known CVEs affecting dependencies. When operating offline, it falls back to a bundled vulnerability list. This dual-mode operation ensures security coverage regardless of network connectivity.
Risk Assessment and Reporting
SkillSpector's risk scoring component calculates a 0-100 score with severity bands (LOW, MEDIUM, HIGH, CRITICAL) and provides actionable recommendations (SAFE, CAUTION, DO_NOT_INSTALL). The scoring logic is centralized in src/skillspector/constants.py.
The output formatting module supports four distinct formats:
- Terminal: Human-readable colored output
- JSON: Machine-readable for CI integration
- Markdown: Documentation-friendly reports
- SARIF: Standard format for IDE integration
Integration and Automation Features
MCP Server Mode
The MCP server component in src/skillspector/mcp_server.py exposes a Model-Context-Protocol endpoint (scan_skill) that AI agents can call to gate skill installations at runtime. It supports both HTTP and STDIO transports, enabling integration into agent orchestration workflows.
Baseline Suppression
The suppression engine in src/skillspector/suppression.py handles baseline files and glob-rules to hide known findings. This allows teams to suppress false positives and only surface new issues, with fingerprints stored in .skillspector-baseline.yaml files.
CLI and Exit Codes
The command-line interface in src/skillspector/cli.py provides sensible defaults with flags for --no-llm, --baseline, --format, and --yara-rules-dir. It implements a strict exit-code contract: 0 for safe/cautionary scans, 1 for risky scans, and 2 for errors, making it ideal for CI/CD gating.
Core Source Files and Architecture
The main components of SkillSpector are implemented across these key files:
src/skillspector/cli.py: Command-line interface, argument parsing, and pipeline dispatchsrc/skillspector/graph.py: LangGraph workflow orchestrating static analysis, LLM evaluation, scoring, and output formattingsrc/skillspector/models.py: Data models for findings, risk assessment, and report serializationsrc/skillspector/suppression.py: Baseline handling, fingerprinting, and false-positive suppression logicsrc/skillspector/mcp_server.py: MCPscan_skilltool implementation with HTTP/STDIO transportssrc/skillspector/llm_utils.py: Provider-agnostic LLM request handling and prompt managementsrc/skillspector/constants.py: Centralized configuration for pattern IDs, severity mapping, and default scoressrc/skillspector/providers/registry.py: Auto-discovery and registration of LLM providers and model registries
Usage Examples
Scan a local skill with full analysis:
skillspector scan ./my-skill/
Static-only scan for CI environments:
skillspector scan ./my-skill/ --no-llm
Generate JSON output for automation:
skillspector scan ./my-skill/ --format json --output report.json
Run as MCP server:
skillspector mcp --transport http --host 127.0.0.1 --port 8000
Python API usage:
from skillspector import graph
result = graph.invoke({
"input_path": "/path/to/skill",
"output_format": "json",
"use_llm": True,
})
print(f"Score: {result['risk_score']}/100")
print(f"Recommendation: {result['risk_recommendation']}")
Summary
- SkillSpector combines static analysis with optional LLM evaluation to secure AI-agent skills before installation.
- The two-stage engine in
src/skillspector/graph.pyorchestrates 68 vulnerability patterns, OSV lookups, and semantic analysis. - Risk scoring provides 0-100 ratings with severity bands and actionable recommendations.
- Integration components include CLI with standard exit codes, MCP server mode, and Python API for flexible deployment.
- Baseline suppression in
src/skillspector/suppression.pyallows teams to manage false positives and track only new findings.
Frequently Asked Questions
What detection methods does SkillSpector use?
SkillSpector employs regex pattern matching, AST-based code analysis, YARA signatures, and live OSV.dev lookups to detect 68 vulnerability patterns across 17 categories. The optional LLM layer provides semantic analysis to reduce false positives. These methods are coordinated through the LangGraph workflow in src/skillspector/graph.py.
Can SkillSpector run without an LLM API key?
Yes. The --no-llm flag disables the semantic analysis layer, allowing fully offline operation using only static analysis. This mode is recommended for CI pipelines where API keys may not be available, though it may produce more false positives than the two-stage analysis.
How does the MCP server integration work?
The MCP server exposes a scan_skill tool that agents can call via HTTP or STDIO transports. Implemented in src/skillspector/mcp_server.py, it allows AI agents to request security scans before installing skills, implementing runtime security gating. The server returns structured risk scores and recommendations that agents can use to make installation decisions.
What is the purpose of baseline files in SkillSpector?
Baseline files (.skillspector-baseline.yaml) store fingerprints of known findings that teams have reviewed and accepted. The suppression engine in src/skillspector/suppression.py compares current findings against the baseline, hiding previously identified issues and only surfacing new vulnerabilities. This prevents alert fatigue while maintaining security coverage for novel threats.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →