SkillSpector Content Analysis Limitations: 8 Constraints Every Developer Should Know

SkillSpector is a static, non-executing security scanner that cannot analyze encrypted binaries, image-based payloads, non-English source code, or runtime behaviors, and it reports these gaps explicitly via the analysis_completeness field in its JSON output.

NVIDIA's SkillSpector combines regex-based static analysis with optional LLM-driven semantic review to inspect skill source code for security issues. Because the tool deliberately avoids execution and relies on text parsing, it carries specific content analysis limitations that security teams must understand to interpret scan results accurately.

Understanding SkillSpector's Analysis Architecture

SkillSpector operates through three distinct stages that define what content it can and cannot inspect.

Static Analysis Pipeline

The core engine runs 11 regex-based detectors, YARA rules, and AST-based behavioral analyzers. These scanners only operate on plaintext source files that can be loaded into memory. According to the source code in src/skillspector/nodes/analyzers/static_yara.py and related static analyzer modules, any file that cannot be parsed as text is automatically excluded from pattern matching.

Optional LLM Semantic Layer

The semantic analysis stage sends file contents to a configured LLM provider for contextual risk assessment. This layer can be disabled with --no-llm or may fail if the provider is unreachable. When unavailable, the tool loses its false-positive filtering capability, as implemented in src/skillspector/llm_utils.py.

Supply-Chain (SC4) Checks

This stage queries api.osv.dev for known CVEs in declared dependencies. If the machine lacks outbound HTTPS connectivity, the check falls back to a minimal static list, significantly reducing vulnerability coverage.

8 Critical Limitations of SkillSpector Content Analysis

The _build_analysis_completeness function in src/skillspector/nodes/report.py explicitly tracks these eight constraint categories:

Non-English Language Support

Regex and YARA patterns are tuned for English-language constructs. Source code written in other languages or containing non-English variable names and comments are likely to be missed by the static detectors. This limitation is documented in the README at lines 13-20.

Image-Based Payloads

Text embedded within images—such as screenshots of malicious code—is not parsed. The static analysis pipeline has no OCR capability, so pattern matching never occurs against visual content.

Encrypted and Binary Files

Compiled binaries, obfuscated payloads, and encrypted files are treated as "unreadable" by the parsers. These files are skipped entirely rather than decrypted or disassembled.

Runtime Behavior Blindness

Because SkillSpector never executes the target skill, dynamic behaviors remain invisible. Network calls made only at runtime, dynamic code loading, and environment-dependent execution paths cannot be observed.

Network-Dependent Supply-Chain Checks

When the scanning environment cannot reach api.osv.dev, the vulnerability lookup fails back to a tiny static list. This reduces the completeness of dependency vulnerability detection, as noted in the README limitations section.

In-Memory Cache Constraints

Files that exceed memory limits or cannot be read into the in-memory cache are omitted from analysis. The _build_analysis_completeness function records these as "skipped" components in the final report (lines 11-14 of src/skillspector/nodes/report.py).

Optional LLM Meta-Analysis

When --no-llm is passed or the LLM provider is unavailable, the semantic layer is omitted. The report generator adds "LLM meta-analysis was disabled" to the limitations list (lines 15-18 of src/skillspector/nodes/report.py), indicating reduced false-positive filtering.

Post-Filtering Finding Reduction

After LLM analysis, the meta-analyzer may drop findings deemed low-risk. The count of filtered findings is recorded as a limitation in the report, alerting users that some potential issues were removed by automated triage (lines 19-22 of src/skillspector/nodes/report.py).

How Limitations Surface in Reports

SkillSpector surfaces these constraints through the analysis_completeness object in JSON, SARIF, and Markdown outputs. The is_complete boolean flag indicates whether any limitations were encountered.

Example: Complete Analysis

When all stages execute successfully:

from skillspector import graph

result = graph.invoke({
    "input_path": "example_skill/",
    "output_format": "json",
    "use_llm": True,
})

print(result["analysis_completeness"])

Output:

{
  "total_components": 12,
  "scanned_components": 12,
  "coverage_percent": 100.0,
  "llm_analysis": "applied",
  "findings_before_filtering": 5,
  "findings_after_filtering": 4,
  "limitations": null,
  "is_complete": true
}

Example: Disabled LLM Analysis

Running with --no-llm exposes the limitation explicitly:

skill-spector scan ./my_skill --no-llm --output-format json

Result fragment:

{
  "llm_analysis": "skipped",
  "limitations": [
    "LLM meta-analysis was disabled (--no-llm)",
    "2 finding(s) filtered by meta-analyzer or heuristics"
  ],
  "is_complete": false
}

Example: Offline Supply-Chain Fallback

In air-gapped environments without internet access:

skill-spector scan ./my_skill --no-llm

The analysis_completeness field will indicate:


"LLM meta-analysis unavailable: Unable to reach OSV.dev"

This signals that the SC4 stage fell back to the static vulnerability list, reducing coverage.

Summary

  • SkillSpector is strictly static: It cannot execute code, decrypt payloads, or interpret images, creating inherent blind spots for runtime behaviors and binary content.
  • Language limitations: Regex and YARA patterns target English constructs, potentially missing issues in non-English source code.
  • Network dependencies: Supply-chain vulnerability checks require connectivity to api.osv.dev; offline scans use a minimal static list.
  • Explicit reporting: The _build_analysis_completeness function in src/skillspector/nodes/report.py documents every skipped file, disabled LLM, and filtered finding in the final report's limitations array.
  • Configurable trade-offs: Using --no-llm disables semantic analysis, while memory constraints may force component skipping, both of which reduce scan completeness.

Frequently Asked Questions

Can SkillSpector analyze compiled binaries or encrypted source code?

No. SkillSpector treats encrypted files, compiled binaries, and obfuscated payloads as "unreadable" according to the README documentation. These files are skipped entirely and added to the "skipped components" count in the analysis_completeness report, leaving potential threats in binary form undetected.

Why does SkillSpector miss vulnerabilities in non-English code?

The static analysis pipeline uses regex and YARA patterns tuned for English-language constructs. Because the patterns in src/skillspector/nodes/analyzers/static_*.py modules target English keywords and strings, source code written in other languages or using non-English variable names likely evades detection.

What happens when the LLM provider is unreachable?

When the LLM provider is unavailable or when --no-llm is specified, SkillSpector omits the semantic analysis stage. The _build_analysis_completeness function records this as "LLM meta-analysis was disabled" in the limitations list. Without this layer, the tool loses its ability to filter false positives contextually, potentially increasing noise in the results.

Does SkillSpector detect runtime threats like malicious network calls?

No. Because SkillSpector is deliberately non-executing, it cannot observe dynamic behaviors such as network calls, file system modifications, or code executed only at runtime. Any malicious behavior that requires execution to manifest remains invisible to the static analysis engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →