Comparing Accuracy Between Static‑Only and Full LLM Analysis in SkillSpector

Static‑only analysis delivers deterministic, high‑precision detections with minimal false positives but limited recall, while full LLM analysis leverages context‑aware reasoning to surface subtle threats at the cost of potential hallucinations and lower precision.

SkillSpector, NVIDIA’s security scanning framework for AI skills, implements a dual‑mode analysis pipeline that toggles between deterministic pattern matching and generative LLM‑driven inspection. Understanding the accuracy trade‑offs between static‑only and full LLM modes—controlled via the use_llm flag in SkillspectorState—lets you optimize scan configurations for precision, recall, and operational constraints.

How Static‑Only Analysis Works

Static‑only mode runs a collection of deterministic pattern‑matchers that scan raw file contents without invoking external language models. In src/skillspector/nodes/analyzers/static_patterns_prompt_injection.py and sibling modules, the analyzer executes regular expressions, Unicode‑tag scans, and other hard‑coded heuristics to identify known vulnerability signatures.

Because every finding maps directly to a concrete regex match or pattern hit, this mode offers very high precision with rare false positives. However, it suffers from limited recall: it cannot detect novel threat variants, complex data‑flow bugs, or logic flaws that fall outside the encoded pattern set. When use_llm is set to False (via the --no-llm CLI flag), the graph bypasses all LLM calls and routes execution through the static‑runner node (static_runner.py), returning only pattern‑based findings.

How Full LLM Analysis Works

Full LLM analysis utilizes the LLMAnalyzerBase class defined in src/skillspector/llm_analyzer_base.py to process files through chat models. For each file or chunk, the node constructs a prompt using BASE_ANALYSIS_PROMPT, sends it to the configured model, and parses the structured response into Finding objects.

This mode achieves higher recall because the model reasons about code context, comments, and developer intent, surfacing issues that static regexes never anticipate. The trade‑off is lower precision: the LLM may hallucinate findings, produce overly‑generic results, or misinterpret ambiguous prompts. The meta‑analyzer node (meta_analyzer.py) orchestrates this process by first running static checks, then optionally invoking LLM passes, and merging the result sets.

Controlling the Analysis Mode

The switch between modes is governed by the use_llm boolean stored in the graph state. In src/skillspector/state.py (line 80), the SkillspectorState dataclass defines this flag, which defaults to True but can be overridden via the command line:


# Static‑only scan (no API key required)

skill-spector scan /path/to/skill --no-llm -o report.json

# Full LLM scan (requires OPENAI_API_KEY or equivalent)

export OPENAI_API_KEY=sk-....
skill-spector scan /path/to/skill -o report.json

When use_llm is True but API calls fail (missing keys, rate limits, or timeouts), the reporting node in src/skillspector/nodes/report.py tracks these failures and injects an _llm_degradation_notice into the final SARIF/JSON output, alerting you that results may be incomplete.

Accuracy Characteristics in Practice

The accuracy gap between modes is observable in SkillSpector’s test suite:

  • test_semantic_quality_policy.py verifies that with use_llm=False, a clean file returns zero findings, while with use_llm=True, the LLM may still return findings based on semantic interpretation.
  • test_meta_analyzer_use_llm.py demonstrates that missing API credentials raise a ValueError when use_llm=True, but the scan succeeds (falling back to static‑only results) when use_llm=False.

Static‑only provides a baseline guarantee: deterministic execution, no hallucinations, fast runtime, and zero external dependencies. Full LLM adds context‑aware coverage capable of detecting complex injection vectors and logic bugs, but requires API credentials and manual triage to filter false positives.

Running Both Modes Programmatically

You can invoke either mode directly via Python:


# Full LLM analysis (single file)

from skillspector.llm_analyzer_base import LLMAnalyzerBase

prompt = "Detect insecure patterns in the following Python file."
llm = LLMAnalyzerBase(base_prompt=prompt, model="gpt-4")
batch = llm.get_batches(
    file_paths=["example.py"],
    file_cache={"example.py": open("example.py").read()},
)
results = llm.run_batches(batch)
findings = llm.collect_findings(results)  # list[Finding]

print(f"LLM discovered {len(findings)} issues")

# Static‑only analysis (no LLM)

from skillspector.nodes.analyzers import static_patterns_prompt_injection as spi
from skillspector.state import SkillspectorState

state = SkillspectorState(
    skill_path="/tmp/skill",
    file_cache={"README.md": "# Example"},

    use_llm=False,
)
static_findings = spi.node(state)["findings"]
print(f"Static scan found {len(static_findings)} issues")

Summary

  • Static‑only analysis in SkillSpector offers high precision and zero false positives through deterministic regex and pattern matching in modules like static_patterns_prompt_injection.py, but cannot detect novel or context‑dependent threats.
  • Full LLM analysis via LLMAnalyzerBase increases recall by reasoning about code semantics and developer intent, though it introduces potential hallucinations and requires external API credentials.
  • The use_llm flag in SkillspectorState (defined in state.py) and the --no-llm CLI option provide granular control over which accuracy profile to apply per scan.
  • When LLM calls fail during a requested full scan, report.py emits a degradation notice so you know results have fallen back to static‑only coverage.

Frequently Asked Questions

When should I choose static‑only analysis over full LLM analysis?

Choose static‑only when you need fast, deterministic scans without API dependencies or when operating in air‑gapped environments. Use it for CI pipelines where false positives break builds and you only need to catch known vulnerability signatures. Choose full LLM when auditing unfamiliar codebases, reviewing complex data‑flow scenarios, or hunting for novel attack patterns that lack existing regex signatures.

How does SkillSpector handle LLM failures during a scan?

When use_llm=True but API calls fail (network errors, invalid keys, or rate limits), the meta‑analyzer (meta_analyzer.py) continues execution but the report node (report.py) appends an _llm_degradation_notice to the output. This notice flags that the scan completed without LLM coverage, ensuring you are not unknowingly viewing incomplete results.

Can I combine static and LLM analysis in the same scan?

Yes. By default, use_llm=True runs both modes sequentially: the meta‑analyzer first executes static pattern checks, then submits files to the LLM, and merges the findings. You receive the union of deterministic detections and model‑generated insights, with the static results serving as a reliable baseline even if the LLM produces noisy output.

What types of vulnerabilities can only the LLM detect?

The LLM excels at identifying semantic vulnerabilities such as subtle prompt injection vectors hidden in natural‑language comments, logic flaws requiring cross‑file context, and insecure defaults that do not match known regex patterns. Static analysis cannot detect these because they require reasoning about intent and context rather than matching predetermined signatures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →