# LLM Semantic Analysis in SkillSpector: How AI Detects Security Issues Beyond Static Code

> Discover how LLM semantic analysis in NVIDIA SkillSpector uncovers security risks like prompt injection and intent mismatches that static code analysis misses. Enhance your security.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: deep-dive
- Published: 2026-07-12

---

**SkillSpector uses Large Language Model (LLM) driven semantic analyzers to interpret the meaning and intent of skill manifests, code, and documentation, detecting security risks like prompt injection and intent mismatches that traditional static analysis cannot catch.**

NVIDIA's SkillSpector employs **LLM semantic analysis** to identify nuanced security vulnerabilities in AI skills. Unlike traditional scanners that rely on pattern matching, these analyzers understand natural language context to uncover hidden risks. This article explores the three semantic analyzers and their implementation based on the actual source code in the NVIDIA/SkillSpector repository.

## The Three Semantic Analyzers

SkillSpector instantiates three specialized analyzers that share a common foundation but target distinct security domains. Each analyzer is implemented as a separate node in the analysis graph.

### semantic_security_discovery (SSD)

The **semantic_security_discovery** analyzer, located in [`src/skillspector/nodes/analyzers/semantic_security_discovery.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/semantic_security_discovery.py), detects intent- and attack-phrasing risks that rely on natural-language semantics rather than literal keywords. According to lines 30-68 of the source file, it identifies four primary categories:

- **Prompt-injection patterns** buried in natural language
- **Paraphrased attacks** that evade keyword filters
- **Covert exfiltration instructions** disguised as legitimate functionality
- **Narrative deception** that misleads users about actual behavior

### semantic_developer_intent (SDI)

The **semantic_developer_intent** analyzer validates consistency between what a skill claims to do and what the code actually implements. Defined in [`src/skillspector/nodes/analyzers/semantic_developer_intent.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/semantic_developer_intent.py) (lines 35-53), it checks for:

- **Description-behavior mismatch** between manifest claims and implementation
- **Unjustified capabilities** that exceed declared functionality
- **Scope-creep** beyond stated permissions
- **Contradictory comments and docstrings** that misrepresent code behavior

### semantic_quality_policy (SQP)

The **semantic_quality_policy** analyzer audits overall quality and policy compliance from a natural-language perspective. Implemented in [`src/skillspector/nodes/analyzers/semantic_quality_policy.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/semantic_quality_policy.py) (lines 35-54), it flags:

- **Vague triggers** that might activate unexpectedly
- **Missing user warnings** for dangerous operations
- **Policy-violation language** that static scanners miss

## How LLM Semantic Analysis Works

All three analyzers share a common implementation foundation built around the **LLMAnalyzerBase** class.

### The LLMAnalyzerBase Foundation

The `LLMAnalyzerBase` class in [`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py) provides the core infrastructure for semantic analysis. This reusable helper builds prompts, batches files, calls the selected LLM, and parses the model's JSON-formatted findings. Each analyzer instantiates this base class with a specific system prompt and model configuration.

### Prompt Construction

Each analyzer defines a system prompt (`ANALYZER_PROMPT`) describing the categories to detect and the required output format. The prompt is sent to the LLM together with the file contents. For example, the semantic_security_discovery prompt (lines 30-68 in its source file) instructs the model to look for specific attack patterns while the semantic_developer_intent prompt (lines 35-53) focuses on manifest-code consistency.

### Model Selection and Configuration

The analyzer resolves which LLM model to use via the **model-config** slot system. As defined in [`src/skillspector/constants.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/constants.py) (lines 41-51), the system checks the `MODEL_CONFIG` dictionary for analyzer-specific overrides. If no specific model is configured, it falls back to `_SKILLSPECTOR_DEFAULT_MODEL`.

### Execution Flow

The execution pattern is consistent across all three analyzers. The following logic appears in the `node` functions (see lines 71-93 in [`semantic_security_discovery.py`](https://github.com/NVIDIA/SkillSpector/blob/main/semantic_security_discovery.py), lines 57-82 in [`semantic_developer_intent.py`](https://github.com/NVIDIA/SkillSpector/blob/main/semantic_developer_intent.py), and lines 30-51 in [`semantic_quality_policy.py`](https://github.com/NVIDIA/SkillSpector/blob/main/semantic_quality_policy.py)):

```python
if not state.get("use_llm", True):
    return {"findings": []}

# Gather file cache & model

analyzer = LLMAnalyzerBase(base_prompt=ANALYZER_PROMPT, model=model)
batches = analyzer.get_batches(files, file_cache)
results = asyncio.run(analyzer.arun_batches(batches))
findings = analyzer.collect_findings(results)

```

### Result Integration

Findings are injected into the **SkillspectorState** graph and logged via `llm_call_record`. The [`src/skillspector/state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/state.py) file defines the state structure and `AnalyzerNodeResponse` type used throughout the system. The meta-analyzer in [`src/skillspector/nodes/meta_analyzer.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/meta_analyzer.py) later aggregates these LLM calls to present a unified report alongside static and behavioral analysis results.

## Implementation Examples

### Running Semantic Analysis via CLI

Scan a skill directory with all semantic analyzers enabled using the default configuration:

```bash

# Scan a skill with all semantic analyzers enabled (default)

skillspector scan path/to/skill_dir

```

### Programmatic Use of Semantic Analyzers

You can invoke individual analyzers programmatically by constructing a proper `SkillspectorState`:

```python
from skillspector.state import SkillspectorState
from skillspector.nodes.analyzers.semantic_security_discovery import node

# Prepare a minimal state (file_cache holds file paths → contents)

state = SkillspectorState(
    file_cache={"skill.py": open("skill.py").read()},
    use_llm=True,
    model_config={"semantic_security_discovery": "nvidia/openai/gpt-oss-120b"},
)

result = node(state)
print(result["findings"])   # → list of SSD findings (if any)

```

### Creating a Custom Semantic Analyzer

Extend the system by creating new analyzers using `LLMAnalyzerBase`:

```python
from skillspector.llm_analyzer_base import LLMAnalyzerBase
from skillspector.logging_config import get_logger

ANALYZER_ID = "my_semantic_check"
PROMPT = """You are a security auditor ... (describe categories)"""

def node(state):
    logger = get_logger(__name__)
    model = state.get("model_config", {}).get(ANALYZER_ID, "default-model")
    analyzer = LLMAnalyzerBase(base_prompt=PROMPT, model=model)
    batches = analyzer.get_batches(state.get("components"), state.get("file_cache"))
    results = analyzer.run_batches(batches)
    findings = analyzer.collect_findings(results)
    logger.info("%s: %d findings", ANALYZER_ID, len(findings))
    return {"findings": findings}

```

## Summary

- **LLM semantic analysis** in SkillSpector interprets natural language meaning to catch security issues that static scanners miss.
- Three specialized analyzers target distinct risks: **semantic_security_discovery** (attack patterns), **semantic_developer_intent** (intent mismatches), and **semantic_quality_policy** (compliance issues).
- The `LLMAnalyzerBase` class in [`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py) provides shared infrastructure for prompt construction, batching, and result parsing.
- Model selection uses the `MODEL_CONFIG` slot system with fallback to `_SKILLSPECTOR_DEFAULT_MODEL` defined in [`src/skillspector/constants.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/constants.py).
- Results integrate into `SkillspectorState` via `llm_call_record` and aggregate through the meta-analyzer for unified reporting.

## Frequently Asked Questions

### What is the difference between LLM semantic analysis and static analysis in SkillSpector?

**Static analysis** relies on pattern matching and literal keyword detection to find known vulnerabilities, while **LLM semantic analysis** interprets the meaning and intent of code and documentation. According to the SkillSpector source code, semantic analyzers can detect paraphrased attacks, narrative deception, and description-behavior mismatches that would evade traditional static scanners.

### How does SkillSpector configure which LLM model to use for semantic analysis?

SkillSpector uses a model-config slot system defined in [`src/skillspector/constants.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/constants.py). Each analyzer checks the `MODEL_CONFIG` dictionary for a specific entry (e.g., `"semantic_security_discovery"`), and if not found, falls back to `_SKILLSPECTOR_DEFAULT_MODEL`. You can override models per analyzer or set a global default through the state configuration.

### What types of security issues does semantic_security_discovery detect?

The **semantic_security_discovery** analyzer detects four categories of semantic risks: prompt-injection patterns hidden in natural language, paraphrased attacks that bypass keyword filters, covert exfiltration instructions disguised as legitimate functionality, and narrative deception that misrepresents actual behavior. These require understanding context and intent rather than matching attack signatures.

### Can I create custom semantic analyzers in SkillSpector?

Yes. You can create custom analyzers by importing `LLMAnalyzerBase` from [`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py) and defining a node function that instantiates the base class with a custom system prompt. Your analyzer should follow the established pattern: check `use_llm` state, batch files via `get_batches()`, run `arun_batches()` asynchronously, and return findings through `collect_findings()`.