# How the Semantic Developer Intent Analyzer Works in NVIDIA SkillSpector

> Discover how the semantic developer intent analyzer in NVIDIA SkillSpector works. It uses LLMs to find intent violations by comparing skill manifests to code behavior, identifying four categories of issues.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: how-to-guide
- Published: 2026-06-24

---

**The Semantic Developer Intent Analyzer is an LLM-driven validation node that detects mismatches between a skill's declared manifest metadata and its actual code behavior, identifying four specific categories of intent violations (SDI-1 through SDI-4) by cross-referencing declarations against source file contents.**

The semantic developer intent analyzer serves as a critical integrity layer within the NVIDIA SkillSpector framework, ensuring that AI skill manifests accurately represent implementation reality. This component leverages large language models to validate that declared attributes—such as permissions, triggers, and descriptions—align with the actual source code to identify deceptive or inconsistent patterns. According to the NVIDIA/SkillSpector source code, the analyzer operates as a registered node that processes batched source files through a sophisticated pipeline to generate actionable findings.

## Entry Point and State Requirements

The analyzer exposes its functionality through the **`node(state)`** function defined in [`src/skillspector/nodes/analyzers/semantic_developer_intent.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/semantic_developer_intent.py) (lines 57-62). This entry point expects a `SkillspectorState` dictionary containing four critical keys:

- **`use_llm`**: Boolean flag enabling LLM analysis (defaults to `True`)
- **`file_cache`**: Mapping of file paths to source code content for every file in the skill
- **`manifest`**: Parsed skill manifest containing name, description, triggers, and permissions
- **`model_config`**: Optional per-analyzer model overrides

If any required keys are missing, the node immediately returns an empty findings list, ensuring graceful degradation without pipeline failure.

## Model Selection Cascade

Before processing begins, the analyzer determines which LLM to invoke through a prioritized cascade implemented in lines 66-73:

```python
model = (
    model_config.get(ANALYZER_ID)          # explicit per-analyzer override

    or model_config.get("default")         # user-provided default

    or MODEL_CONFIG.get(ANALYZER_ID)       # project-wide defaults

    or _SKILLSPECTOR_DEFAULT_MODEL         # fallback

)

```

This four-tier selection mechanism allows granular control over model choice, from specific analyzer overrides to global fallbacks.

## Prompt Construction and Manifest Formatting

The analyzer builds its LLM prompt in two stages. First, **`_format_manifest(manifest)`** (lines 35-55) transforms the manifest dictionary into a human-readable block containing the skill's name, description, and permissions. If the manifest is empty, the function substitutes a placeholder string to maintain prompt structure.

Second, the static **`ANALYZER_PROMPT`** template (lines 35-124) incorporates the formatted manifest via the `{manifest_section}` placeholder at line 176. This template defines the four SDI rule categories that the LLM must evaluate against the source code to identify intent mismatches.

## LLM Batch Processing Pipeline

The analyzer delegates heavy processing to **`LLMAnalyzerBase`** from [`skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/skillspector/llm_analyzer_base.py), implementing a sophisticated batching strategy to handle large codebases:

**Batch Creation**: The `get_batches()` method (lines 90-135) partitions source files into processing batches. For each file, it evaluates whether the entire content fits within the model's token budget; oversized files are automatically chunked by line numbers using `number_lines()` (lines 46-52).

**Prompt Assembly**: The `build_prompt()` method (lines 39-45) injects the base analyzer prompt and numbered file content into **`BASE_ANALYSIS_PROMPT`** (lines 27-45), producing structured input that preserves line number references for accurate finding attribution.

**Execution and Parsing**: The `run_batches()` method (lines 81-95) iterates over batches, invoking the configured chat model via `get_chat_model()` and parsing responses through `parse_response()`. The analyzer expects JSON-compatible output conforming to the **`LLMAnalysisResult`** schema, converting each `LLMFinding` into a generic `Finding` object (lines 60-62).

**Aggregation**: Finally, `collect_findings()` (line 81) merges per-batch results into a unified findings list attached to the state.

## Error Handling and Registry Integration

Robustness is ensured through exception handling at lines 84-87, where any LLM call failures trigger a warning log and return an empty findings list rather than crashing the pipeline.

The analyzer integrates into SkillSpector's modular architecture through registration in [`src/skillspector/nodes/analyzers/__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/__init__.py). Here, the identifier `"semantic_developer_intent"` maps to the `node` function and appears in **`ANALYZER_IDS`** (line 102), enabling the graph builder to instantiate the node automatically when users request semantic-developer-intent analysis.

## Practical Implementation Examples

### Direct Python Invocation

For programmatic access, import the node directly and populate the required state:

```python
from skillspector.nodes.analyzers.semantic_developer_intent import node
from skillspector.state import SkillspectorState

state: SkillspectorState = {
    "use_llm": True,
    "manifest": {"name": "Example Skill", "description": "...", "permissions": ["read"]},
    "file_cache": {"src/main.py": "import os\n...", "src/utils.py": "..."},
    "model_config": {"semantic_developer_intent": "gpt-4o-mini"},
}

result = node(state)
print(result["findings"])  # List of SDI-1 through SDI-4 findings

```

### Integration with SkillSpector Pipeline

When using the full framework, the analyzer runs automatically via the registry:

```python
from skillspector.graph import SkillGraph
from skillspector.input_handler import load_skill

skill_path = "examples/hello_world_skill"
graph = SkillGraph()
state = load_skill(skill_path)

graph.run(state)
print(state["findings"]["semantic_developer_intent"])

```

The `load_skill` function populates the state with manifest and file cache data, while `SkillGraph` orchestrates the analyzer via the registry defined in [`__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/__init__.py).

## Summary

- The **Semantic Developer Intent Analyzer** validates skill manifest accuracy against implementation code using LLM-driven analysis of four specific mismatch categories (SDI-1 through SDI-4).
- Entry occurs through the **`node(state)`** function in [`semantic_developer_intent.py`](https://github.com/NVIDIA/SkillSpector/blob/main/semantic_developer_intent.py) (lines 57-62), which requires specific state keys including `file_cache` and `manifest`.
- **Model selection** follows a four-tier cascade (lines 66-73) from analyzer-specific overrides to global defaults.
- **Prompt construction** combines formatted manifest data (lines 35-55) with static templates defining SDI evaluation criteria (lines 35-124).
- **Batch processing** via `LLMAnalyzerBase` handles token limits through intelligent file chunking (`get_batches()`, lines 90-135) and line numbering (lines 46-52).
- The analyzer is **registered** in [`__init__.py`](https://github.com/NVIDIA/SkillSpector/blob/main/__init__.py) (line 102) for automatic pipeline integration.
- **Error handling** (lines 84-87) ensures pipeline continuity by returning empty findings on LLM failures rather than raising exceptions.

## Frequently Asked Questions

### What are the four SDI categories the analyzer detects?

The analyzer identifies four specific semantic developer intent mismatches labeled SDI-1 through SDI-4. These categories evaluate whether the skill's manifest accurately describes its code behavior regarding permissions, triggers, data handling, and functional capabilities. The specific definitions reside in the `ANALYZER_PROMPT` template (lines 35-124) that instructs the LLM evaluation criteria.

### How does the analyzer handle large source files that exceed token limits?

The `LLMAnalyzerBase.get_batches()` method (lines 90-135) automatically partitions files based on token budgets. Files exceeding the limit are chunked by line numbers using `number_lines()` (lines 46-52), ensuring the LLM receives manageable batches while preserving line number references for accurate finding attribution in the final report.

### Can I use a different LLM model specifically for the semantic developer intent analysis?

Yes, the model selection cascade (lines 66-73) checks for an analyzer-specific override first. Pass a `model_config` dictionary with the key `"semantic_developer_intent"` set to your preferred model identifier (e.g., `"gpt-4o-mini"` or `"claude-3-opus"`) in the state object to override defaults for this specific analyzer.

### What happens if the LLM service is unavailable during analysis?

The analyzer implements defensive error handling at lines 84-87. If the LLM call raises an exception, the node logs a warning and returns an empty findings list. This design ensures the SkillSpector pipeline continues processing other analyzers rather than failing entirely, maintaining workflow continuity.