# How to Debug LLM Analysis Failures and Provider Issues in SkillSpector: A Complete Guide

> Debug LLM analysis failures and provider issues in SkillSpector. Set log level to DEBUG and validate registry entries/credentials to fix empty findings or timeouts.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: how-to-guide
- Published: 2026-07-09

---

**Set `SKILLSPECTOR_LOG_LEVEL=DEBUG` to surface token estimates and provider errors, then validate your model registry entries and credentials to resolve empty findings or connection timeouts.**

SkillSpector is NVIDIA’s open-source, LLM-driven static analysis framework that batches source files into prompts and parses structured responses to identify security vulnerabilities. When analysis fails silently, returns empty findings, or throws provider exceptions, the root cause typically lies in one of three architectural layers: the analyzer core, the provider credential system, or the token-budget registry. Understanding how these components interact—specifically within [`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py) and the provider modules—is essential for rapid troubleshooting.

## Understanding the SkillSpector Architecture

SkillSpector stitches together three distinct systems to process code:

1. **LLM Analyzer Core** – Located in [`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py), this component estimates token budgets (`get_batches`), constructs numbered prompts (`build_prompt`), and parses structured responses (`parse_response`).
2. **Provider Layer** – Creates concrete LangChain chat models (e.g., `ChatOpenAI`, `ChatBedrockConverse`) via factory methods like `OpenAIProvider.create_chat_model` and `BedrockProvider.create_chat_model`.
3. **Logging Configuration** – Centralized through `logging_config.get_logger`, controlled by the environment variable `SKILLSPECTOR_LOG_LEVEL`.

### Common Failure Symptoms

When analysis breaks, you will typically encounter one of these patterns:

- **Empty or missing findings** – The LLM returned an empty `findings` list or a malformed payload that `parse_response` cannot convert to the `LLMAnalysisResult` schema.
- **Provider exceptions** – Missing API keys or AWS credentials cause `resolve_credentials` to return `None`, triggering a fallback or hard failure.
- **Token-budget errors** – The registry lookup for `context_length` or `max_output_tokens` returns `None`, forcing the analyzer to default to 1024 tokens and over-chunk files.
- **Silent timeouts** – Default `WARNING` log level suppresses debug messages, hiding network timeouts or retry loops.

## Enabling Debug Logging

Before modifying code, raise the log level to expose the analyzer’s internal state. The logger is configured in [`src/skillspector/logging_config.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/logging_config.py) (lines 31-38), where `_configure` reads `SKILLSPECTOR_LOG_LEVEL` on first use.

```bash
export SKILLSPECTOR_LOG_LEVEL=DEBUG
skill-spector analyze /path/to/your/code

```

This setting surfaces:
- Token estimates from `estimate_tokens(prompt)` during batch creation
- Provider-specific connection errors during `create_chat_model`
- Raw prompt sizes and chunking decisions in `get_batches`

## Resolving Provider Credential Issues

Provider failures occur when `resolve_credentials` cannot locate valid authentication material. Each provider implements distinct resolution logic.

### OpenAI Configuration

OpenAI credentials resolve in [`src/skillspector/providers/openai/provider.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/providers/openai/provider.py) (lines 52-58):

```python
from skillspector.providers.openai.provider import OpenAIProvider

provider = OpenAIProvider()
creds = provider.resolve_credentials()  # Returns None if OPENAI_API_KEY missing

print(f"Credentials resolved: {creds is not None}")

```

**Required environment variables:**
- `OPENAI_API_KEY`
- `OPENAI_BASE_URL` (optional)

If `resolve_credentials()` returns `None`, the analyzer may fall back to a No-Op model, resulting in "No content available" messages.

### AWS Bedrock Configuration

Bedrock resolution logic lives in [`src/skillspector/providers/bedrock/provider.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/providers/bedrock/provider.py) (lines 64-84):

```python
from skillspector.providers.bedrock.provider import BedrockProvider

provider = BedrockProvider()
creds = provider.resolve_credentials()  # Returns None if AWS chain fails

```

Ensure your AWS credential chain is active via `aws configure`, or set `AWS_PROFILE` and `AWS_REGION` environment variables. Unlike OpenAI, Bedrock explicitly returns `None` from `resolve_credentials` when the chain fails, which the CLI surfaces as a credential error.

## Fixing Token Budget and Registry Errors

The analyzer queries [`src/skillspector/providers/registry.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/providers/registry.py) (lines 68-81) via `lookup_context_length` and `lookup_max_output_tokens` to determine how large each batch can be. If you specify a model via `SKILLSPECTOR_MODEL` that does not exist in the bundled [`model_registry.yaml`](https://github.com/NVIDIA/SkillSpector/blob/main/model_registry.yaml), these lookups return `None`, forcing a conservative fallback of **1024 tokens**. This causes large files to be over-chunked or truncated.

Verify your model configuration:

```bash
cat src/skillspector/providers/openai/model_registry.yaml

```

Ensure your model appears with valid `context_length` and `max_output_tokens` values. If the model is missing, either add it to the registry or select a supported model to avoid performance degradation.

## Debugging Response Parsing Failures

When the LLM returns JSON that does not conform to the `LLMAnalysisResult` schema, `parse_response` in [`src/skillspector/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/llm_analyzer_base.py) (lines 53-64) raises a `NotImplementedError`. This often happens when the model hallucinates field names or omits required `LLMFinding` attributes.

To diagnose:
1. Temporarily insert `print(response)` at the start of `parse_response` to capture the raw payload
2. Verify the response against the `LLMAnalysisResult` Pydantic schema
3. Check that the model isn't outputting markdown code blocks around the JSON

## Isolating Failures with a Minimal Reproducer

When CLI output is ambiguous, isolate the pipeline in a standalone script:

```python
import os
os.environ["OPENAI_API_KEY"] = "sk-..."
os.environ["SKILLSPECTOR_LOG_LEVEL"] = "DEBUG"

from skillspector.providers.openai.provider import OpenAIProvider
from skillspector.llm_analyzer_base import LLMAnalyzerBase, Batch

provider = OpenAIProvider()
model = provider.resolve_model()

analyzer = LLMAnalyzerBase(
    base_prompt="You are a security analyst. Identify vulnerable patterns.",
    model=model
)

with open("target.py") as f:
    content = f.read()

batch = Batch(file_path="target.py", content=content)
results = analyzer.run_batches([batch])
print(results)  # [(Batch(...), [Finding(...), ...])]

```

This bypasses CLI argument parsing and lets you step through `get_batches`, `build_prompt`, and `parse_response` with a debugger attached.

## Summary

- **Enable debug logging** immediately by setting `SKILLSPECTOR_LOG_LEVEL=DEBUG` to see token estimates, prompt sizes, and provider initialization details.
- **Validate provider credentials** by checking `resolve_credentials()` in `OpenAIProvider` (lines 52-58) or `BedrockProvider` (lines 64-84) before running analysis.
- **Verify model registry entries** to ensure `SKILLSPECTOR_MODEL` exists in [`model_registry.yaml`](https://github.com/NVIDIA/SkillSpector/blob/main/model_registry.yaml) with accurate `context_length` and `max_output_tokens`.
- **Inspect parsing logic** in `parse_response` (lines 53-64 of [`llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/llm_analyzer_base.py)) when `NotImplementedError` indicates schema mismatches.
- **Use isolated scripts** to test the `LLMAnalyzerBase` pipeline without CLI overhead, making it easier to attach debuggers.

## Frequently Asked Questions

### Why does SkillSpector return empty findings for files that contain clear vulnerabilities?

The LLM may return an empty `findings` list due to the "precision-over-recall" guidance embedded in `BASE_ANALYSIS_PROMPT`, which instructs the model to prefer false negatives over false positives. Alternatively, the response may be malformed and fail validation in `parse_response`. Set `SKILLSPECTOR_LOG_LEVEL=DEBUG` to distinguish between these cases by inspecting the raw model output before schema conversion.

### How do I resolve "Could not resolve credentials" errors when using OpenAI?

Ensure `OPENAI_API_KEY` is exported in your environment and detected by `OpenAIProvider.resolve_credentials()` in [`src/skillspector/providers/openai/provider.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/providers/openai/provider.py) (lines 52-58). The method returns `None` when the key is missing, causing the analyzer to fail. You may also set `OPENAI_BASE_URL` if using a proxy or custom endpoint.

### Why are large source files being split into hundreds of tiny chunks?

This occurs when the token-budget lookup fails for your specified model, forcing a fallback to 1024 tokens in `get_batches`. Check that your `SKILLSPECTOR_MODEL` value exists in [`model_registry.yaml`](https://github.com/NVIDIA/SkillSpector/blob/main/model_registry.yaml) with accurate `context_length` entries, or the analyzer in [`src/skillspector/providers/registry.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/providers/registry.py) (lines 68-81) cannot calculate appropriate batch sizes.

### Where does SkillSpector output debug information during analysis?

Diagnostic information flows through a package-wide logger configured in [`src/skillspector/logging_config.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/logging_config.py). The `_configure` function reads `SKILLSPECTOR_LOG_LEVEL` (defaulting to `WARNING`) on first use. Setting this variable to `DEBUG` surfaces internal operations including the `estimate_tokens()` calculations, provider creation attempts, and raw response payloads.