How to Debug LLM Analysis Failures and Provider Issues in SkillSpector: A Complete Guide

Set SKILLSPECTOR_LOG_LEVEL=DEBUG to surface token estimates and provider errors, then validate your model registry entries and credentials to resolve empty findings or connection timeouts.

SkillSpector is NVIDIA’s open-source, LLM-driven static analysis framework that batches source files into prompts and parses structured responses to identify security vulnerabilities. When analysis fails silently, returns empty findings, or throws provider exceptions, the root cause typically lies in one of three architectural layers: the analyzer core, the provider credential system, or the token-budget registry. Understanding how these components interact—specifically within src/skillspector/llm_analyzer_base.py and the provider modules—is essential for rapid troubleshooting.

Understanding the SkillSpector Architecture

SkillSpector stitches together three distinct systems to process code:

  1. LLM Analyzer Core – Located in src/skillspector/llm_analyzer_base.py, this component estimates token budgets (get_batches), constructs numbered prompts (build_prompt), and parses structured responses (parse_response).
  2. Provider Layer – Creates concrete LangChain chat models (e.g., ChatOpenAI, ChatBedrockConverse) via factory methods like OpenAIProvider.create_chat_model and BedrockProvider.create_chat_model.
  3. Logging Configuration – Centralized through logging_config.get_logger, controlled by the environment variable SKILLSPECTOR_LOG_LEVEL.

Common Failure Symptoms

When analysis breaks, you will typically encounter one of these patterns:

  • Empty or missing findings – The LLM returned an empty findings list or a malformed payload that parse_response cannot convert to the LLMAnalysisResult schema.
  • Provider exceptions – Missing API keys or AWS credentials cause resolve_credentials to return None, triggering a fallback or hard failure.
  • Token-budget errors – The registry lookup for context_length or max_output_tokens returns None, forcing the analyzer to default to 1024 tokens and over-chunk files.
  • Silent timeouts – Default WARNING log level suppresses debug messages, hiding network timeouts or retry loops.

Enabling Debug Logging

Before modifying code, raise the log level to expose the analyzer’s internal state. The logger is configured in src/skillspector/logging_config.py (lines 31-38), where _configure reads SKILLSPECTOR_LOG_LEVEL on first use.

export SKILLSPECTOR_LOG_LEVEL=DEBUG
skill-spector analyze /path/to/your/code

This setting surfaces:

  • Token estimates from estimate_tokens(prompt) during batch creation
  • Provider-specific connection errors during create_chat_model
  • Raw prompt sizes and chunking decisions in get_batches

Resolving Provider Credential Issues

Provider failures occur when resolve_credentials cannot locate valid authentication material. Each provider implements distinct resolution logic.

OpenAI Configuration

OpenAI credentials resolve in src/skillspector/providers/openai/provider.py (lines 52-58):

from skillspector.providers.openai.provider import OpenAIProvider

provider = OpenAIProvider()
creds = provider.resolve_credentials()  # Returns None if OPENAI_API_KEY missing

print(f"Credentials resolved: {creds is not None}")

Required environment variables:

  • OPENAI_API_KEY
  • OPENAI_BASE_URL (optional)

If resolve_credentials() returns None, the analyzer may fall back to a No-Op model, resulting in "No content available" messages.

AWS Bedrock Configuration

Bedrock resolution logic lives in src/skillspector/providers/bedrock/provider.py (lines 64-84):

from skillspector.providers.bedrock.provider import BedrockProvider

provider = BedrockProvider()
creds = provider.resolve_credentials()  # Returns None if AWS chain fails

Ensure your AWS credential chain is active via aws configure, or set AWS_PROFILE and AWS_REGION environment variables. Unlike OpenAI, Bedrock explicitly returns None from resolve_credentials when the chain fails, which the CLI surfaces as a credential error.

Fixing Token Budget and Registry Errors

The analyzer queries src/skillspector/providers/registry.py (lines 68-81) via lookup_context_length and lookup_max_output_tokens to determine how large each batch can be. If you specify a model via SKILLSPECTOR_MODEL that does not exist in the bundled model_registry.yaml, these lookups return None, forcing a conservative fallback of 1024 tokens. This causes large files to be over-chunked or truncated.

Verify your model configuration:

cat src/skillspector/providers/openai/model_registry.yaml

Ensure your model appears with valid context_length and max_output_tokens values. If the model is missing, either add it to the registry or select a supported model to avoid performance degradation.

Debugging Response Parsing Failures

When the LLM returns JSON that does not conform to the LLMAnalysisResult schema, parse_response in src/skillspector/llm_analyzer_base.py (lines 53-64) raises a NotImplementedError. This often happens when the model hallucinates field names or omits required LLMFinding attributes.

To diagnose:

  1. Temporarily insert print(response) at the start of parse_response to capture the raw payload
  2. Verify the response against the LLMAnalysisResult Pydantic schema
  3. Check that the model isn't outputting markdown code blocks around the JSON

Isolating Failures with a Minimal Reproducer

When CLI output is ambiguous, isolate the pipeline in a standalone script:

import os
os.environ["OPENAI_API_KEY"] = "sk-..."
os.environ["SKILLSPECTOR_LOG_LEVEL"] = "DEBUG"

from skillspector.providers.openai.provider import OpenAIProvider
from skillspector.llm_analyzer_base import LLMAnalyzerBase, Batch

provider = OpenAIProvider()
model = provider.resolve_model()

analyzer = LLMAnalyzerBase(
    base_prompt="You are a security analyst. Identify vulnerable patterns.",
    model=model
)

with open("target.py") as f:
    content = f.read()

batch = Batch(file_path="target.py", content=content)
results = analyzer.run_batches([batch])
print(results)  # [(Batch(...), [Finding(...), ...])]

This bypasses CLI argument parsing and lets you step through get_batches, build_prompt, and parse_response with a debugger attached.

Summary

  • Enable debug logging immediately by setting SKILLSPECTOR_LOG_LEVEL=DEBUG to see token estimates, prompt sizes, and provider initialization details.
  • Validate provider credentials by checking resolve_credentials() in OpenAIProvider (lines 52-58) or BedrockProvider (lines 64-84) before running analysis.
  • Verify model registry entries to ensure SKILLSPECTOR_MODEL exists in model_registry.yaml with accurate context_length and max_output_tokens.
  • Inspect parsing logic in parse_response (lines 53-64 of llm_analyzer_base.py) when NotImplementedError indicates schema mismatches.
  • Use isolated scripts to test the LLMAnalyzerBase pipeline without CLI overhead, making it easier to attach debuggers.

Frequently Asked Questions

Why does SkillSpector return empty findings for files that contain clear vulnerabilities?

The LLM may return an empty findings list due to the "precision-over-recall" guidance embedded in BASE_ANALYSIS_PROMPT, which instructs the model to prefer false negatives over false positives. Alternatively, the response may be malformed and fail validation in parse_response. Set SKILLSPECTOR_LOG_LEVEL=DEBUG to distinguish between these cases by inspecting the raw model output before schema conversion.

How do I resolve "Could not resolve credentials" errors when using OpenAI?

Ensure OPENAI_API_KEY is exported in your environment and detected by OpenAIProvider.resolve_credentials() in src/skillspector/providers/openai/provider.py (lines 52-58). The method returns None when the key is missing, causing the analyzer to fail. You may also set OPENAI_BASE_URL if using a proxy or custom endpoint.

Why are large source files being split into hundreds of tiny chunks?

This occurs when the token-budget lookup fails for your specified model, forcing a fallback to 1024 tokens in get_batches. Check that your SKILLSPECTOR_MODEL value exists in model_registry.yaml with accurate context_length entries, or the analyzer in src/skillspector/providers/registry.py (lines 68-81) cannot calculate appropriate batch sizes.

Where does SkillSpector output debug information during analysis?

Diagnostic information flows through a package-wide logger configured in src/skillspector/logging_config.py. The _configure function reads SKILLSPECTOR_LOG_LEVEL (defaulting to WARNING) on first use. Setting this variable to DEBUG surfaces internal operations including the estimate_tokens() calculations, provider creation attempts, and raw response payloads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →