# What Data Does SkillSpector Send to LLM Providers and OSV.dev During Analysis?

> Learn what data SkillSpector sends to LLM providers and OSV.dev during analysis. Discover the limited dependency identifiers and code snippets shared, ensuring your secrets remain private.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: internals
- Published: 2026-07-08

---

**SkillSpector transmits only dependency identifiers (name, version, ecosystem) to OSV.dev and relevant source code snippets with minimal metadata to LLM providers, never sending secrets, environment variables, or full project files.**

Understanding what data leaves your environment during security analysis is critical for compliance and privacy. When NVIDIA SkillSpector scans a software project, it makes targeted external calls to LLM providers (OpenAI, Anthropic, AWS Bedrock, NVIDIA Build) and the OSV.dev vulnerability database. This article examines exactly what SkillSpector sends to these external services based on the source code implementation in the NVIDIA/SkillSpector repository.

## Data Sent to LLM Providers

When SkillSpector requires natural language explanation for a finding, it constructs a chat completion request through LangChain-compatible clients. The transmission occurs only when findings need contextual analysis, not during the initial scan.

### Message Structure and Content

The analyzer builds a JSON payload containing three distinct components:

- **System message**: Defines the model’s role (e.g., "You are a security code reviewer analyzing potential vulnerabilities")
- **User messages**: Contain the **relevant source code fragment** that triggered the finding, typically the specific file or line range from the SARIF model
- **Contextual metadata**: Includes the file path, programming language, and the list of SARIF findings the model should evaluate

### Token Limiting and Truncation

In [`src/skillspector/nodes/analyzers/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/llm_analyzer_base.py), the routine extracts content from the in-memory SARIF representation and truncates it to the provider-specific token limit before transmission. This ensures only the necessary code context is sent, not the entire source file. The provider-specific client instantiation occurs in [`src/skillspector/providers/openai/provider.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/providers/openai/provider.py) via the `create_openai_compatible_chat_model` function, which configures the LangChain `ChatOpenAI` instance.

### Implementation Flow

The payload construction follows this path:

1. The analyzer identifies a finding requiring LLM evaluation
2. [`llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/llm_analyzer_base.py) extracts the file content and formats it into chat messages
3. The LangChain client POSTs the JSON payload to the provider’s chat endpoint

```python

# Conceptual payload structure from llm_analyzer_base.py

{
  "model": "gpt-4",
  "messages": [
    {
      "role": "system",
      "content": "You are a security code reviewer analyzing potential vulnerabilities..."
    },
    {
      "role": "user",
      "content": "File: src/auth.py\nLanguage: python\nFinding: SQL injection\n\nCode:\ndef query(user_input):\n    cursor.execute(f\"SELECT * FROM users WHERE id = {user_input}\")"
    }
  ],
  "temperature": 0.0
}

```

## Data Sent to OSV.dev

For supply chain security analysis, SkillSpector queries the OSV.dev vulnerability database without transmitting any source code or proprietary project files.

### Batch Query Format

The analyzer sends a single POST request to `https://api.osv.dev/v1/querybatch` containing a JSON array of package objects. Each query object is constructed by the `_build_query(name, version, ecosystem)` function in [`osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/osv_client.py) and contains:

- **`package.name`**: The dependency name normalized to lower-case with hyphens
- **`package.ecosystem`**: Either `"PyPI"` or `"npm"`
- **`version`**: The exact version string extracted from lock files

### Vulnerability Details Retrieval

After receiving vulnerability IDs from the batch query, SkillSpector makes individual GET requests to `https://api.osv.dev/v1/vulns/<ID>` to retrieve detailed vulnerability fields including summary, severity scores, and CVE aliases.

### Client Implementation

The [`src/skillspector/nodes/analyzers/osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/osv_client.py) file contains the `query_batch` method that handles the HTTP transmission. The supply chain analyzer in [`src/skillspector/nodes/analyzers/static_patterns_supply_chain.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/static_patterns_supply_chain.py) drives this process by extracting dependency tuples from [`requirements.txt`](https://github.com/NVIDIA/SkillSpector/blob/main/requirements.txt) or [`package.json`](https://github.com/NVIDIA/SkillSpector/blob/main/package.json) and passing them to the OSV client.

```python

# Query structure from _build_query in osv_client.py

{
  "package": {
    "name": "django",
    "ecosystem": "PyPI"
  },
  "version": "3.1.0"
}

```

## Data Flow Overview

SkillSpector follows a strict data minimization approach for external communications:

1. **Dependency Discovery**: The supply chain analyzer parses [`requirements.txt`](https://github.com/NVIDIA/SkillSpector/blob/main/requirements.txt), [`package.json`](https://github.com/NVIDIA/SkillSpector/blob/main/package.json), and other lock files to extract (name, version, ecosystem) tuples
2. **OSV.dev Transmission**: Only the dependency identifiers are sent to the batch query endpoint. No proprietary source code, project files, or internal paths leave the host
3. **LLM Analysis**: When findings require explanation, only the specific code snippet triggering the alert is transmitted, truncated to provider token limits

## Summary

- **LLM Providers**: Receive only relevant code snippets, file paths, language identifiers, and SARIF finding context—never secrets, environment variables, or unrelated project files
- **OSV.dev**: Receives only dependency names (normalized), versions, and ecosystem identifiers ("PyPI" or "npm")
- **Token Safety**: All LLM requests enforce provider-specific token limits via truncation in [`llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/llm_analyzer_base.py)
- **No Source Exfiltration**: Full source trees remain local; only necessary fragments transmit for analysis

## Frequently Asked Questions

### Does SkillSpector send my entire codebase to LLM providers?

No. According to the implementation in [`llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/llm_analyzer_base.py), SkillSpector extracts only the specific file or line range associated with a SARIF finding. The content undergoes truncation to fit provider token limits before transmission, ensuring only relevant fragments leave your environment.

### What dependency information does SkillSpector share with OSV.dev?

SkillSpector sends only the package name (lower-cased and hyphen-ified), ecosystem string ("PyPI" or "npm"), and version number. This data is constructed in [`osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/osv_client.py) by the `_build_query` function and transmitted via the `query_batch` method to the OSV.dev batch endpoint.

### Are API keys or environment variables transmitted to external services?

No. The source code in [`providers/openai/provider.py`](https://github.com/NVIDIA/SkillSpector/blob/main/providers/openai/provider.py) and [`providers/chat_models.py`](https://github.com/NVIDIA/SkillSpector/blob/main/providers/chat_models.py) shows that credentials are used locally to configure the LangChain client. Only the constructed analysis payloads—code snippets for LLMs and dependency identifiers for OSV.dev—are transmitted to external endpoints.

### Can I audit what data SkillSpector sends before it leaves my network?

Yes. All external request construction happens in specific analyzer files. You can inspect [`src/skillspector/nodes/analyzers/llm_analyzer_base.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/llm_analyzer_base.py) for LLM payloads and [`src/skillspector/nodes/analyzers/osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/osv_client.py) for OSV queries. Both files build standard JSON payloads that can be logged or intercepted for audit purposes before the HTTP POST occurs.