What Data Does SkillSpector Send to LLM Providers and OSV.dev During Analysis?
SkillSpector transmits only dependency identifiers (name, version, ecosystem) to OSV.dev and relevant source code snippets with minimal metadata to LLM providers, never sending secrets, environment variables, or full project files.
Understanding what data leaves your environment during security analysis is critical for compliance and privacy. When NVIDIA SkillSpector scans a software project, it makes targeted external calls to LLM providers (OpenAI, Anthropic, AWS Bedrock, NVIDIA Build) and the OSV.dev vulnerability database. This article examines exactly what SkillSpector sends to these external services based on the source code implementation in the NVIDIA/SkillSpector repository.
Data Sent to LLM Providers
When SkillSpector requires natural language explanation for a finding, it constructs a chat completion request through LangChain-compatible clients. The transmission occurs only when findings need contextual analysis, not during the initial scan.
Message Structure and Content
The analyzer builds a JSON payload containing three distinct components:
- System message: Defines the model’s role (e.g., "You are a security code reviewer analyzing potential vulnerabilities")
- User messages: Contain the relevant source code fragment that triggered the finding, typically the specific file or line range from the SARIF model
- Contextual metadata: Includes the file path, programming language, and the list of SARIF findings the model should evaluate
Token Limiting and Truncation
In src/skillspector/nodes/analyzers/llm_analyzer_base.py, the routine extracts content from the in-memory SARIF representation and truncates it to the provider-specific token limit before transmission. This ensures only the necessary code context is sent, not the entire source file. The provider-specific client instantiation occurs in src/skillspector/providers/openai/provider.py via the create_openai_compatible_chat_model function, which configures the LangChain ChatOpenAI instance.
Implementation Flow
The payload construction follows this path:
- The analyzer identifies a finding requiring LLM evaluation
llm_analyzer_base.pyextracts the file content and formats it into chat messages- The LangChain client POSTs the JSON payload to the provider’s chat endpoint
# Conceptual payload structure from llm_analyzer_base.py
{
"model": "gpt-4",
"messages": [
{
"role": "system",
"content": "You are a security code reviewer analyzing potential vulnerabilities..."
},
{
"role": "user",
"content": "File: src/auth.py\nLanguage: python\nFinding: SQL injection\n\nCode:\ndef query(user_input):\n cursor.execute(f\"SELECT * FROM users WHERE id = {user_input}\")"
}
],
"temperature": 0.0
}
Data Sent to OSV.dev
For supply chain security analysis, SkillSpector queries the OSV.dev vulnerability database without transmitting any source code or proprietary project files.
Batch Query Format
The analyzer sends a single POST request to https://api.osv.dev/v1/querybatch containing a JSON array of package objects. Each query object is constructed by the _build_query(name, version, ecosystem) function in osv_client.py and contains:
package.name: The dependency name normalized to lower-case with hyphenspackage.ecosystem: Either"PyPI"or"npm"version: The exact version string extracted from lock files
Vulnerability Details Retrieval
After receiving vulnerability IDs from the batch query, SkillSpector makes individual GET requests to https://api.osv.dev/v1/vulns/<ID> to retrieve detailed vulnerability fields including summary, severity scores, and CVE aliases.
Client Implementation
The src/skillspector/nodes/analyzers/osv_client.py file contains the query_batch method that handles the HTTP transmission. The supply chain analyzer in src/skillspector/nodes/analyzers/static_patterns_supply_chain.py drives this process by extracting dependency tuples from requirements.txt or package.json and passing them to the OSV client.
# Query structure from _build_query in osv_client.py
{
"package": {
"name": "django",
"ecosystem": "PyPI"
},
"version": "3.1.0"
}
Data Flow Overview
SkillSpector follows a strict data minimization approach for external communications:
- Dependency Discovery: The supply chain analyzer parses
requirements.txt,package.json, and other lock files to extract (name, version, ecosystem) tuples - OSV.dev Transmission: Only the dependency identifiers are sent to the batch query endpoint. No proprietary source code, project files, or internal paths leave the host
- LLM Analysis: When findings require explanation, only the specific code snippet triggering the alert is transmitted, truncated to provider token limits
Summary
- LLM Providers: Receive only relevant code snippets, file paths, language identifiers, and SARIF finding context—never secrets, environment variables, or unrelated project files
- OSV.dev: Receives only dependency names (normalized), versions, and ecosystem identifiers ("PyPI" or "npm")
- Token Safety: All LLM requests enforce provider-specific token limits via truncation in
llm_analyzer_base.py - No Source Exfiltration: Full source trees remain local; only necessary fragments transmit for analysis
Frequently Asked Questions
Does SkillSpector send my entire codebase to LLM providers?
No. According to the implementation in llm_analyzer_base.py, SkillSpector extracts only the specific file or line range associated with a SARIF finding. The content undergoes truncation to fit provider token limits before transmission, ensuring only relevant fragments leave your environment.
What dependency information does SkillSpector share with OSV.dev?
SkillSpector sends only the package name (lower-cased and hyphen-ified), ecosystem string ("PyPI" or "npm"), and version number. This data is constructed in osv_client.py by the _build_query function and transmitted via the query_batch method to the OSV.dev batch endpoint.
Are API keys or environment variables transmitted to external services?
No. The source code in providers/openai/provider.py and providers/chat_models.py shows that credentials are used locally to configure the LangChain client. Only the constructed analysis payloads—code snippets for LLMs and dependency identifiers for OSV.dev—are transmitted to external endpoints.
Can I audit what data SkillSpector sends before it leaves my network?
Yes. All external request construction happens in specific analyzer files. You can inspect src/skillspector/nodes/analyzers/llm_analyzer_base.py for LLM payloads and src/skillspector/nodes/analyzers/osv_client.py for OSV queries. Both files build standard JSON payloads that can be logged or intercepted for audit purposes before the HTTP POST occurs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →