How SkillSpector Integrates with OSV.dev for Vulnerability Lookups

SkillSpector queries the OSV.dev vulnerability database via batch HTTP requests to identify security issues in dependency files, implementing a two-stage lookup with in-memory caching and a static fallback for offline environments.

The NVIDIA SkillSpector project automates security analysis of AI skill dependencies by integrating with the Open Source Vulnerabilities (OSV) database. This integration enables real-time detection of vulnerable packages in requirements.txt, package.json, and other manifest files. Understanding how SkillSpector integrates with OSV.dev reveals a robust architecture designed for both accuracy and resilience.

The OSV.dev Batch Query Architecture

SkillSpector implements a two-stage lookup process against the OSV.dev API to minimize network overhead while maximizing vulnerability coverage.

Building the Batch Request

In src/skillspector/nodes/analyzers/osv_client.py, the _build_query() method (lines 92–96) constructs a single HTTP POST payload containing all packages extracted from a dependency file. This batch approach allows SkillSpector to query multiple vulnerabilities in one request rather than generating individual network calls for each dependency.

Executing the HTTP POST

The query_batch() function (lines 122–135) transmits the assembled payload to https://api.osv.dev/v1/querybatch using the httpx library. This method handles the bulk vulnerability lookup, returning a list of vulnerability IDs associated with the submitted package names and versions. The implementation parses the JSON response and prepares the data for detailed enrichment.

Fetching Detailed Vulnerability Data

For each vulnerability ID returned by the batch query, _fetch_vuln_details() (lines 86–96) performs a subsequent GET request to https://api.osv.dev/v1/vulns/<ID>. This second stage retrieves comprehensive vulnerability metadata including severity scores, summary descriptions, and alias mappings (such as GHSA or CVE identifiers).

Caching and Offline Resilience

The OSV client implements multiple strategies to ensure reliable operation across network conditions.

In-Memory Cache Implementation

To prevent redundant API calls within a single analysis session, SkillSpector maintains an in-memory cache stored in _cache (lines 57–61). The cache uses a composite key of (name, version, ecosystem) and respects a one-hour TTL defined by _CACHE_TTL_SECS. When processing multiple files or repeated analyses, cached results return immediately without additional network requests.

Static Fallback Database

When the OSV API is unreachable or returns no data, the analyzer falls back to curated static vulnerability lists. The _FALLBACK_VULNERABLE_PYPI and _FALLBACK_VULNERABLE_NPM dictionaries in src/skillspector/nodes/analyzers/static_patterns_supply_chain.py (lines 104–132) provide offline detection capabilities for air-gapped environments. This ensures that known vulnerable packages are still reported even without internet connectivity.

Severity Normalization and Mapping

Raw OSV data requires normalization to match SkillSpector's internal severity classifications.

The _severity_from_vuln() function (lines 49–57) extracts the most reliable severity indicator from available sources, prioritizing GHSA ratings, ecosystem-specific scores, or CVSS vectors. The helper _osv_severity_to_app() (lines 60–68) then maps these values to the internal Severity enum, normalizing them to one of four levels: CRITICAL, HIGH, MEDIUM, or LOW.

Integration with the Supply Chain Analyzer

The OSV client integrates into SkillSpector's broader supply-chain security analysis through the static_patterns_supply_chain node.

The _sc4_from_osv() function (lines 88–108) orchestrates the live vulnerability lookup by invoking query_batch() and converting results into AnalyzerFinding objects. Each finding contains the package name, version, advisory count, human-readable summary generated by _format_vuln_ids(), and the normalized severity derived from OSV data (lines 120–133).

For packages not returned by OSV (either due to API failures or absence of known vulnerabilities), _sc4_from_fallback() (lines 136–176) processes the static vulnerability lists, ensuring comprehensive coverage regardless of network availability.

Practical Implementation Example

You can query OSV.dev directly using SkillSpector's client from a Python REPL:

from skillspector.nodes.analyzers.osv_client import query_batch, ECOSYSTEM_PYPI

# Example packages extracted from a requirements.txt file

packages = [("requests", "2.28.1"), ("django", None)]

# Perform the batch lookup

results = query_batch(packages, ECOSYSTEM_PYPI)

for pkg, vulns in zip(packages, results):
    name, version = pkg
    print(f"{name}{'==' + version if version else ''}:")
    for v in vulns:
        print(f"  - {v.vuln_id} [{v.severity}] {v.summary}")

Within the analyzer pipeline, the integration appears as follows:

def _sc4_from_osv(
    packages: list[tuple[str, str | None, int]],
    ecosystem: str,
    file_path: str,
    tag: list[str],
) -> tuple[list[AnalyzerFinding], set[str]]:
    # Build (name, version) tuples for the OSV client

    pkg_pairs = [(name, version) for name, version, _ in packages]

    # Query OSV.dev (live lookup)

    osv_results = query_batch(pkg_pairs, ecosystem)

    # Convert OSV results → AnalyzerFinding objects

    ...

Summary

  • Batch API Architecture: SkillSpector uses _build_query() and query_batch() in osv_client.py to send single HTTP POST requests to https://api.osv.dev/v1/querybatch, reducing network overhead when checking multiple dependencies.
  • Two-Stage Lookup: The system first queries for vulnerability IDs, then uses _fetch_vuln_details() to retrieve full metadata including severity and aliases via https://api.osv.dev/v1/vulns/<ID>.
  • Resilience Mechanisms: A one-hour in-memory cache (_CACHE_TTL_SECS) prevents redundant requests, while static fallback lists in static_patterns_supply_chain.py ensure offline operation.
  • Severity Normalization: The _severity_from_vuln() and _osv_severity_to_app() functions map OSV severity ratings to standardized CRITICAL, HIGH, MEDIUM, or LOW classifications.
  • Supply Chain Integration: The static_patterns_supply_chain node calls _sc4_from_osv() to generate AnalyzerFinding objects, with _sc4_from_fallback() providing coverage when OSV is unavailable.

Frequently Asked Questions

How does SkillSpector handle network failures when querying OSV.dev?

When the OSV API is unreachable, SkillSpector automatically falls back to static vulnerability lists defined in _FALLBACK_VULNERABLE_PYPI and _FALLBACK_VULNERABLE_NPM within static_patterns_supply_chain.py. This ensures that known vulnerable packages are still detected even in air-gapped environments without internet access.

What ecosystem does SkillSpector support for OSV.dev lookups?

The OSV client supports multiple ecosystems including PyPI (ECOSYSTEM_PYPI) and npm, as evidenced by the fallback lists for both Python and Node.js packages. The query_batch() function accepts an ecosystem parameter that identifies the package registry to query against the OSV database.

How long does SkillSpector cache OSV.dev query results?

Results are cached in memory for one hour using the _CACHE_TTL_SECS constant. The cache uses a composite key of package name, version, and ecosystem to ensure that repeated queries for the same dependency within the same analysis session do not generate redundant API calls.

What severity scoring system does SkillSpector use for OSV vulnerabilities?

SkillSpector normalizes OSV severity data to four standardized levels: CRITICAL, HIGH, MEDIUM, and LOW. The _severity_from_vuln() function extracts the most reliable available metric (prioritizing GHSA, ecosystem-specific, or CVSS scores), which _osv_severity_to_app() then maps to the internal Severity enum used throughout the application.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →