How SkillSpector Integrates with OSV.dev to Query Vulnerable Dependencies

SkillSpector performs batched HTTP POST requests to the OSV.dev API, caches vulnerability results for one hour, and falls back to static curated lists when the API is unreachable.

SkillSpector, NVIDIA's automated skill analysis framework, identifies compromised third-party dependencies by querying the Open Source Vulnerabilities (OSV) database. The integration lives in osv_client.py and static_patterns_supply_chain.py, where a custom client_batches package queries, normalizes severity scores, and degrades gracefully when running in air-gapped environments. This implementation allows SkillSpector to map every dependency in a requirements.txt or package.json to known CVEs without manual intervention.

Batch Query Architecture

The OSV.dev client minimizes network overhead by bundling all dependency lookups into a single HTTP request. This architecture reduces latency when analyzing large dependency trees.

Building the OSV Request

In src/skillspector/nodes/analyzers/osv_client.py, the private function _build_query() (lines 92‑96) constructs the JSON payload for the batch API. It transforms every (name, version) tuple extracted from the dependency file into a structured query object that specifies the package ecosystem (e.g., PyPI or npm).

Executing the Batch Query

The public entry point query_batch() (lines 122‑135) sends the compiled payload to https://api.osv.dev/v1/querybatch using the httpx library. The function returns a list of vulnerability identifiers grouped by package. If the request fails or returns no data, the client immediately triggers the fallback mechanism.

Fetching Detailed Vulnerability Data

For each returned vulnerability ID, _fetch_vuln_details() (lines 86‑96) performs a secondary GET request to https://api.osv.dev/v1/vulns/<ID>. This endpoint provides the full advisory text, aliases (such as CVE identifiers), and severity scores required to generate actionable findings.

Caching and Offline Fallback Strategy

SkillSpector implements a two-tier resilience strategy to ensure consistent results regardless of network conditions.

In-Memory Result Caching

The module maintains a private _cache dictionary (lines 57‑61) keyed by (name, version, ecosystem). Results expire after _CACHE_TTL_SECS (one hour), preventing redundant API calls when scanning multiple skills that share the same dependency manifest. Cache hits bypass both the batch POST and the detail-fetching stage entirely.

Static Vulnerability Lists

When the OSV API is unreachable or returns no match, the analyzer falls back to hard-curated lists _FALLBACK_VULNERABLE_PYPI and _FALLBACK_VULNERABLE_NPM located in static_patterns_supply_chain.py (lines 104‑132). These lists contain high-impact packages known to be vulnerable, ensuring that critical security issues are still flagged even in offline environments.

Severity Normalization

OSV.dev aggregates severity scores from multiple sources—including GHSA, ecosystem-specific advisories, and CVSS vectors—which SkillSpector must normalize to a consistent internal scale.

Extracting CVSS and GHSA Scores

The function _severity_from_vuln() (lines 49‑57) parses the OSV response and extracts the most reliable severity indicator available. It prioritizes GHSA-provided scores, falling back to ecosystem-specific ratings or CVSS vectors when GHSA data is absent.

Mapping to Internal Severity Levels

_osv_severity_to_app() (lines 60‑68) converts the extracted OSV strings into SkillSpector's internal Severity enum, mapping values to CRITICAL, HIGH, MEDIUM, or LOW. This normalization ensures that findings from different ecosystems (Python, JavaScript, etc.) are comparable and actionable.

Integration with Supply Chain Analysis

The OSV client plugs into SkillSpector's supply-chain analyzer through two primary entry points in static_patterns_supply_chain.py.

The _sc4_from_osv Entry Point

_sc4_from_osv() (lines 88‑108) orchestrates the live lookup. It converts the parsed dependency list into (name, version) tuples, invokes query_batch(), and transforms the returned data into AnalyzerFinding objects. Each finding contains the package name, affected version, number of advisories, a human-readable summary generated by _format_vuln_ids(), and the normalized severity from _osv_severity_to_app().

Fallback Processing with _sc4_from_fallback

When OSV returns no results for a package, _sc4_from_fallback() (lines 136‑176) checks the static vulnerability lists. If the package appears in _FALLBACK_VULNERABLE_PYPI or _FALLBACK_VULNERABLE_NPM, the analyzer generates a finding with pre-defined severity levels, ensuring zero false negatives when the OSV.dev API is unavailable.

Practical Code Examples

Running a Manual OSV Lookup

You can query the OSV.dev integration directly from a Python REPL to audit a dependency list:

from skillspector.nodes.analyzers.osv_client import query_batch, ECOSYSTEM_PYPI

# Example packages extracted from a requirements.txt file

packages = [("requests", "2.28.1"), ("django", None)]

# Perform the batch lookup

results = query_batch(packages, ECOSYSTEM_PYPI)

for pkg, vulns in zip(packages, results):
    name, version = pkg
    print(f"{name}{'==' + version if version else ''}:")
    for v in vulns:
        print(f"  - {v.vuln_id} [{v.severity}] {v.summary}")

Integrating Results into the Analyzer

The following excerpt from static_patterns_supply_chain.py demonstrates how the supply-chain node consumes the OSV client:

def _sc4_from_osv(
    packages: list[tuple[str, str | None, int]],
    ecosystem: str,
    file_path: str,
    tag: list[str],
) -> tuple[list[AnalyzerFinding], set[str]]:
    # Build (name, version) tuples for the OSV client

    pkg_pairs = [(name, version) for name, version, _ in packages]

    # Query OSV.dev (live lookup)

    osv_results = query_batch(pkg_pairs, ecosystem)

    # Convert OSV results → AnalyzerFinding objects

    ...

Summary

  • Batched API queries: SkillSpector sends one HTTP POST to https://api.osv.dev/v1/querybatch containing all dependencies, then fetches individual vulnerability details via https://api.osv.dev/v1/vulns/<ID>.
  • One-hour cache: Results are cached in memory by (name, version, ecosystem) for 3600 seconds to avoid redundant network traffic.
  • Graceful degradation: If the OSV.dev API is unreachable, the analyzer falls back to static lists _FALLBACK_VULNERABLE_PYPI and _FALLBACK_VULNERABLE_NPM.
  • Severity normalization: The client extracts GHSA, ecosystem-specific, or CVSS scores and maps them to internal CRITICAL, HIGH, MEDIUM, or LOW levels.
  • Supply-chain integration: _sc4_from_osv() converts raw OSV data into AnalyzerFinding objects, while _sc4_from_fallback() handles offline scenarios.

Frequently Asked Questions

How does SkillSpector handle network failures when querying OSV.dev?

If the httpx request to https://api.osv.dev/v1/querybatch fails or returns no data, the analyzer immediately falls back to static curated lists defined in static_patterns_supply_chain.py. This ensures that known vulnerable packages are still reported even in air-gapped environments.

What severity scoring systems does SkillSpector support from OSV.dev?

The _severity_from_vuln() function in osv_client.py prioritizes GHSA-provided scores, then ecosystem-specific advisories, and finally CVSS vectors. These are normalized to SkillSpector's internal Severity enum via _osv_severity_to_app().

How long does SkillSpector cache OSV.dev vulnerability results?

Results are cached in the _cache dictionary for one hour (_CACHE_TTL_SECS = 3600 seconds). The cache is keyed by package name, version, and ecosystem, allowing subsequent scans to skip redundant API calls for identical dependencies.

Can SkillSpector analyze dependencies from package managers other than PyPI and npm?

Yes. While the static fallback lists currently cover PyPI and npm (_FALLBACK_VULNERABLE_PYPI and _FALLBACK_VULNERABLE_NPM), the query_batch() function accepts any ecosystem string supported by OSV.dev, including Maven, Go, Rust, and NuGet, enabling live vulnerability lookups for any OSV-compatible package manager.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →