# How OSV.dev Integration Queries Work for CVE Lookups in SkillSpector

> Discover how SkillSpector integrates with OSV.dev for fast CVE lookups using in-memory caching and fallback databases to ensure real-time vulnerability scanning.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: how-to-guide
- Published: 2026-07-09

---

**SkillSpector queries the OSV.dev batch API with intelligent in-memory caching to perform real-time CVE lookups, falling back to static vulnerability databases when the service is unreachable.**

NVIDIA's SkillSpector performs live vulnerability detection by integrating with the Open Source Vulnerabilities (OSV) database. The integration leverages OSV.dev's batch API to minimize network overhead while providing accurate CVE identification for Python and JavaScript dependencies. This architecture ensures that supply chain scans reflect current threat intelligence without compromising performance or reliability.

## OSV.dev Client Architecture

The OSV.dev integration resides in [`src/skillspector/nodes/analyzers/osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/osv_client.py), which exposes two primary public functions: `is_available()` for connectivity checks and `query_batch()` for vulnerability lookups. The client implements a lightweight in-memory cache keyed by `(name, version, ecosystem)` tuples with a one-hour TTL (`_CACHE_TTL_SECS = 3600`) to prevent redundant API calls during large-scale repository scans.

The module defines ecosystem constants `ECOSYSTEM_PYPI` and `ECOSYSTEM_NPM` to identify package registries. Internal helpers handle JSON query construction, response parsing, and severity estimation without requiring heavy external dependencies.

## Batch Query Workflow

The CVE lookup process follows a deterministic eight-step pipeline that optimizes for both accuracy and latency.

### Package-Level Request Generation

For each dependency, the analyzer constructs a JSON query containing the package name, optional version, and ecosystem identifier. The `_build_query` helper formats these into the structure required by OSV.dev's specification, ensuring compatibility with both PyPI and npm ecosystems.

### In-Memory Caching Strategy

Before issuing network requests, the client checks an in-process cache via `_get_cached`. Cache hits return immediately with previously retrieved `VulnResult` objects. Misses proceed to the batch API, with results subsequently stored via `_put_cache` for the duration of the scan. This caching layer reduces API traffic by deduplicating requests for popular dependencies across multiple analysis passes.

### Batch API Requests

Uncached queries are aggregated into a single POST request to `https://api.osv.dev/v1/querybatch` (defined as `_OSV_BATCH_URL`). The batch endpoint accepts multiple package queries in one payload, returning a list of vulnerability IDs for each submitted dependency. This approach minimizes HTTP overhead compared to individual API calls.

### Vulnerability Detail Fetching

For each returned vulnerability ID, the client retrieves full advisory details via GET requests to `https://api.osv.dev/v1/vulns/<ID>`. To prevent latency spikes in packages with numerous CVEs, the implementation limits detail fetching to the first ten vulnerability IDs. The `_fetch_vuln_details` helper parses the JSON response into structured `VulnResult` objects containing `vuln_id`, `summary`, `severity`, and `aliases` (which include CVE numbers when available).

### Severity Estimation

OSV records may contain explicit severity scores or raw CVSS vectors. When only CVSS vectors are present, `_estimate_cvss_severity` calculates coarse severity bands (CRITICAL, HIGH, MEDIUM, LOW) without importing full CVSS calculation libraries. This keeps the dependency footprint minimal while providing actionable risk ratings.

### Graceful Degradation and Fallbacks

If the OSV request fails due to network errors, timeouts, or air-gapped environments, `query_batch` returns empty result lists. The supply chain analyzer ([`static_patterns_supply_chain.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_patterns_supply_chain.py)) detects these failures and automatically falls back to built-in static vulnerability lists (`_FALLBACK_VULNERABLE_PYPI` and `_FALLBACK_VULNERABLE_NPM`). The `is_available()` method performs a lightweight connectivity check using a dummy query before attempting full batch operations.

## Supply Chain Analyzer Integration

The [`static_patterns_supply_chain.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_patterns_supply_chain.py) analyzer orchestrates the OSV.dev integration by extracting dependency tuples from [`requirements.txt`](https://github.com/NVIDIA/SkillSpector/blob/main/requirements.txt), [`pyproject.toml`](https://github.com/NVIDIA/SkillSpector/blob/main/pyproject.toml), [`setup.py`](https://github.com/NVIDIA/SkillSpector/blob/main/setup.py), `Pipfile`, or [`package.json`](https://github.com/NVIDIA/SkillSpector/blob/main/package.json) files. It invokes `query_batch` with the extracted `(name, version)` pairs and appropriate ecosystem constants.

For each package returning vulnerabilities, the analyzer generates SC4 findings that include the most severe result and a human-readable list of advisory IDs. The `_format_vuln_ids` helper processes aliases to ensure CVE numbers appear alongside OSV identifiers in the final report.

## Implementation Examples

Check OSV.dev availability before performing lookups:

```python
from skillspector.nodes.analyzers.osv_client import (
    ECOSYSTEM_PYPI,
    query_batch,
    is_available,
)

if is_available():
    packages = [("requests", "2.28.0"), ("urllib3", "1.26.0")]
    results = query_batch(packages, ECOSYSTEM_PYPI)
    
    for (name, version), vulns in zip(packages, results):
        if vulns:
            print(f"{name}@{version}: {len(vulns)} vulnerabilities")

```

The supply chain analyzer automatically invokes this logic during dependency scans:

```python

# Inside static_patterns_supply_chain.py

pkg_pairs = [(name, version) for name, version, _ in dependencies]
osv_results = query_batch(pkg_pairs, ECOSYSTEM_PYPI)

# osv_results is List[List[VulnResult]] parallel to pkg_pairs

for pkg, vuln_list in zip(pkg_pairs, osv_results):
    if vuln_list:
        # Generate SC4 finding with highest severity

        finding = build_finding(pkg, vuln_list)

```

## Key Implementation Files

- **[`src/skillspector/nodes/analyzers/osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/osv_client.py)**: Implements HTTP client, caching layer, severity parsing, and connectivity testing.
- **[`src/skillspector/nodes/analyzers/static_patterns_supply_chain.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/static_patterns_supply_chain.py)**: Orchestrates dependency extraction, calls `query_batch`, and formats findings with fallback handling.
- **[`tests/unit/test_osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/tests/unit/test_osv_client.py)**: Unit tests validating cache behavior, API error paths, and severity estimation logic.

## Summary

- SkillSpector uses OSV.dev's batch API endpoint (`v1/querybatch`) to query multiple packages in a single HTTP request, minimizing network latency.
- An in-memory cache with a one-hour TTL prevents redundant lookups for dependencies encountered across different files in the same scan.
- Vulnerability details are fetched individually (capped at ten per package) and parsed into `VulnResult` objects containing CVE aliases and severity estimates.
- When OSV.dev is unreachable, the system automatically falls back to static vulnerability lists built into the analyzer.
- The implementation requires no external authentication and supports both PyPI and npm ecosystems through standardized query structures.

## Frequently Asked Questions

### How does SkillSpector handle network failures when querying OSV.dev?

The `query_batch` function catches network exceptions and timeouts, returning empty result lists that trigger the supply chain analyzer's fallback mechanism. This mechanism uses static vulnerability definitions stored in `_FALLBACK_VULNERABLE_PYPI` and `_FALLBACK_VULNERABLE_NPM` to ensure scans complete deterministically even in air-gapped environments.

### What information does SkillSpector extract from OSV.dev vulnerability records?

For each vulnerability ID returned by the batch query, SkillSpector retrieves the full advisory and extracts the `vuln_id`, human-readable `summary`, severity score (or estimated CVSS severity), and any `aliases` which typically include CVE identifiers. This data populates `VulnResult` objects used to generate security findings.

### Does SkillSpector cache CVE lookup results between scans?

SkillSpector maintains an in-process cache with a one-hour TTL (`_CACHE_TTL_SECS = 3600`) that stores `VulnResult` lists keyed by package name, version, and ecosystem. This cache persists for the duration of a single process execution, deduplicating requests when the same dependency appears in multiple manifest files, but does not persist across separate SkillSpector invocations.

### Which package ecosystems does the OSV.dev integration support?

The current implementation supports **PyPI** (`ECOSYSTEM_PYPI`) for Python packages and **npm** (`ECOSYSTEM_NPM`) for JavaScript packages. The query builder in [`osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/osv_client.py) formats requests according to OSV.dev's specification for these ecosystems, enabling accurate version matching and vulnerability correlation for dependencies listed in [`requirements.txt`](https://github.com/NVIDIA/SkillSpector/blob/main/requirements.txt), [`package.json`](https://github.com/NVIDIA/SkillSpector/blob/main/package.json), and similar manifest files.