# How SkillSpector Integrates with OSV.dev for Vulnerability Lookup

> Discover how SkillSpector integrates with OSV.dev using its Python client to batch-query the OSV.dev API, cache results, and enrich dependency graphs with live vulnerability data.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: how-to-guide
- Published: 2026-07-13

---

**SkillSpector integrates with OSV.dev through a dedicated Python client module that batch-queries the public OSV.dev API, caches results for one hour, and extracts severity ratings to enrich dependency graphs with live vulnerability data.**

The NVIDIA SkillSpector project performs supply-chain security analysis by integrating with OSV.dev to identify vulnerabilities in project dependencies. This integration leverages a specialized client module that handles HTTP communication, response caching, and severity estimation without requiring heavy external libraries. Understanding how SkillSpector integrates with OSV.dev reveals a robust architecture designed for performance, reliability, and graceful degradation when network services are unavailable.

## OSV.dev Client Architecture

The integration centers on [`src/skillspector/nodes/analyzers/osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/osv_client.py), which encapsulates all OSV.dev communication. This module defines constants `_OSV_BATCH_URL` and `_OSV_VULN_URL` that point to the batch query endpoint (`/v1/querybatch`) and single vulnerability endpoint (`/v1/vulns`) respectively. These endpoints enable the client to submit HTTP POST requests for multiple packages simultaneously and retrieve detailed CVE information.

### Batch Query Implementation

The `query_batch(packages, ecosystem)` function serves as the primary interface for the rest of the codebase. It accepts a list of `(name, version)` tuples and an ecosystem identifier such as `"PyPI"` or `"npm"`. The function constructs requests via `_build_query` and posts them to `https://api.osv.dev/v1/querybatch`.

Results are cached in-memory using the `_cache` dictionary with a TTL of `_CACHE_TTL_SECS` (3600 seconds). Cache keys are normalized using the format `(name.lower().replace("_","-"), version, ecosystem)` to ensure consistent lookups. When a package has no vulnerabilities, the client stores an empty list and emits a log entry. If the API is unreachable, the function returns empty lists for all packages and sets `_last_query_ok = False`, enabling the supply-chain analyzer to detect the failure.

## Severity Extraction Without Heavy Dependencies

Rather than importing comprehensive CVSS libraries, the client implements `_estimate_cvss_severity` and `_severity_from_vuln` to parse CVSS strings directly. These helpers analyze CVSS:3.1 vectors and return discrete severity labels—**"CRITICAL"**, **"HIGH"**, **"MEDIUM"**, or **"LOW"**—based on the proportion of high-severity metrics present. This approach keeps the dependency footprint light while providing actionable severity data for the UI.

## Health Monitoring and Configuration

The client provides two mechanisms for monitoring OSV.dev availability. The `is_available()` method performs a minimal POST request with a dummy package (`pip` on PyPI) and returns `True` only if the response status is 200. After batch queries, the `was_osv_reachable()` function reports whether the most recent call succeeded, allowing analyzers to surface fallback warnings when the API is down.

Operators can tune network behavior using the environment variable `SKILLSPECTOR_OSV_TIMEOUT`, which sets the HTTP timeout in seconds (defaulting to 30). The client reads this variable at import time and falls back to the default if conversion fails.

## Practical Implementation Examples

The following patterns demonstrate how SkillSpector components interact with the OSV.dev client.

### Checking Service Availability

Before launching analysis, call `is_available()` to verify the endpoint is reachable:

```python
from skillspector.nodes.analyzers.osv_client import is_available
import logging

logger = logging.getLogger(__name__)

if not is_available():
    logger.warning("OSV.dev endpoint is unreachable – live look‑ups disabled")

```

### Executing Batch Vulnerability Queries

To query vulnerabilities for multiple dependencies, prepare a list of `(name, version)` tuples and specify the ecosystem:

```python
from skillspector.nodes.analyzers.osv_client import query_batch

# Example dependencies for a Python project

deps = [
    ("numpy", "1.26.0"),
    ("torch", None),          # version‑unspecified – OSV will look for any vuln

    ("requests", "2.31.0"),
]

# Query OSV.dev batch endpoint

vuln_lists = query_batch(deps, ecosystem="PyPI")

# Process results (vuln_lists is parallel to deps)

for (name, version), vulns in zip(deps, vuln_lists):
    if not vulns:
        continue
    for v in vulns:
        print(
            f"[{v.severity}] {name}{'==' + version if version else ''}: "
            f"{v.vuln_id} – {v.summary}"
        )

```

### Handling Failures and Cache Management

Detect API failures to provide fallback messaging, and manage the cache when necessary:

```python
from skillspector.nodes.analyzers.osv_client import was_osv_reachable, clear_cache

# Check if the last batch query succeeded

if not was_osv_reachable():
    logger.info(
        "OSV.dev could not be contacted – displayed results may be incomplete"
    )

# Flush cached entries (advanced usage)

clear_cache()

```

## Integration with Supply Chain Analysis

The supply-chain analyzer, located under `src/skillspector/nodes/analyzers/`, consumes the OSV client to implement the SC4 live vulnerability lookup feature documented in [`docs/SC4-osv-live-vulnerability-lookups.md`](https://github.com/NVIDIA/SkillSpector/blob/main/docs/SC4-osv-live-vulnerability-lookups.md). It extracts dependencies from project manifests such as [`requirements.txt`](https://github.com/NVIDIA/SkillSpector/blob/main/requirements.txt) or [`package-lock.json`](https://github.com/NVIDIA/SkillSpector/blob/main/package-lock.json), then calls `query_batch` with the extracted packages. Each returned `VulnResult` object—containing `vuln_id`, `summary`, `severity`, and `aliases`—is injected into the dependency graph.

When `was_osv_reachable()` returns `False`, the analyzer falls back to a static vulnerability list while flagging the incomplete results for the user. This architecture isolates external service dependencies, ensuring that the core analysis logic remains testable and resilient. Unit tests in [`tests/unit/test_osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/tests/unit/test_osv_client.py) verify batch query behavior, caching logic, and failure handling.

## Summary

- **SkillSpector integrates with OSV.dev** through a dedicated client module at [`src/skillspector/nodes/analyzers/osv_client.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/osv_client.py) that handles all HTTP communication.
- **Batch queries** via `query_batch()` support ecosystems like PyPI and npm, with built-in TTL caching (3600 seconds) to reduce network traffic.
- **Lightweight severity extraction** parses CVSS:3.1 vectors without heavy dependencies, returning standardized labels (CRITICAL, HIGH, MEDIUM, LOW).
- **Health monitoring** through `is_available()` and `was_osv_reachable()` enables graceful degradation when the OSV.dev service is unreachable.
- **Environment configuration** via `SKILLSPECTOR_OSV_TIMEOUT` allows operators to adjust network timeouts from the default 30 seconds.

## Frequently Asked Questions

### How does SkillSpector handle network failures when querying OSV.dev?

When the OSV.dev API is unreachable, the `query_batch()` function catches `httpx.HTTPError` exceptions and returns empty lists for all requested packages. It sets the internal flag `_last_query_ok = False`, which callers can detect via `was_osv_reachable()`. The supply-chain analyzer uses this signal to fall back to a static vulnerability list and warn users that live results may be incomplete.

### What caching mechanism does SkillSpector use for OSV.dev responses?

The client implements a simple in-memory cache using the `_cache` dictionary with a TTL of 3600 seconds (`_CACHE_TTL_SECS`). Cache keys are normalized to lowercase with underscores replaced by hyphens, ensuring consistent lookups across different package naming conventions. This design prevents repeated network calls for identical dependency queries within the same analysis session.

### How does SkillSpector determine vulnerability severity without using CVSS libraries?

The client includes `_estimate_cvss_severity` and `_severity_from_vuln` helper functions that parse raw CVSS:3.1 vectors as strings. These functions calculate severity by analyzing the proportion of high-impact metrics in the vector, returning discrete labels (CRITICAL, HIGH, MEDIUM, LOW) without importing heavy CVSS calculation libraries. This approach minimizes dependencies while providing actionable severity data for the UI.

### Can operators configure the OSV.dev integration timeout settings?

Yes, the client reads the `SKILLSPECTOR_OSV_TIMEOUT` environment variable at import time to set the HTTP timeout for all OSV.dev requests. If the variable is unset or contains an invalid value, the client defaults to 30 seconds. This allows operators to adjust network behavior based on their infrastructure and latency requirements.