How OSV.dev Live Vulnerability Lookup Works in NVIDIA SkillSpector
TLDR: SkillSpector's SC4 supply-chain analyzer performs real-time vulnerability lookups against the OSV.dev database using a batched HTTP client with in-memory caching, automatic severity scoring, and graceful fallback to static vulnerability lists when the API is unreachable.
NVIDIA's SkillSpector framework analyzes AI skills for supply chain security risks through its SC4 analyzer. The OSV.dev live vulnerability lookup module, implemented in src/skillspector/nodes/analyzers/osv_client.py, replaces hard-coded vulnerability lists with dynamic queries to the Open Source Vulnerabilities (OSV) database, enabling up-to-date security assessments for Python and npm dependencies.
The OSV.dev Client Architecture
The core functionality resides in src/skillspector/nodes/analyzers/osv_client.py, which provides a Pythonic interface to the OSV.dev REST API with built-in caching and error resilience.
Batch Query Processing
The query_batch function serves as the primary entry point. It accepts a list of (package_name, version) tuples plus an ecosystem identifier ("PyPI" or "npm"), then orchestrates the lookup process according to the implementation in lines 27-34 and 55-63.
First, the client checks an in-memory _cache dictionary keyed by normalized (name-version-ecosystem) tuples. Cached results return immediately, while uncached queries are collected for a single HTTP POST to https://api.osv.dev/v1/querybatch (lines 68-71).
The batch endpoint returns vulnerability IDs per package, which the client then resolves by calling GET /v1/vulns/{id} for each result. To maintain low latency, the client caps detail fetches at 10 IDs per package (lines 99-110). Each vulnerability is parsed into a VulnResult dataclass containing the ID, summary, severity label, and aliases (lines 60-68, 90-98).
Severity Determination
The helper method _severity_from_vuln extracts severity through a prioritized fallback chain implemented in lines 118-132:
- Database-specific severity values
- Ecosystem-specific severity values
- CVSS vectors parsed by
_estimate_cvss_severity - Default to
"HIGH"if no severity data exists
Caching and TTL
Results are stored in _cache with a 1-hour TTL (_CACHE_TTL_SECS = 3600.0) as defined in lines 73-79 and 81-89. Subsequent scans of identical dependencies reuse cached data, eliminating redundant network traffic across repeated analyses of the same skill.
Graceful Degradation
If the batch request fails due to network errors, timeouts, or malformed responses, the client logs a warning, sets _last_query_ok = False, and returns empty results (lines 96-103, 96-106). This allows the SC4 analyzer to fall back to its static vulnerability list rather than failing the entire scan.
The is_available() method provides a lightweight connectivity probe by sending a test POST to the batch endpoint (lines 104-113), enabling the analyzer to verify OSV.dev reachability before initiating full scans.
Integration with SC4 Supply Chain Analysis
The SC4 analyzer in src/skillspector/nodes/analyzers/static_patterns_supply_chain.py integrates the OSV client through the _sc4_from_osv() function.
When scanning dependencies, the analyzer calls osv_client.query_batch() and maps the returned VulnResult objects back to their originating packages. It aggregates findings by selecting the worst severity per dependency and emits a single SC4 finding for each vulnerable package.
If query_batch() returns empty results or the OSV service is unavailable, the analyzer automatically invokes _sc4_from_fallback(), which uses the original static vulnerability list to ensure the supply chain analysis completes without interruption.
Practical Code Examples
Querying Vulnerabilities Directly
from skillspector.nodes.analyzers.osv_client import query_batch, ECOSYSTEM_PYPI
# Dependencies extracted from requirements.txt
deps = [
("requests", "2.25.0"),
("jinja2", "2.11.3"),
]
# Perform live lookup
vulns_by_pkg = query_batch(deps, ECOSYSTEM_PYPI)
for (name, version), vulns in zip(deps, vulns_by_pkg):
if vulns:
print(f"{name}=={version} has {len(vulns)} known vulnerabilities:")
for v in vulns:
print(f" • {v.vuln_id} – {v.summary} (severity={v.severity})")
else:
print(f"{name}=={version} is clean")
Checking OSV.dev Availability
from skillspector.nodes.analyzers.osv_client import is_available
if is_available():
print("OSV.dev reachable – live lookups enabled")
else:
print("OSV.dev unreachable – using static fallback data")
Managing the Cache
from skillspector.nodes.analyzers.osv_client import clear_cache
clear_cache() # Reset in-memory cache before fresh scan
Summary
- Location: Core implementation in
src/skillspector/nodes/analyzers/osv_client.pywith SC4 integration instatic_patterns_supply_chain.py - API: Uses OSV.dev
querybatchendpoint for bulk lookups and individualGET /v1/vulns/{id}calls for detail retrieval - Performance: In-memory caching with 1-hour TTL prevents redundant API calls across repeated scans
- Reliability: Graceful degradation to static vulnerability lists when OSV.dev is unreachable or returns errors
- Coverage: Supports PyPI and npm ecosystems with automatic severity normalization
Frequently Asked Questions
How does SkillSpector handle rate limiting when querying OSV.dev?
The client minimizes API load by batching package queries into a single HTTP POST request to the querybatch endpoint. Additionally, it caches results for one hour (_CACHE_TTL_SECS = 3600.0) and caps detail fetches at 10 vulnerability IDs per package, significantly reducing the total number of individual API calls required per scan.
What happens if OSV.dev is offline during a scan?
If the OSV.dev API is unreachable or returns errors, the query_batch function sets _last_query_ok = False and returns empty results. The SC4 analyzer then automatically falls back to static_patterns_supply_chain.py, which uses an internal static vulnerability list to ensure the supply chain analysis completes without interruption.
Which package ecosystems does the OSV.dev lookup support?
According to the source code in osv_client.py, the implementation explicitly supports "PyPI" for Python packages and "npm" for Node.js packages. The ecosystem identifier is passed to the query_batch function and normalized into the cache key alongside package name and version.
How is vulnerability severity determined when OSV entries lack CVSS scores?
The _severity_from_vuln helper implements a fallback hierarchy: it first checks for database-specific severity values, then ecosystem-specific values, then attempts to parse CVSS vectors using _estimate_cvss_severity. If no severity data exists, it defaults to "HIGH" to ensure conservative security reporting.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →