SkillSpector Risk Scoring Algorithm and Severity Calculation Explained

SkillSpector's risk scoring algorithm converts security findings into a 0-100 numeric score using a weighted severity system, applies a 1.3x multiplier for executable scripts, and maps the result to severity bands and installation recommendations.

NVIDIA's SkillSpector employs a deterministic approach to quantify security risks from analyzed AI skills. The entire logic resides in the _compute_risk_score function within src/skillspector/nodes/report.py, transforming lists of qualitative findings into actionable numeric scores that drive the final report generation.

How the Risk Score Is Computed

The algorithm processes findings through a multi-stage pipeline that weights severity levels, adjusts for executable content, and normalizes the output.

Base Points by Severity Level

Each finding contributes fixed points based on its annotated severity. In lines 82-92 of src/skillspector/nodes/report.py, the function assigns:

  • CRITICAL → 50 points
  • HIGH → 25 points
  • MEDIUM → 10 points
  • LOW → 5 points

The function iterates through the findings list, summing these base values into a subtotal before applying modifiers.

The Executable Script Multiplier

If the analyzed skill contains executable scripts (indicated by the boolean flag has_executable_scripts), the subtotal is multiplied by 1.3 and rounded down. This adjustment, implemented in lines 92-94 of src/skillspector/nodes/report.py, elevates the risk profile of skills that can execute arbitrary code.

Score Normalization

The final calculation clamps the resulting value to the range [0, 100]. This ensures that even skills with numerous critical findings receive a capped score that fits the standardized scale.

Mapping Scores to Severity Bands and Recommendations

Once normalized, the numeric score maps to human-readable severity bands using the _RISK_SEVERITY_BANDS threshold table defined in lines 54-56 of src/skillspector/nodes/report.py:

Threshold Severity Band
≥ 81 CRITICAL
≥ 51 HIGH
≥ 21 MEDIUM
≥ 0 LOW

The algorithm selects the first band whose threshold is less than or equal to the computed score.

Recommendation Mapping

The severity band then translates into specific action recommendations via the _RISK_RECOMMENDATION static mapping (lines 56-60). These recommendations typically include classifications like "SAFE", "CAUTION", or "DO_NOT_INSTALL", providing clear guidance for security teams evaluating skill deployment.

Implementation Details and Code Example

The _compute_risk_score function returns a tuple of (score, severity_band, recommendation), which the public report(state) helper consumes to assemble SARIF outputs, JSON reports, and terminal displays.

from skillspector.nodes.report import _compute_risk_score
from skillspector.models import Finding

# Example findings with varying severity

findings = [
    Finding(rule_id="R1", severity="HIGH", ...),      # +25 points

    Finding(rule_id="R2", severity="LOW", ...),       # +5 points

    Finding(rule_id="R3", severity="CRITICAL", ...),  # +50 points

]

# Flag indicating presence of executable scripts

has_executable_scripts = True

# Calculate risk metrics

score, severity, recommendation = _compute_risk_score(findings, has_executable_scripts)

print(f"Score: {score}")                    # Output: 104 → capped to 100

print(f"Severity band: {severity}")         # Output: "CRITICAL"

print(f"Recommendation: {recommendation}")  # Output: "DO_NOT_INSTALL"

The report(state) function internally calls this logic at lines 63-66 of src/skillspector/nodes/report.py, integrating the computed values into the final analysis artifact.

Summary

  • Weighted scoring: Critical findings contribute 50 points, High 25, Medium 10, and Low 5.
  • Executable multiplier: Skills containing scripts receive a 1.3x multiplier on the base score.
  • Normalization: All scores are clamped to the 0-100 range.
  • Threshold bands: Scores map to severity bands at thresholds 81, 51, 21, and 0.
  • Actionable output: The algorithm returns a tuple containing the numeric score, severity band, and installation recommendation used throughout the reporting pipeline.

Frequently Asked Questions

How does SkillSpector calculate the final risk score?

SkillSpector sums base points from all findings (Critical=50, High=25, Medium=10, Low=5), applies a 1.3x multiplier if executable scripts are present, then clamps the result to 0-100. This algorithm is implemented in the _compute_risk_score function within src/skillspector/nodes/report.py.

What is the executable script multiplier in SkillSpector?

The executable script multiplier is a 1.3x adjustment applied to the subtotal score when the skill contains executable scripts (has_executable_scripts=True). This factor increases the risk score for skills capable of arbitrary code execution, as implemented in lines 92-94 of the report node.

How are risk severity bands determined?

Risk severity bands use ordered thresholds defined in _RISK_SEVERITY_BANDS: scores ≥81 become CRITICAL, ≥51 become HIGH, ≥21 become MEDIUM, and ≥0 become LOW. The algorithm selects the first band where the threshold is less than or equal to the computed score.

Where is the risk scoring logic implemented in the codebase?

The core risk scoring logic resides in src/skillspector/nodes/report.py, specifically in the _compute_risk_score function (lines 82-95). Supporting structures like the Finding model are defined in src/skillspector/models.py, which provides the severity annotations driving the calculation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →