# How SkillSpector's Baseline and False-Positive Suppression Work

> Discover how SkillSpector's baseline and false-positive suppression system leverages glob rules and cryptographic fingerprints to deliver accurate scan results. Learn more.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: internals
- Published: 2026-06-25

---

**SkillSpector uses a dual-layer suppression system that combines human-readable glob rules with cryptographically secure fingerprints to identify and filter known false positives from scan results before they reach the final report.**

NVIDIA's SkillSpector employs a baseline file to instruct the report generation step which findings should be ignored. According to the source code in [`src/skillspector/suppression.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/suppression.py), the baseline implements two complementary mechanisms that work together to keep results clean while maintaining flexibility for developers.

## The Dual-Mechanism Suppression Architecture

The suppression system relies on two distinct approaches defined in [`src/skillspector/suppression.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/suppression.py): **glob-based rules** for flexible pattern matching and **cryptographic fingerprints** for exact identification.

### Glob-Based Rules for Flexible Pattern Matching

The `rules` mechanism uses human-written patterns that match a finding's rule ID, file path, and/or message. In [`src/skillspector/suppression.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/suppression.py), the `SuppressionRule` class (lines 106-126) provides the `matches` method that compares these fields against each finding. If every field specified in a rule matches the finding, the finding is suppressed.

This mechanism is tolerant to small changes such as added comments or line-number shifts, making it ideal for policy-style suppression like ignoring all `SQP-1` findings in any file.

### SHA-256 Fingerprints for Exact Identification

The `fingerprints` mechanism uses machine-generated hashes that uniquely identify a finding based on rule ID, file, line range, and message. The `finding_fingerprint()` function (lines 85-104) computes a SHA-256 hash of the concatenated fields, storing these in the `Baseline.fingerprints` dictionary.

When a scan produces the same fingerprint, the finding is considered a known false-positive and is dropped from results. This provides reproducible, exact suppression useful for CI pipelines where only new issues should surface.

## The Baseline Processing Pipeline

The suppression workflow follows a three-stage pipeline implemented across [`src/skillspector/suppression.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/suppression.py) and [`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py).

### Loading and Parsing the Baseline File

The `load_baseline()` function parses YAML or JSON baseline files, with `baseline_from_dict()` performing the heavy lifting (lines 109-124). This builds a `Baseline` object containing both the `SuppressionRule` list and the fingerprint dictionary.

### Partitioning Findings During Report Generation

During the reporting phase, the `report()` function in [`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py) (lines 58-62) retrieves the baseline from the execution state and passes it to `partition_findings()` along with the list of findings.

The `partition_findings()` function (lines 126-146) iterates over each `Finding`, querying the baseline via `Baseline.reason_for()`. If a rule matches, the reason is returned; otherwise, the function checks the fingerprint map. Suppressed findings become `SuppressedFinding` objects, while non-suppressed findings proceed to scoring and SARIF output.

### Controlling Suppressed Finding Visibility

By default, suppressed findings are omitted from the risk score and SARIF payload. However, the `report()` function checks for the `--show-suppressed` flag (lines 43-46 and 54-58) to optionally display a markdown table of suppressed findings in the terminal output.

## Creating and Managing Baselines

SkillSpector provides both manual and automated methods for creating baselines.

### Sample Baseline Configuration

A baseline file supports both rules and fingerprints:

```yaml
version: 1
rules:
  - id: "SQP-1"
    reason: "Trigger-phrase description, not a vulnerability"
  - id: "SSD-2"
    path: "*deploy-topology*/SKILL.md"
    message: "*run the exploit*"
    reason: "False positive – lab test phrase"
fingerprints:
  - hash: "sha256:1a2b3c4d5e6f7081"
    rule_id: "SDI-2"
    file: "baas-build-analysis/SKILL.md"
    reason: "Accepted 2026-06-19 – first-party env detection"

```

### Generating Baselines via CLI

You can automatically fingerprint all findings from a scan:

```bash
skillspector scan path/to/skill --baseline baseline.yaml

```

This command internally calls `build_baseline_dict()` (lines 250-266) to produce the fingerprint list, then writes the file using `dump_baseline()`.

### Programmatic Baseline Usage

For custom integrations, load and apply baselines programmatically:

```python
from pathlib import Path
from skillspector.suppression import load_baseline, partition_findings
from skillspector.models import Finding

baseline = load_baseline(Path("baseline.yaml"))
findings: list[Finding] = [...]

kept, suppressed = partition_findings(findings, baseline)

print(f"Active findings: {len(kept)}")
print(f"Suppressed findings: {len(suppressed)}")
for sf in suppressed:
    print(f"- {sf.finding.rule_id} suppressed because: {sf.reason}")

```

To display suppressed findings in terminal reports, use:

```bash
skillspector report --show-suppressed --baseline baseline.yaml

```

## Why SkillSpector Uses Two Suppression Mechanisms

**Glob rules** provide tolerance to minor code changes like line shifts or comment additions, making them ideal for expressing broad policies such as "ignore all `SQP-1` findings in test files." **Fingerprints** deliver exact, reproducible identification through SHA-256 hashing, ensuring that specific known false positives remain suppressed even as the surrounding code evolves.

Together, these mechanisms allow developers to maintain a clean, incremental baseline while retaining the flexibility to write expressive, human-readable suppression rules.

## Summary

- SkillSpector's baseline system resides in [`src/skillspector/suppression.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/suppression.py) and implements dual suppression mechanisms.
- **Glob-based rules** (`SuppressionRule` class with `matches` method) allow pattern-based suppression using wildcards for rule IDs, paths, and messages.
- **SHA-256 fingerprints** (`finding_fingerprint()`) provide exact matching based on content hashes of rule ID, file, line range, and message.
- The `partition_findings()` function separates findings into active and suppressed sets based on baseline matching.
- The `--show-suppressed` flag in [`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py) controls visibility of suppressed findings in terminal output.
- Baselines can be generated automatically via the CLI (`skillspector scan --baseline`) or created manually as YAML/JSON files.

## Frequently Asked Questions

### How do I create a new baseline file from an existing scan?

Run the scan with the `--baseline` flag to automatically generate fingerprints for all current findings. The CLI calls `build_baseline_dict()` (lines 250-266) to create the fingerprint entries and `dump_baseline()` to write the YAML file. This captures the current state as known good, ensuring future scans only report new issues.

### What is the difference between a rule and a fingerprint in the baseline?

Rules use glob patterns to match findings flexibly across file paths and messages, tolerating minor code changes. Fingerprints use SHA-256 hashes to match exact content strings, providing deterministic suppression for specific findings that are known to be false positives regardless of line number shifts.

### Can I see suppressed findings in my terminal output?

Yes. Pass the `--show-suppressed` flag to the `report` command. This toggles the display branch in [`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py) (lines 43-46) that prints a markdown table of suppressed findings. Without this flag, suppressed findings are excluded from both the terminal output and the SARIF payload.

### Where is the suppression logic implemented in the SkillSpector codebase?

The core implementation lives in [`src/skillspector/suppression.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/suppression.py), which contains the `SuppressionRule` class, `finding_fingerprint()` function, and `partition_findings()` helper. The reporting integration occurs in [`src/skillspector/nodes/report.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/report.py) where the baseline is loaded and applied during the reporting phase.