How Behavioral AST Analysis Detects Dangerous Code Patterns in SkillSpector
SkillSpector detects dangerous code patterns by parsing Python files into abstract syntax trees (AST) and walking the node structure to identify risky function calls, dynamic imports, and nested execution chains that could lead to arbitrary code execution.
SkillSpector, NVIDIA's security analysis framework, inspects Python source files by transforming them into abstract syntax trees and applying behavioral rules to identify potentially harmful execution patterns. The core detection engine resides in behavioral_ast.py, where it systematically evaluates function calls against curated threat signatures to prevent runtime compromise.
Parsing and Traversing the Abstract Syntax Tree
The analysis workflow begins in src/skillspector/nodes/analyzers/behavioral_ast.py. The analyzer first validates that the target file size is within the configurable limit defined by MAX_FILE_BYTES in src/skillspector/nodes/analyzers/static_runner.py.
It then uses Python's built-in ast.parse to convert source code into a tree structure, followed by ast.walk(tree) to iterate over every node. The engine specifically isolates ast.Call nodes (function invocations) for detailed inspection, ignoring other syntactic elements.
Eight Core Detection Rules
SkillSpector categorizes dangerous patterns into eight distinct rule identifiers (AST1 through AST8), each targeting specific execution risks. The resolve_call_name helper function from src/skillspector/nodes/analyzers/common.py extracts fully-qualified names (e.g., os.system, subprocess.run) from attribute chains to match against predefined threat signatures.
Direct Dangerous Builtins (AST1, AST2, AST3, AST6)
The analyzer checks call names against the _DANGEROUS_BUILTINS set to flag fundamental code execution risks:
- AST1: Direct
exec()calls capable of executing arbitrary Python code - AST2: Direct
eval()calls that evaluate string expressions - AST3: Dynamic imports via
__import__()that bypass static import analysis - AST6: Use of
compile()to create executable code objects from strings
System and Process Execution (AST4, AST5)
Operating system interaction patterns trigger AST4 and AST5 findings:
- AST4: Calls to the
subprocessmodule that launch new processes (e.g.,subprocess.run,subprocess.Popen), matched against the_SUBPROCESS_CALLSset - AST5: Execution via the
osmodule's exec-family functions (e.g.,os.system,os.execve), identified through the_OS_EXEC_CALLSset
Dynamic Attribute Access (AST7)
AST7 detects getattr() invocations where the attribute parameter is not a constant string literal. This pattern indicates potential method invocation based on user-controlled input, allowing attackers to access restricted object methods.
Dangerous Execution Chains (AST8)
AST8 identifies chained dangerous calls where exec, eval, or compile wrap secondary risky sources. This includes patterns like exec(eval("...")) or eval(subprocess.run(...)), representing critical severity escalation vectors.
Execution Chain Analysis
The _contains_dangerous_source function in behavioral_ast.py performs deep inspection of nested calls. When the analyzer encounters exec, eval, or compile, it examines the first argument (ast_node.args[0]) to determine if that argument itself contains calls to subprocess modules, OS execution functions, __import__, or obfuscation loaders like base64 decoders.
If the nested argument contains any dangerous source, the analyzer emits an AST8 finding with critical severity, indicating a multi-layered execution chain capable of bypassing superficial static checks.
From Detection to Report
Each match creates an AnalyzerFinding object defined in src/skillspector/models.py. These findings encapsulate:
- The rule identifier (AST1 through AST8)
- Severity levels mapped in
_RULE_SEVERITIES - Confidence values from
_RULE_CONFIDENCES - Location metadata (file path, start/end lines)
- Source context snippets extracted via
get_source_segment
The analyzer_finding_to_finding function in static_runner.py converts these internal representations into public Finding objects, which the SkillSpector pipeline processes through src/skillspector/graph.py alongside other analysis results.
Practical Code Examples
Direct Dangerous Builtin
# vulnerable.py
exec(user_input) # flagged as AST1 (high severity)
This produces a finding with the message "exec() call detected".
Subprocess Execution
import subprocess
subprocess.run("rm -rf /", shell=True) # flagged as AST4 (medium severity)
Dangerous Execution Chain
import os
exec(eval("os.system('whoami')")) # flagged as AST8 (critical severity)
The outer exec triggers AST8 because _contains_dangerous_source discovers the nested os.system call within the eval argument.
Dynamic Attribute Access
obj = SomeClass()
method_name = get_user_input()
getattr(obj, method_name)() # flagged as AST7 (low severity)
Summary
- Behavioral AST analysis in SkillSpector converts Python source into abstract syntax trees using
ast.parseand walks the structure withast.walkto examine every function call. - Eight specific rules (AST1 through AST8) categorize risks ranging from direct
exec()calls to chained execution patterns involving nested dangerous functions. - The
resolve_call_namehelper extracts fully-qualified function names to match against sets including_DANGEROUS_BUILTINS,_SUBPROCESS_CALLS, and_OS_EXEC_CALLS. - AST8 specifically detects dangerous execution chains by analyzing the first argument of
exec,eval, orcompilefor nested risky sources. - Findings are converted from
AnalyzerFindingobjects to genericFindingmodels viaanalyzer_finding_to_findinginstatic_runner.py, preserving severity, confidence, and source context.
Frequently Asked Questions
What is the difference between AST1 and AST8 findings?
AST1 flags direct exec() calls where the function itself is invoked, while AST8 detects when exec, eval, or compile wrap a secondary dangerous source (such as exec(eval("os.system(...)"))). AST8 represents chained execution with critical severity, whereas AST1 indicates single-layer builtin usage.
How does SkillSpector handle import aliases in behavioral analysis?
While the current AST rules use resolve_call_name to extract dotted names like os.system, common.py provides utilities including _build_import_aliases and build_type_map that enable future extensions to resolve aliased imports (e.g., import subprocess as sp; sp.run(...)).
What severity levels does SkillSpector assign to dangerous code patterns?
Severity values are defined in _RULE_SEVERITIES within behavioral_ast.py. Direct builtins like exec() and eval() typically receive high severity, subprocess and OS calls receive medium severity, dynamic attribute access receives low severity, and chained dangerous calls (AST8) receive critical severity.
Can SkillSpector detect obfuscated code chains?
Yes, the _contains_dangerous_source function examines arguments to exec, eval, and compile for nested dangerous calls, including potential obfuscation patterns using __import__, base64 decoders, or urllib-based loaders that might hide malicious payloads within seemingly benign outer calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →