How SkillSpector Performs AST-Based Behavioral Analysis on Python Code
SkillSpector performs AST-based behavioral analysis by walking Python abstract syntax trees to detect dangerous execution patterns like exec(), eval(), and subprocess calls, converting findings into structured AnalyzerFinding objects with severity ratings and source context.
NVIDIA's SkillSpector is a static analysis tool that inspects AI-generated code for security vulnerabilities. Its AST-based behavioral analysis module specifically targets dangerous runtime execution patterns by parsing Python source into an abstract syntax tree and applying heuristic rule sets to identify code that could execute arbitrary commands.
Entry Point and File Discovery
The behavioral analysis begins at the analyzer node defined in [src/skillspector/nodes/analyzers/behavioral_ast.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/behavioral_ast.py). The node(state) function receives a SkillspectorState object containing two critical keys: components (the list of discovered file paths) and file_cache (a mapping of paths to source code content).
For each Python file that fits within the size constraints defined by MAX_FILE_BYTES in [static_runner.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/static_runner.py), the node invokes the internal helper _analyze_python(content, path). This design keeps the analysis pipeline modular, allowing the AST walker to process files independently while maintaining access to the full project context.
Parsing and Walking the Abstract Syntax Tree
The analysis core converts source code into an AST using Python's built-in ast module:
tree = ast.parse(content, filename=file_path)
If ast.parse raises a SyntaxError, the analyzer logs the failure and skips the file rather than crashing the entire scan. Once parsed, the walker iterates over every node in the tree using ast.walk(tree) to find all ast.Call instances.
For each call node, the analyzer invokes resolve_call_name(ast_node) from [common.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/common.py). This utility translates complex call expressions into dotted strings like os.system, subprocess.run, or exec, enabling simple string matching against known dangerous patterns.
Detection Rules and Pattern Matching
The behavioral analyzer checks resolved call names against several pre-computed rule sets to identify hazardous constructs.
Dangerous Built-in Functions
The _DANGEROUS_BUILTINS set contains high-risk functions including exec, eval, compile, and __import__. When the walker encounters these, it emits findings with rule IDs AST1 through AST3 and AST6, flagging calls that allow dynamic code execution or unsafe imports.
Process Execution and Subprocess Calls
The analyzer distinguishes between different execution mechanisms using two specialized sets:
_SUBPROCESS_CALLScapturessubprocess.run,subprocess.Popen,subprocess.check_output, and related methods (rule AST4)_OS_EXEC_CALLSidentifiesos.system,os.execl,os.posix_spawn, and other OS-level execution functions (rule AST5)
This categorization helps security teams understand whether code uses high-level subprocess wrappers or low-level system calls.
Dynamic Attribute Access
The walker specifically detects getattr calls where the attribute argument is not a constant string literal. This triggers AST7 ("Dynamic attribute access via getattr()"), alerting users to code that could invoke arbitrary methods based on user input.
Dangerous Execution Chains
When exec, eval, or compile wraps another dangerous call—such as eval("subprocess.run(['rm', '-rf', '/']")—the analyzer invokes _contains_dangerous_source to walk the argument subtree. If nested dangerous calls are detected, it emits AST8 ("Dangerous execution chain"), capturing complex obfuscation techniques that simple string matching would miss.
Building and Emitting Findings
Each detection invokes the _emit(rule_id, lineno, end_lineno, msg_override=None) helper to construct an AnalyzerFinding object. Defined in [src/skillspector/models.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/models.py), this dataclass captures:
- Rule ID (e.g.,
AST1,AST4) - Severity from
_RULE_SEVERITIES(HIGH, MEDIUM, etc.) - Confidence from
_RULE_CONFIDENCES - Location with file path and line numbers
- Context extracted via
get_context_from_linesincommon.py(3-line code snippets) - Matched text truncated to 200 characters using
get_source_segment
The static_runner.py file then converts these internal AnalyzerFinding objects into public Finding types via analyzer_finding_to_finding, enriching them with remediation data, tags, and user-friendly categories for the final report.
Integration with the Static Analysis Pipeline
The AST-based behavioral analysis integrates into SkillSpector's broader static analysis pipeline through four stages:
- File Discovery: The
graph.Scannercollects all project files and populatesstate["file_cache"]with source content - Pattern Execution:
run_static_patternsinstatic_runner.pyiterates over analyzer modules, includingbehavioral_ast - State Mutation: The
nodefunction returns{"findings": all_findings}, updating the shared graph state - Report Generation: Downstream nodes aggregate findings and render outputs in SARIF or CLI formats
This architecture ensures that AST-based behavioral analysis runs alongside other static pattern checks while maintaining consistent output formatting.
Practical Code Examples
Detecting Simple Dangerous Calls
The following code triggers AST1 because it uses exec:
# unsafe.py
exec("print('Hello world')")
SkillSpector reports:
[AST1] exec() call detected
Location: unsafe.py:1
Context:
exec("print('Hello world')")
Identifying Dynamic Imports
Using __import__ dynamically generates AST3:
mod = __import__('os')
Result:
[AST3] Dynamic import via __import__()
Location: unsafe.py:1
Chaining Dangerous Calls
Nested dangerous calls trigger AST8:
eval("subprocess.run(['rm', '-rf', '/'])")
Output:
[AST8] Dangerous chain: eval() wrapping subprocess.run
Location: unsafe.py:1
Running the Analyzer from CLI
Execute the behavioral AST analyzer standalone:
skill-spector scan ./my_project \
--analyzers behavioral_ast \
--output sarif
This command scans all Python files, applies the AST walker, and outputs a SARIF report containing rule IDs, severity levels, and source code contexts.
Summary
- AST-based behavioral analysis in SkillSpector walks Python syntax trees to find dangerous execution patterns.
- The analyzer lives in
behavioral_ast.pyand processes files from theSkillspectorStateobject populated by the graph scanner. - Detection relies on
resolve_call_nameto matchast.Callnodes against sets of dangerous builtins, subprocess calls, and OS execution functions. - Rule IDs AST1-AST8 cover specific risks from simple
exec()calls to complex nested execution chains. - Findings are emitted as
AnalyzerFindingobjects with severity, confidence, and 3-line code context, then converted to publicFindingtypes instatic_runner.py. - The module integrates into the static pipeline via
run_static_patternsand supports SARIF and CLI output formats.
Frequently Asked Questions
What is AST-based behavioral analysis in SkillSpector?
AST-based behavioral analysis is a static analysis technique where SkillSpector parses Python source code into an abstract syntax tree and walks the tree to identify dangerous runtime behaviors. Unlike simple regex pattern matching, this method understands code structure, enabling detection of nested dangerous calls and dynamic attribute access that text-based scanners might miss.
Which dangerous patterns does the behavioral AST analyzer detect?
According to the source code in behavioral_ast.py, the analyzer detects eight specific patterns: exec() (AST1), eval() (AST2), compile() (AST6), __import__() (AST3), subprocess module calls (AST4), OS execution family calls (AST5), dynamic getattr() usage (AST7), and dangerous execution chains where builtins wrap other risky calls (AST8).
How does SkillSpector handle Python files with syntax errors?
The analyzer wraps ast.parse() in a try-except block that catches SyntaxError. When parsing fails, the tool logs the failure and skips the file, allowing the scan to continue processing other components rather than terminating the entire analysis pipeline.
What information is included in an AST-based behavioral finding?
Each finding includes the rule ID (e.g., AST4), severity level (HIGH, MEDIUM, etc.), confidence score, exact file location with line numbers, a 3-line source code context snippet extracted via get_context_from_lines, and the first 200 characters of the matched source segment. This metadata structure is defined in models.py and enriched during the conversion to the public Finding type.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →