How SkillSpector Performs AST-Based Behavioral Analysis on Python Code

SkillSpector performs AST-based behavioral analysis by walking Python abstract syntax trees to detect dangerous execution patterns like exec(), eval(), and subprocess calls, converting findings into structured AnalyzerFinding objects with severity ratings and source context.

NVIDIA's SkillSpector is a static analysis tool that inspects AI-generated code for security vulnerabilities. Its AST-based behavioral analysis module specifically targets dangerous runtime execution patterns by parsing Python source into an abstract syntax tree and applying heuristic rule sets to identify code that could execute arbitrary commands.

Entry Point and File Discovery

The behavioral analysis begins at the analyzer node defined in [src/skillspector/nodes/analyzers/behavioral_ast.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/behavioral_ast.py). The node(state) function receives a SkillspectorState object containing two critical keys: components (the list of discovered file paths) and file_cache (a mapping of paths to source code content).

For each Python file that fits within the size constraints defined by MAX_FILE_BYTES in [static_runner.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/static_runner.py), the node invokes the internal helper _analyze_python(content, path). This design keeps the analysis pipeline modular, allowing the AST walker to process files independently while maintaining access to the full project context.

Parsing and Walking the Abstract Syntax Tree

The analysis core converts source code into an AST using Python's built-in ast module:

tree = ast.parse(content, filename=file_path)

If ast.parse raises a SyntaxError, the analyzer logs the failure and skips the file rather than crashing the entire scan. Once parsed, the walker iterates over every node in the tree using ast.walk(tree) to find all ast.Call instances.

For each call node, the analyzer invokes resolve_call_name(ast_node) from [common.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/nodes/analyzers/common.py). This utility translates complex call expressions into dotted strings like os.system, subprocess.run, or exec, enabling simple string matching against known dangerous patterns.

Detection Rules and Pattern Matching

The behavioral analyzer checks resolved call names against several pre-computed rule sets to identify hazardous constructs.

Dangerous Built-in Functions

The _DANGEROUS_BUILTINS set contains high-risk functions including exec, eval, compile, and __import__. When the walker encounters these, it emits findings with rule IDs AST1 through AST3 and AST6, flagging calls that allow dynamic code execution or unsafe imports.

Process Execution and Subprocess Calls

The analyzer distinguishes between different execution mechanisms using two specialized sets:

  • _SUBPROCESS_CALLS captures subprocess.run, subprocess.Popen, subprocess.check_output, and related methods (rule AST4)
  • _OS_EXEC_CALLS identifies os.system, os.execl, os.posix_spawn, and other OS-level execution functions (rule AST5)

This categorization helps security teams understand whether code uses high-level subprocess wrappers or low-level system calls.

Dynamic Attribute Access

The walker specifically detects getattr calls where the attribute argument is not a constant string literal. This triggers AST7 ("Dynamic attribute access via getattr()"), alerting users to code that could invoke arbitrary methods based on user input.

Dangerous Execution Chains

When exec, eval, or compile wraps another dangerous call—such as eval("subprocess.run(['rm', '-rf', '/']")—the analyzer invokes _contains_dangerous_source to walk the argument subtree. If nested dangerous calls are detected, it emits AST8 ("Dangerous execution chain"), capturing complex obfuscation techniques that simple string matching would miss.

Building and Emitting Findings

Each detection invokes the _emit(rule_id, lineno, end_lineno, msg_override=None) helper to construct an AnalyzerFinding object. Defined in [src/skillspector/models.py](https://github.com/NVIDIA/SkillSpector/blob/main/src/skillspector/models.py), this dataclass captures:

  • Rule ID (e.g., AST1, AST4)
  • Severity from _RULE_SEVERITIES (HIGH, MEDIUM, etc.)
  • Confidence from _RULE_CONFIDENCES
  • Location with file path and line numbers
  • Context extracted via get_context_from_lines in common.py (3-line code snippets)
  • Matched text truncated to 200 characters using get_source_segment

The static_runner.py file then converts these internal AnalyzerFinding objects into public Finding types via analyzer_finding_to_finding, enriching them with remediation data, tags, and user-friendly categories for the final report.

Integration with the Static Analysis Pipeline

The AST-based behavioral analysis integrates into SkillSpector's broader static analysis pipeline through four stages:

  1. File Discovery: The graph.Scanner collects all project files and populates state["file_cache"] with source content
  2. Pattern Execution: run_static_patterns in static_runner.py iterates over analyzer modules, including behavioral_ast
  3. State Mutation: The node function returns {"findings": all_findings}, updating the shared graph state
  4. Report Generation: Downstream nodes aggregate findings and render outputs in SARIF or CLI formats

This architecture ensures that AST-based behavioral analysis runs alongside other static pattern checks while maintaining consistent output formatting.

Practical Code Examples

Detecting Simple Dangerous Calls

The following code triggers AST1 because it uses exec:


# unsafe.py

exec("print('Hello world')")

SkillSpector reports:


[AST1] exec() call detected
Location: unsafe.py:1
Context:
exec("print('Hello world')")

Identifying Dynamic Imports

Using __import__ dynamically generates AST3:

mod = __import__('os')

Result:


[AST3] Dynamic import via __import__()
Location: unsafe.py:1

Chaining Dangerous Calls

Nested dangerous calls trigger AST8:

eval("subprocess.run(['rm', '-rf', '/'])")

Output:


[AST8] Dangerous chain: eval() wrapping subprocess.run
Location: unsafe.py:1

Running the Analyzer from CLI

Execute the behavioral AST analyzer standalone:

skill-spector scan ./my_project \
  --analyzers behavioral_ast \
  --output sarif

This command scans all Python files, applies the AST walker, and outputs a SARIF report containing rule IDs, severity levels, and source code contexts.

Summary

  • AST-based behavioral analysis in SkillSpector walks Python syntax trees to find dangerous execution patterns.
  • The analyzer lives in behavioral_ast.py and processes files from the SkillspectorState object populated by the graph scanner.
  • Detection relies on resolve_call_name to match ast.Call nodes against sets of dangerous builtins, subprocess calls, and OS execution functions.
  • Rule IDs AST1-AST8 cover specific risks from simple exec() calls to complex nested execution chains.
  • Findings are emitted as AnalyzerFinding objects with severity, confidence, and 3-line code context, then converted to public Finding types in static_runner.py.
  • The module integrates into the static pipeline via run_static_patterns and supports SARIF and CLI output formats.

Frequently Asked Questions

What is AST-based behavioral analysis in SkillSpector?

AST-based behavioral analysis is a static analysis technique where SkillSpector parses Python source code into an abstract syntax tree and walks the tree to identify dangerous runtime behaviors. Unlike simple regex pattern matching, this method understands code structure, enabling detection of nested dangerous calls and dynamic attribute access that text-based scanners might miss.

Which dangerous patterns does the behavioral AST analyzer detect?

According to the source code in behavioral_ast.py, the analyzer detects eight specific patterns: exec() (AST1), eval() (AST2), compile() (AST6), __import__() (AST3), subprocess module calls (AST4), OS execution family calls (AST5), dynamic getattr() usage (AST7), and dangerous execution chains where builtins wrap other risky calls (AST8).

How does SkillSpector handle Python files with syntax errors?

The analyzer wraps ast.parse() in a try-except block that catches SyntaxError. When parsing fails, the tool logs the failure and skips the file, allowing the scan to continue processing other components rather than terminating the entire analysis pipeline.

What information is included in an AST-based behavioral finding?

Each finding includes the rule ID (e.g., AST4), severity level (HIGH, MEDIUM, etc.), confidence score, exact file location with line numbers, a 3-line source code context snippet extracted via get_context_from_lines, and the first 200 characters of the matched source segment. This metadata structure is defined in models.py and enriched during the conversion to the public Finding type.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →