How Behavioral AST Analysis Identifies Dangerous Code Execution Patterns in SkillSpector

SkillSpector uses behavioral AST analysis to detect dangerous code execution patterns by parsing Python source files into abstract syntax trees and traversing them to identify risky function calls, dynamic imports, and nested execution chains.

NVIDIA's SkillSpector is a security analysis framework that inspects Python code for potentially harmful execution behaviors. The behavioral AST analysis node, implemented in src/skillspector/nodes/analyzers/behavioral_ast.py, serves as the primary detection engine that converts source code into inspectable structures and applies rule-based pattern matching to flag code capable of arbitrary execution or system compromise.

The Behavioral AST Analysis Pipeline

The analysis begins by ingesting Python source files under a configurable size limit. In src/skillspector/nodes/analyzers/static_runner.py, the MAX_FILE_BYTES constant defines the file size threshold to prevent resource exhaustion on excessively large files.

The analyzer calls Python's built-in ast.parse method to convert source code into an abstract syntax tree. It then invokes ast.walk(tree) to iterate over every node in the tree, filtering specifically for ast.Call nodes that represent function invocations. This targeted approach ensures the behavioral AST analysis focuses exclusively on execution points where dangerous behavior can manifest.

Detecting Dangerous Patterns with Behavioral AST Analysis

For each function call discovered during the AST traversal, the analyzer uses resolve_call_name from src/skillspector/nodes/analyzers/common.py to extract the fully-qualified name (e.g., os.system, subprocess.run). It then compares this resolved name against predefined rule sets to categorize the risk:

  • _DANGEROUS_BUILTINS: Flags direct calls to exec(), eval(), compile(), and __import__()
  • _SUBPROCESS_CALLS: Identifies process spawning via the subprocess module
  • _OS_EXEC_CALLS: Detects operating system execution functions like os.system
  • Dynamic attribute access: Spots getattr() usage with non-constant attribute names

Each rule maps to a specific rule identifier (AST1 through AST8) with associated severity and confidence values defined in _RULE_SEVERITIES and _RULE_CONFIDENCES.

AST1-AST3: Dangerous Built-in Functions

AST1 triggers on direct exec() calls, while AST2 detects eval() invocations. AST3 identifies dynamic imports via __import__(). These rules catch the most straightforward code execution vectors that allow runtime evaluation of strings as Python code.


# vulnerable.py

exec(user_input)          # flagged as AST1 (high severity)

AST4-AST6: System and Process Execution

AST4 monitors calls to the subprocess module that launch new processes, such as subprocess.run or subprocess.Popen. AST5 tracks the os module's exec-family functions including os.system and os.execve. AST6 flags usage of compile() to create executable code objects, which often precede dynamic execution.

import subprocess

subprocess.run("rm -rf /", shell=True)   # flagged as AST4 (medium severity)

AST7: Dynamic Attribute Resolution

AST7 identifies risky dynamic behavior where getattr() retrieves an attribute name from user-controlled input rather than a constant string, enabling arbitrary method invocation:

obj = SomeClass()
method_name = get_user_input()
getattr(obj, method_name)()   # flagged as AST7 (low severity)

Identifying Dangerous Execution Chains

AST8 represents the most critical severity level in behavioral AST analysis, detecting when dangerous calls are nested within each other. The analyzer examines the first argument (ast_node.args[0]) of exec, eval, or compile calls using the _contains_dangerous_source helper. If this argument contains a call to subprocess, os exec, __import__, or similar risky sources, the system flags a dangerous execution chain:

import os

exec(eval("os.system('whoami')"))   # flagged as AST8 (critical severity)

The outer exec triggers AST8 because _contains_dangerous_source discovers the nested os.system call, revealing a multi-layered attack vector.

Emitting Security Findings

When the behavioral AST analysis identifies a match, it creates an AnalyzerFinding object defined in src/skillspector/models.py. Each finding includes:

  • The rule identifier (AST1-AST8)
  • A human-readable message (e.g., "exec() call detected" or "Dangerous chain: exec() wrapping subprocess.run")
  • Severity and confidence classifications
  • Precise location information (file path, start/end line numbers)
  • A source code snippet via get_source_segment and surrounding context via get_context_from_lines

The analyzer_finding_to_finding function in static_runner.py converts these analyzer-specific objects into generic Finding models that integrate with the broader SkillSpector pipeline orchestrated by src/skillspector/graph.py.

Extensibility and Alias Resolution

While the current behavioral AST analysis focuses on direct calls to dangerous functions, src/skillspector/nodes/analyzers/common.py provides infrastructure for future enhancements. The _build_import_aliases and build_type_map utilities enable resolution of import aliases (e.g., import subprocess as sp; sp.run(...)), allowing the analysis to handle obfuscated imports as the detection logic evolves.

Summary

  • Behavioral AST analysis in SkillSpector converts Python source into an AST using ast.parse and traverses it with ast.walk to examine ast.Call nodes.
  • The system checks resolved function names against _DANGEROUS_BUILTINS, _SUBPROCESS_CALLS, and _OS_EXEC_CALLS to identify eight distinct dangerous patterns (AST1-AST8).
  • AST8 specifically detects dangerous execution chains where exec, eval, or compile wrap secondary risky sources like subprocess calls or dynamic imports.
  • Findings are enriched with source context and location data, then converted to standard Finding objects via analyzer_finding_to_finding in static_runner.py.
  • The analysis respects resource limits through MAX_FILE_BYTES and leverages resolve_call_name from common.py for accurate function identification.

Frequently Asked Questions

How does SkillSpector's behavioral AST analysis parse Python files without executing them?

SkillSpector uses Python's built-in ast.parse method to generate an abstract syntax tree from source code, allowing static inspection of the code structure without runtime execution. The analyzer enforces a MAX_FILE_BYTES limit defined in static_runner.py to prevent memory exhaustion, then walks the tree using ast.walk(tree) to identify potentially dangerous patterns before any code executes.

What distinguishes AST8 from other detection rules in the behavioral AST analysis?

AST8 specifically identifies dangerous execution chains where high-risk functions like exec, eval, or compile receive arguments that themselves contain calls to dangerous sources. While AST1-AST7 detect direct usage of risky functions, AST8 triggers when these functions wrap secondary threats like subprocess.run or os.system, creating multi-layered attack vectors that bypass simple pattern matching.

How does resolve_call_name determine the target of a function call in the AST?

The resolve_call_name helper in common.py constructs fully-qualified names (e.g., os.system, subprocess.Popen) by analyzing ast.Call nodes and their attribute chains. It handles both simple names and complex attribute access patterns, enabling the behavioral AST analysis to accurately match calls against rule sets regardless of whether functions are called directly or via module attributes.

Can SkillSpector detect dangerous code when imports use aliases?

The current behavioral AST analysis focuses on direct calls to dangerous functions. However, common.py includes infrastructure such as _build_import_aliases and build_type_map designed to support alias resolution in future implementations. For now, the system primarily detects patterns like import os; os.system() rather than import os as o; o.system(), though the node architecture supports extending this capability.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →