# How dcg Performs Heredoc and Inline Script Extraction Using AST Scanning

> Learn how dcg uses AST scanning to extract heredoc and inline scripts, employing regex, content extraction, and ast-grep for safe shell command execution.

- Repository: [Jeff Emanuel/destructive_command_guard](https://github.com/Dicklesworthstone/destructive_command_guard)
- Tags: internals
- Published: 2026-07-21

---

**Destructive Command Guard (dcg) protects shell commands by scanning for heredoc operators and inline script flags in a three-tier pipeline that uses regex-based trigger detection, bounded content extraction, and language-aware AST pattern matching via ast-grep-core to identify destructive API calls.**

The `destructive_command_guard` (dcg) crate prevents agents from executing dangerous commands by deeply inspecting embedded scripts. When a command contains heredoc operators like `<<` or inline flags like `python -c`, dcg uses a carefully orchestrated extraction process followed by structural AST scanning to detect risky API calls that simple regex patterns would miss.

## The Three-Tier Extraction and Analysis Pipeline

The dcg pipeline is deliberately split so that cheap, allocation-free checks can reject safe commands quickly, while heavyweight structural analysis only runs when a potential risk is detected. The system processes every command through three sequential tiers with strict latency budgets.

### Tier 1: Fast Trigger Detection

The first tier spots any heredoc operator (`<<`, `<<<`, `<<-`, `<<~`) or interpreter flags that might introduce an inline script. In [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs) (lines 46-62), a pre-compiled **RegexSet** holds 17 patterns covering `python -c`, `bash -c`, PowerShell `-Command`, `cmd /c`, and similar variants.

The scanner also runs a hand-written parser, `contains_active_heredoc_operator`, that searches for the `<<` operator in a zero-allocation pass. This tier typically executes in **less than 10 µs** for non-matches and under 100 µs when a trigger is found. If `check_triggers` returns `TriggerResult::NoTrigger`, the command bypasses all further analysis.

### Tier 2: Bounded Content Extraction

When triggers are detected, the `extract_content` function pulls the actual script text from the command string while guarding against resource exhaustion. This stage uses compiled regexes such as `HERESTRING_SINGLE_QUOTE`, `INLINE_SCRIPT_DOUBLE_QUOTE`, and `CMD_INLINE_SCRIPT` to parse here-strings, inline-script flags, and classic heredocs.

The extractor respects `ExtractionLimits` defined in [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs), enforcing:
- Maximum 1 MiB content size
- Maximum 10,000 lines
- Maximum 10 heredocs per command
- 50 ms timeout deadline

Binary content detection via `check_binary_content` aborts extraction if null bytes or high non-printable ratios are detected. If limits are exceeded, the system *fail-opens* (allows the command) but logs the `SkipReason` for audit purposes. Typical extraction completes in **under 1 ms**.

### Tier 3: Structural AST Analysis

The final tier analyzes extracted scripts using `AstMatcher` from [`src/ast_matcher.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/ast_matcher.rs). This component pre-compiles destructive patterns for supported languages: **Python**, **JavaScript**, **TypeScript**, **Ruby**, **Bash**, **Go**, and **Php**.

The matcher uses **ast-grep-core** (`AstGrep`) to parse the script into an AST, then walks the tree using `root.find_all(&compiled.pattern)`. Unlike regexes, this detects indirect executions such as:

```javascript
const cp = require('child_process');
cp.execSync('rm -rf /');

```

A hard timeout of **20 ms** (5 seconds in test configurations) guarantees the hook never exceeds its deadline. If the language lacks an ast-grep grammar (e.g., `Perl`), the system falls back to `scan_executing_sink_fallback`, a lightweight regex scan for known destructive function calls like `execSync`, `os.system`, or Ruby `system`.

## Language Detection and Normalization

Before AST analysis begins, `ScriptLanguage::detect` infers the script language from three signals in descending priority:
1. **Command-prefix** (`from_command`): e.g., `python3 -c …`
2. **Shebang** (`from_shebang`): e.g., `#!/usr/bin/env ruby`
3. **Content heuristics** (`from_content`): lightweight scans of the first few lines

The [`src/normalize.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/normalize.rs) module prepares command strings by removing wrapper prefixes and stripping paths before they reach the heredoc pipeline. The detected language drives both delimiter selection during extraction and the AST matcher configuration via `script_language_to_ast_lang`.

## Safety Mechanisms and Timeouts

Both Tier 2 and Tier 3 enforce strict timeout policies using `record_timeout_if_needed` and `run_ast_match_with_timeout`. These functions monitor a shared deadline; on expiry, the command is allowed but emits a warning. This fail-open design ensures that performance degradation never blocks legitimate agent operations.

The `check_binary_content` guard prevents the system from attempting to parse binary data as source code, while `ExtractionLimits` provide deterministic resource bounds regardless of input size.

## Practical Code Examples

Detecting triggers in a command string:

```rust
use destructive_command_guard::heredoc::{check_triggers, TriggerResult};

assert_eq!(check_triggers("git status"), TriggerResult::NoTrigger);
assert_eq!(check_triggers("cat <<EOF"), TriggerResult::Triggered);
assert_eq!(check_triggers("python -c 'import os'"), TriggerResult::Triggered);

```

Extracting content from a heredoc with limits:

```rust
use destructive_command_guard::heredoc::{extract_content, ExtractionLimits};

let cmd = r#"cat <<'EOS'
#!/usr/bin/env python3
import os; os.system("rm -rf /")
EOS
"#;
let result = extract_content(cmd, &ExtractionLimits::default());
match result {
    ExtractionResult::Extracted(contents) => {
        for e in contents {
            println!("Language: {:?}, Body length: {}", e.language, e.content.len());
        }
    }
    _ => println!("No extractable content"),
}

```

Running the AST matcher against extracted Python code:

```rust
use destructive_command_guard::ast_matcher::AstMatcher;
use destructive_command_guard::heredoc::ScriptLanguage;

let matcher = AstMatcher::default();
let script = r#"import os; os.system("rm -rf /")"#;
if let Some(m) = matcher.has_blocking_match(script, ScriptLanguage::Python) {
    println!("Blocked by rule {}: {}", m.rule_id, m.reason);
}

```

## Summary

- **Three-tier architecture**: Trigger detection ([`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs)), bounded extraction (`extract_content`), and AST matching ([`src/ast_matcher.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/ast_matcher.rs)) provide defense in depth.
- **Performance budgets**: Sub-100 µs fast-path for safe commands, with 50 ms and 20 ms timeouts for extraction and AST analysis respectively.
- **Language support**: Native AST parsing for Python, JavaScript, TypeScript, Ruby, Bash, Go, and Php, with regex fallback for unsupported languages.
- **Fail-open design**: Binary detection, resource limits, and timeout enforcement ensure the system allows commands rather than crashing or hanging when encountering edge cases.
- **Structural analysis**: ast-grep-core enables detection of destructive calls through aliases and indirect references that pure regex patterns cannot reliably identify.

## Frequently Asked Questions

### How does dcg handle unsupported languages like Perl?

When `ScriptLanguage::detect` identifies a language without an ast-grep grammar (such as Perl), dcg falls back to `scan_executing_sink_fallback` in [`src/ast_matcher.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/ast_matcher.rs). This lightweight regex scanner looks for known destructive function calls—like `system`, `exec`, or `qx//`—and blocks them when they contain destructive payloads, providing protection without full AST parsing.

### What happens if content extraction exceeds the 50ms timeout?

The `extract_content` function monitors elapsed time via `record_timeout_if_needed`. If the 50 ms deadline expires during extraction, the function immediately returns `ExtractionResult::Skipped` with a timeout reason. Following the fail-open policy, dcg allows the command to execute but logs the timeout event for security auditing.

### Why does dcg use AST matching instead of pure regex?

Pure regexes cannot reliably detect indirect executions where dangerous calls are aliased or passed as arguments. For example, `const cp = require('child_process'); cp.execSync('rm -rf /')` requires understanding that `cp.execSync` resolves to the `child_process.execSync` sink. The `AstMatcher` uses ast-grep-core to parse the code structure, identifying the callee regardless of variable names and examining actual argument values—including list arguments like `subprocess.run(["sh","-c","rm -rf /"])`—to eliminate false negatives.

### Which source files implement the heredoc extraction pipeline?

The pipeline spans three primary files: [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs) implements Tier 1 trigger detection and Tier 2 bounded extraction; [`src/ast_matcher.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/ast_matcher.rs) contains the Tier 3 `AstMatcher` with pattern compilation and timeout handling; and [`src/evaluator.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/evaluator.rs) orchestrates the three-tier pipeline when processing commands. Language detection logic resides in [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs) (lines 94-166), while [`src/normalize.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/normalize.rs) prepares command strings before they enter the pipeline.