# How dcg Handles Heredoc Extraction with AST-Based Pattern Matching: A Three-Tier Pipeline

> Discover how dcg uses AST pattern matching and a three-tier pipeline to efficiently extract and analyze heredocs and inline scripts with sub-millisecond latency.

- Repository: [Jeff Emanuel/destructive_command_guard](https://github.com/Dicklesworthstone/destructive_command_guard)
- Tags: internals
- Published: 2026-07-16

---

**dcg (Destructive Command Guard) employs a three-tier detection pipeline—fast regex triggers, bounded content extraction, and AST-based structural analysis—to safely evaluate heredocs and inline scripts with sub-millisecond latency.**

The `destructive_command_guard` repository by Dicklesworthstone implements a high-performance pre-hook for shell commands that must analyze complex heredoc and inline-script payloads without impacting interactive latency. This article examines how the Rust-based `dcg` tool combines regex-based triage with `ast-grep-core` powered pattern matching to extract and analyze embedded code.

## Tier 1: Ultra-Fast Trigger Detection

Implemented in [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs), the [`check_triggers`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs#L39-L48) function serves as the first line of defense. It utilizes a compiled [`RegexSet`](https://docs.rs/regex/latest/regex/struct.RegexSet.html) named `HEREDOC_TRIGGERS` to scan the entire command in a single pass against 17 distinct patterns.

These patterns detect heredoc operators (`<<`, `<<<`), inline script flags (`python -c`, `node -e`), Windows wrappers (`cmd /c`, `powershell -Command`), and pipe-to-interpreter sequences. The function also invokes `contains_active_heredoc_operator`, a hand-written scanner that specifically checks for the `<<` operator to guarantee **zero false-negatives** while permitting false positives—any trigger merely forces progression to Tier 2.

## Tier 2: Bounded Content Extraction

When Tier 1 triggers fire, the [`extract_content`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs#L59-L85) function handles safe parsing with strict resource limits defined by `ExtractionLimits`. These constraints include maximum byte counts, line limits, heredoc quantity caps, and a per-command timeout defaulting to **50 ms**.

The extraction module handles multiple embedding mechanisms:

- **Inline-script flags** via `INLINE_SCRIPT_SINGLE_QUOTE` and `INLINE_SCRIPT_DOUBLE_QUOTE` regexes
- **Windows execution wrappers** through `CMD_INLINE_SCRIPT`, `POWERSHELL_ENCODED_COMMAND`, and `IEX_INLINE_SCRIPT`
- **Heredoc variants** including `<<`, `<<-`, `<<~` via `HEREDOC_EXTRACTOR`
- **Here-strings** via `HERESTRING_SINGLE_QUOTE`, `HERESTRING_DOUBLE_QUOTE`, and `HERESTRING_UNQUOTED`

During extraction, `dcg` monitors for binary-like content, size overruns, and time-budget violations. On any anomaly, it **fails open** (allows the command) while recording a `SkipReason` for observability.

## Tier 3: AST-Based Structural Matching

After successful extraction, content enters the **AST matcher** implemented using the `ast-grep-core` crate (located in [`src/ast_matcher.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/ast_matcher.rs)). This tier parses the extracted snippet into an abstract syntax tree and applies language-specific pattern trees to recognize destructive constructs like `os.system`, `fs::remove_dir_all`, or `rm -rf`.

The [`ScriptLanguage::detect`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs#L442-L456) function in [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs) determines the appropriate parser by examining:

1. **Command-prefix** – The interpreter name from the command head (e.g., `python`, `node`, `powershell`)
2. **Shebang** – `#!/…` lines within the extracted script
3. **Content heuristics** – Regex patterns over the first 20 lines detecting imports, `require`, or `package` statements

Each detection returns a `(ScriptLanguage, DetectionConfidence)` tuple attached to the `ExtractedContent` record. The evaluator ([`src/evaluator.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/evaluator.rs)) then matches these against `SAFE_PATTERNS` and `DESTRUCTIVE_PATTERNS` tables.

## Performance Guarantees and Safety Limits

`dcg` balances thoroughness with latency through strict tiered guarantees:

| Tier | Goal | Typical Latency | Guarantees |
|------|------|----------------|------------|
| 1 – Trigger | Fast-path allow for most commands | < 10 µs (non-match), < 100 µs (match) | Zero false-negatives |
| 2 – Extraction | Safe bounded parsing | < 1 ms | Bounded memory, fail-open on error |
| 3 – AST | Accurate destructive-pattern detection | < 5 ms | Structural matching, language-aware |

## Implementation Examples

The following examples demonstrate the heredoc extraction API.

Fast-path trigger detection:

```rust
use destructive_command_guard::heredoc::{check_triggers, TriggerResult};

assert_eq!(check_triggers("git status"), TriggerResult::NoTrigger);   // allowed
assert_eq!(check_triggers("cat <<EOF"), TriggerResult::Triggered); // continue to Tier 2

```

Extracting inline Python with limits:

```rust
use destructive_command_guard::heredoc::{extract_content, ExtractionLimits};

let cmd = "python3 -c 'import os; os.system(\"rm -rf /\")'";
let res = extract_content(cmd, &ExtractionLimits::default());

match res {
    ExtractionResult::Extracted(contents) => {
        println!("found {} script(s)", contents.len());
        println!("language: {:?}", contents[0].language); // ScriptLanguage::Python
    }
    _ => println!("no extractable content"),
}

```

Language detection from heredoc body:

```rust
let heredoc_body = r#"#!/usr/bin/env python3
import subprocess
subprocess.run(["rm","-rf","/"])
"#;
let (lang, confidence) = ScriptLanguage::detect("", heredoc_body);
assert_eq!(lang, ScriptLanguage::Python);
assert_eq!(confidence, DetectionConfidence::Shebang);

```

## Key Source Files

| File | Role |
|------|------|
| [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs) | Implements the three-tier pipeline (trigger detection, extraction, language detection) |
| [`src/evaluator.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/evaluator.rs) | Applies AST-based destructive patterns to extracted scripts |
| [`src/ast_matcher.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/ast_matcher.rs) | Provides the `ast-grep` matcher for Tier 3 |
| [`src/normalize.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/normalize.rs) | Normalizes commands before heredoc detection |
| [`src/config.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/config.rs) | Holds extraction limits and enabled pattern packs |

## Summary

- **dcg** uses a three-tier pipeline to analyze heredocs and inline scripts: fast regex triggers, bounded extraction, and AST-based matching.
- **Tier 1** guarantees zero false-negatives for heredoc detection using `HEREDOC_TRIGGERS` and `check_triggers` in [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs).
- **Tier 2** extracts content with strict `ExtractionLimits` (50ms timeout, size caps) and fails open on errors to prevent blocking legitimate commands.
- **Tier 3** leverages `ast-grep-core` for structural pattern matching against `DESTRUCTIVE_PATTERNS`, with language detection via `ScriptLanguage::detect`.
- The architecture maintains sub-millisecond latency for the common case while providing deep analysis when needed.

## Frequently Asked Questions

### How does dcg prevent performance degradation when scanning every shell command?

`dcg` employs **Tier 1** trigger detection using a compiled `RegexSet` that scans commands in under 10 microseconds for non-matching inputs. Only commands containing potential heredocs or inline scripts proceed to the more expensive extraction and AST parsing stages, ensuring the common case remains imperceptible to users.

### What happens if the heredoc extraction exceeds the time or size limits?

The system **fails open**, allowing the command to execute while recording a `SkipReason` for audit purposes. The `ExtractionLimits` struct enforces defaults including a 50ms timeout and byte/line thresholds, preventing resource exhaustion attacks or denial-of-service through maliciously large heredocs.

### Which languages does the AST-based pattern matching support?

Language support is determined by `ScriptLanguage::detect` in [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs), which recognizes interpreters via command-prefix analysis (e.g., `python`, `node`, `powershell`), shebang lines (`#!/usr/bin/env python3`), and content heuristics. The `ast-grep-core` integration enables precise structural matching for any language supported by the underlying tree-sitter grammars configured in [`src/ast_matcher.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/ast_matcher.rs).

### Can the destructive pattern rules be customized or extended?

Yes, pattern configuration resides in [`src/config.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/config.rs) and the evaluator references `SAFE_PATTERNS` and `DESTRUCTIVE_PATTERNS` tables. Users can modify enabled pattern packs and extraction limits through the configuration system, though the core AST matching logic in [`src/evaluator.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/evaluator.rs) and [`src/ast_matcher.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/ast_matcher.rs) handles the structural validation.