How dcg Handles Heredoc Extraction with AST-Based Pattern Matching: A Three-Tier Pipeline

dcg (Destructive Command Guard) employs a three-tier detection pipeline—fast regex triggers, bounded content extraction, and AST-based structural analysis—to safely evaluate heredocs and inline scripts with sub-millisecond latency.

The destructive_command_guard repository by Dicklesworthstone implements a high-performance pre-hook for shell commands that must analyze complex heredoc and inline-script payloads without impacting interactive latency. This article examines how the Rust-based dcg tool combines regex-based triage with ast-grep-core powered pattern matching to extract and analyze embedded code.

Tier 1: Ultra-Fast Trigger Detection

Implemented in src/heredoc.rs, the check_triggers function serves as the first line of defense. It utilizes a compiled RegexSet named HEREDOC_TRIGGERS to scan the entire command in a single pass against 17 distinct patterns.

These patterns detect heredoc operators (<<, <<<), inline script flags (python -c, node -e), Windows wrappers (cmd /c, powershell -Command), and pipe-to-interpreter sequences. The function also invokes contains_active_heredoc_operator, a hand-written scanner that specifically checks for the << operator to guarantee zero false-negatives while permitting false positives—any trigger merely forces progression to Tier 2.

Tier 2: Bounded Content Extraction

When Tier 1 triggers fire, the extract_content function handles safe parsing with strict resource limits defined by ExtractionLimits. These constraints include maximum byte counts, line limits, heredoc quantity caps, and a per-command timeout defaulting to 50 ms.

The extraction module handles multiple embedding mechanisms:

  • Inline-script flags via INLINE_SCRIPT_SINGLE_QUOTE and INLINE_SCRIPT_DOUBLE_QUOTE regexes
  • Windows execution wrappers through CMD_INLINE_SCRIPT, POWERSHELL_ENCODED_COMMAND, and IEX_INLINE_SCRIPT
  • Heredoc variants including <<, <<-, <<~ via HEREDOC_EXTRACTOR
  • Here-strings via HERESTRING_SINGLE_QUOTE, HERESTRING_DOUBLE_QUOTE, and HERESTRING_UNQUOTED

During extraction, dcg monitors for binary-like content, size overruns, and time-budget violations. On any anomaly, it fails open (allows the command) while recording a SkipReason for observability.

Tier 3: AST-Based Structural Matching

After successful extraction, content enters the AST matcher implemented using the ast-grep-core crate (located in src/ast_matcher.rs). This tier parses the extracted snippet into an abstract syntax tree and applies language-specific pattern trees to recognize destructive constructs like os.system, fs::remove_dir_all, or rm -rf.

The ScriptLanguage::detect function in src/heredoc.rs determines the appropriate parser by examining:

  1. Command-prefix – The interpreter name from the command head (e.g., python, node, powershell)
  2. Shebang – #!/… lines within the extracted script
  3. Content heuristics – Regex patterns over the first 20 lines detecting imports, require, or package statements

Each detection returns a (ScriptLanguage, DetectionConfidence) tuple attached to the ExtractedContent record. The evaluator (src/evaluator.rs) then matches these against SAFE_PATTERNS and DESTRUCTIVE_PATTERNS tables.

Performance Guarantees and Safety Limits

dcg balances thoroughness with latency through strict tiered guarantees:

Tier Goal Typical Latency Guarantees
1 – Trigger Fast-path allow for most commands < 10 µs (non-match), < 100 µs (match) Zero false-negatives
2 – Extraction Safe bounded parsing < 1 ms Bounded memory, fail-open on error
3 – AST Accurate destructive-pattern detection < 5 ms Structural matching, language-aware

Implementation Examples

The following examples demonstrate the heredoc extraction API.

Fast-path trigger detection:

use destructive_command_guard::heredoc::{check_triggers, TriggerResult};

assert_eq!(check_triggers("git status"), TriggerResult::NoTrigger);   // allowed
assert_eq!(check_triggers("cat <<EOF"), TriggerResult::Triggered); // continue to Tier 2

Extracting inline Python with limits:

use destructive_command_guard::heredoc::{extract_content, ExtractionLimits};

let cmd = "python3 -c 'import os; os.system(\"rm -rf /\")'";
let res = extract_content(cmd, &ExtractionLimits::default());

match res {
    ExtractionResult::Extracted(contents) => {
        println!("found {} script(s)", contents.len());
        println!("language: {:?}", contents[0].language); // ScriptLanguage::Python
    }
    _ => println!("no extractable content"),
}

Language detection from heredoc body:

let heredoc_body = r#"#!/usr/bin/env python3
import subprocess
subprocess.run(["rm","-rf","/"])
"#;
let (lang, confidence) = ScriptLanguage::detect("", heredoc_body);
assert_eq!(lang, ScriptLanguage::Python);
assert_eq!(confidence, DetectionConfidence::Shebang);

Key Source Files

File Role
src/heredoc.rs Implements the three-tier pipeline (trigger detection, extraction, language detection)
src/evaluator.rs Applies AST-based destructive patterns to extracted scripts
src/ast_matcher.rs Provides the ast-grep matcher for Tier 3
src/normalize.rs Normalizes commands before heredoc detection
src/config.rs Holds extraction limits and enabled pattern packs

Summary

  • dcg uses a three-tier pipeline to analyze heredocs and inline scripts: fast regex triggers, bounded extraction, and AST-based matching.
  • Tier 1 guarantees zero false-negatives for heredoc detection using HEREDOC_TRIGGERS and check_triggers in src/heredoc.rs.
  • Tier 2 extracts content with strict ExtractionLimits (50ms timeout, size caps) and fails open on errors to prevent blocking legitimate commands.
  • Tier 3 leverages ast-grep-core for structural pattern matching against DESTRUCTIVE_PATTERNS, with language detection via ScriptLanguage::detect.
  • The architecture maintains sub-millisecond latency for the common case while providing deep analysis when needed.

Frequently Asked Questions

How does dcg prevent performance degradation when scanning every shell command?

dcg employs Tier 1 trigger detection using a compiled RegexSet that scans commands in under 10 microseconds for non-matching inputs. Only commands containing potential heredocs or inline scripts proceed to the more expensive extraction and AST parsing stages, ensuring the common case remains imperceptible to users.

What happens if the heredoc extraction exceeds the time or size limits?

The system fails open, allowing the command to execute while recording a SkipReason for audit purposes. The ExtractionLimits struct enforces defaults including a 50ms timeout and byte/line thresholds, preventing resource exhaustion attacks or denial-of-service through maliciously large heredocs.

Which languages does the AST-based pattern matching support?

Language support is determined by ScriptLanguage::detect in src/heredoc.rs, which recognizes interpreters via command-prefix analysis (e.g., python, node, powershell), shebang lines (#!/usr/bin/env python3), and content heuristics. The ast-grep-core integration enables precise structural matching for any language supported by the underlying tree-sitter grammars configured in src/ast_matcher.rs.

Can the destructive pattern rules be customized or extended?

Yes, pattern configuration resides in src/config.rs and the evaluator references SAFE_PATTERNS and DESTRUCTIVE_PATTERNS tables. Users can modify enabled pattern packs and extraction limits through the configuration system, though the core AST matching logic in src/evaluator.rs and src/ast_matcher.rs handles the structural validation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →