How dcg's Three-Tier Heredoc Scanning Architecture Detects Destructive Commands

dcg employs a three-tier heredoc scanning architecture that combines zero-allocation regex triggers, bounded-memory content extraction, and planned AST-based analysis to identify destructive scripts embedded in shell heredocs while maintaining sub-millisecond latency for the vast majority of commands.

The destructive_command_guard (dcg) repository protects AI agents from executing dangerous commands hidden within heredocs or inline interpreter invocations. To balance detection accuracy with performance, the system implements a cascading three-tier scanning architecture in src/heredoc.rs that avoids memory allocation on the hot path while ensuring complex payloads never slip through undetected.

The Three-Tier Detection Pipeline

The architecture splits detection into three successive stages, each triggered only if the previous tier signals a potential threat. This design ensures that >99% of everyday shell invocations receive a verdict in under 100 microseconds without allocating memory.

Tier 1 – Trigger Detection with Zero Allocation

The first tier acts as a high-speed filter using a compiled RegexSet containing 17 broad patterns that match heredoc operators (<<<) or interpreter names (python, perl, powershell) combined with script execution flags (-c, -e).

The implementation leverages memchr to short-circuit evaluation when the < character is absent from the command string. If HEREDOC_TRIGGERS.is_match(command) returns false, dcg instantly allows the command. This path executes in <10 µs for non-matches and <100 µs when a match occurs, with zero heap allocations.

Tier 2 – Content Extraction and Script Isolation

When Tier 1 detects a potential threat, the system invokes contains_active_heredoc_operator to perform a bounded-memory parse of the command. This recursive scanner handles quoted strings, $() substitutions, backticks, and escaped newlines to extract the exact script text—whether from a heredoc body, a -c argument, or a base64-encoded payload.

If extraction succeeds, the isolated script content moves to Tier 3. If the parser encounters malformed syntax or timeouts, dcg gracefully degrades by allowing the command while emitting a warning, ensuring the hot path never crashes. This tier typically completes in <1 ms.

Tier 3 – AST Pattern Matching for Deep Inspection

The final tier performs language-aware analysis using ast-grep-core (planned implementation) to parse the extracted script into an abstract syntax tree. This allows dcg to run language-specific patterns—such as detecting rm -rf / or git reset --hard—with semantic understanding that regex alone cannot provide.

A destructive match produces a structured BLOCK decision, while non-matches result in allow. This deep inspection targets a latency budget of <5 ms, acceptable for the rare commands that reach this stage.

How the Tiers Interact in Practice

The interaction between tiers follows a strict cascading flow implemented in src/heredoc.rs and called from src/evaluator.rs:

// Tier 1: quick check (used internally by dcg)
if HEREDOC_TRIGGERS.is_match(command) {
    // Tier 2 extraction
    if let Some(script) = extract_script(command) {
        // Tier 3 AST matching (future-ready)
        if ast_matcher::is_destructive(&script) {
            deny_with_json(...);
        } else {
            allow();
        }
    } else {
        // extraction failed → allow with a warning
        warn!("heredoc extraction error, allowing command");
        allow();
    }
} else {
    // No trigger → fast-path allow
    allow();
}

Fast-path optimization: When Tier 1 finds no triggers, the command bypasses all heavier processing, preserving the <10 µs latency for benign inputs.

Fail-safe design: Tier 1 guarantees zero false negatives—every possible heredoc or inline script triggers the extraction phase, ensuring no destructive payload bypasses the initial filter due to pattern limitations.

Performance Guarantees and Safety Properties

The three-tier architecture provides three critical guarantees that distinguish dcg from simpler regex-based approaches:

  • Zero allocations on the non-match path: The entire pipeline remains allocation-free for the overwhelming majority of commands, preventing memory pressure during high-throughput scenarios.
  • Zero false negatives for Tier 1: The broad 17-pattern RegexSet ensures every heredoc or inline script invocation reaches Tier 2 for content extraction.
  • Graceful fallback: Malformed input or extraction timeouts never cause hard failures; the system allows the command but logs a warning for visibility.

Key Implementation Files

The heredoc scanning system spans several critical files in the Dicklesworthstone/destructive_command_guard repository:

Summary

  • dcg's three-tier heredoc scanning architecture uses a cascading filter approach to minimize latency for safe commands while ensuring deep inspection of suspicious inputs.
  • Tier 1 employs a RegexSet with memchr optimization to provide <10 µs zero-allocation screening for >99% of commands.
  • Tier 2 implements a bounded-memory parser that extracts script content from heredocs and inline interpreters, with graceful degradation on parse failures.
  • Tier 3 (planned) will use AST-based pattern matching via ast-grep-core to detect semantically destructive operations impossible to catch with regex alone.
  • The system guarantees zero false negatives at the trigger level and never crashes on malformed input, allowing commands with warnings when extraction fails.

Frequently Asked Questions

What makes dcg's heredoc detection faster than single-pass regex scanning?

Single-pass regex scanning requires evaluating complex patterns against every input, often involving backtracking and allocations. dcg's three-tier heredoc scanning architecture uses a fast RegexSet with 17 simple patterns in Tier 1 that short-circuit via memchr, eliminating 99% of inputs in <10 µs without allocation. Only potential threats trigger the heavier Tier 2 and Tier 3 processing, maintaining average latency orders of magnitude lower than deep inspection of every command.

How does dcg handle malformed heredoc syntax or extraction timeouts?

The system implements graceful degradation in Tier 2. If the bounded-memory parser encounters unclosed quotes, invalid escape sequences, or exceeds time limits during extraction, it does not block the command or panic. Instead, dcg allows the command to proceed while emitting a structured warning log, ensuring that edge-case syntax never causes availability issues for AI agents.

Why does Tier 1 use RegexSet instead of a single complex regex?

A single complex regex combining 17 different patterns for heredocs, python, perl, powershell, and various flags would suffer from catastrophic backtracking and poor cache locality. The RegexSet implementation in src/heredoc.rs evaluates all patterns simultaneously in a single pass using finite automata, providing O(n) linear scanning with memchr short-circuiting when the < character is absent, achieving the <10 µs latency target impossible with monolithic regex patterns.

What types of destructive patterns will Tier 3 detect that Tier 1 cannot?

While Tier 1 detects the presence of script content, Tier 3 uses ast-grep-core to detect semantically destructive operations like rm -rf / disguised by variable interpolation, git reset --hard hidden in shell functions, or os.system() calls in Python heredocs. AST matching understands language syntax trees, eliminating false positives from Tier 1 (like benign python -c "print('hello')") while catching obfuscated commands that would bypass literal string matching.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →