How the Three-Tier Heredoc Scanning Architecture Works in dcg

The three-tier heredoc scanning architecture in dcg uses a fast regex trigger, followed by content extraction, and finally AST pattern matching to detect destructive scripts embedded in heredocs or inline interpreter calls while maintaining sub-millisecond latency for the hot path.

The destructive_command_guard (dcg) repository protects AI agents from shell commands that conceal destructive scripts within heredocs and inline interpreter invocations. To balance security with performance, the detection logic in src/heredoc.rs implements a three-tier heredoc scanning architecture that escalates analysis depth only when necessary. This design ensures that over 99% of benign commands bypass expensive processing through zero-allocation fast paths.

Overview of the Tiered Detection Strategy

The architecture divides detection into three successive stages, each optimized for specific latency budgets and analysis depth. According to the implementation in src/heredoc.rs, the system processes commands through Tier 1 (Trigger Detection), Tier 2 (Content Extraction), and Tier 3 (AST Pattern Matching) only when the previous tier signals a potential threat.

This progressive approach guarantees that the hot path remains allocation-free for typical commands while ensuring zero false negatives for heredoc-based attacks. The benchmarks in benches/heredoc_perf.rs verify that each tier meets its strict latency constraints.

Tier 1: Trigger Detection (The Fast Path)

The first tier acts as a high-speed filter using a compiled RegexSet containing 17 broad patterns. These patterns match heredoc operators (<<<), interpreter names (such as python, perl, and powershell), and execution flags (-c, -e, etc.).

To maximize speed, the implementation leverages memchr to short-circuit evaluation when the input lacks the < character, avoiding regex overhead entirely for most commands. This tier operates in < 100 µs for potential matches and < 10 µs for non-matches, performing zero allocations on the rejection path.

When HEREDOC_TRIGGERS.is_match(command) returns false, dcg immediately allows the command without further processing.

Tier 2: Content Extraction

When Tier 1 detects a potential script, the architecture invokes the contains_active_heredoc_operator function (defined in src/heredoc.rs and called from src/evaluator.rs). This tier employs a bounded-memory parser that traverses the command string while handling quoted strings, $() substitutions, backticks, and escaped newlines.

The extraction logic targets:

  • Heredoc bodies (content between << delimiters)
  • Inline script arguments (e.g., python -c '...')
  • Base64-encoded payloads

If extraction succeeds within the < 1 ms latency budget, the cleaned script text passes to Tier 3. If parsing fails or times out, the system gracefully degrades by allowing the command with a warning, ensuring the hot path never crashes. The specific warning messages are documented in docs/heredoc-error-messages.md.

Tier 3: AST Pattern Matching

The final tier performs deep, language-aware analysis using the ast-grep-core library (planned implementation). This stage parses the extracted script into an abstract syntax tree to identify destructive patterns such as rm -rf / or git reset --hard.

Operating within a < 5 ms budget, this tier eliminates false positives from earlier stages by analyzing concrete code structure rather than surface syntax. A match produces a BLOCK decision with structured JSON denial, while non-matches result in immediate allow.

How the Tiers Interact

The interaction between tiers follows a strict escalation protocol:

  1. Fast-path allow: If Tier 1 finds no triggers, dcg allows the command instantly.
  2. Extraction attempt: When Tier 1 fires, the system recursively scans for active heredoc operators and inline scripts.
  3. Deep inspection: Successfully extracted snippets undergo AST analysis; failures allow with warnings.

This flow ensures that malformed input or resource exhaustion never triggers hard failures, maintaining availability while logging suspicious attempts. The fuzz harness in fuzz/fuzz_targets/heredoc_fuzz.rs continuously validates this robustness against edge-case inputs.

Performance Guarantees and Safety Properties

The three-tier heredoc scanning architecture provides specific safety and performance guarantees:

  • Zero allocations on the non-match path, keeping the hot path memory-efficient
  • Zero false negatives at Tier 1—every possible heredoc or inline script reaches Tier 2
  • Graceful fallback through warning-based allowing when Tier 2 extraction fails
  • Bounded latency with strict budgets: <100 µs (Tier 1), <1 ms (Tier 2), <5 ms (Tier 3)

The unit tests in tests/heredoc_pack_gap.rs verify that the tiered flow correctly handles complex command structures without breaking the latency contracts.

Implementation Example

The following Rust logic (adapted from src/heredoc.rs) illustrates the tiered decision flow:

// Tier 1: quick check (used internally by dcg)
if HEREDOC_TRIGGERS.is_match(command) {
    // Tier 2 extraction
    if let Some(script) = extract_script(command) {
        // Tier 3 AST matching (future-ready)
        if ast_matcher::is_destructive(&script) {
            deny_with_json(...);
        } else {
            allow();
        }
    } else {
        // extraction failed → allow with a warning
        warn!("heredoc extraction error, allowing command");
        allow();
    }
} else {
    // No trigger → fast-path allow
    allow();
}

Summary

  • The three-tier heredoc scanning architecture in dcg escalates analysis depth only when triggers indicate potential threats.
  • Tier 1 uses a RegexSet with 17 patterns and memchr optimization to filter commands in under 100 µs without allocations.
  • Tier 2 extracts script content via contains_active_heredoc_operator using bounded-memory parsing, handling heredocs, inline flags, and encodings.
  • Tier 3 (planned) will leverage AST pattern matching with ast-grep-core to eliminate false positives through structural analysis.
  • The system guarantees zero allocation on rejection paths, zero false negatives for trigger detection, and graceful degradation for malformed input.

Frequently Asked Questions

What makes the three-tier heredoc scanning architecture faster than single-pass detection?

The architecture separates concerns by latency budget: Tier 1 rejects benign commands in under 10 microseconds using simple regex patterns without allocations, while expensive parsing and AST analysis are reserved for the <1% of commands that actually contain heredocs or inline scripts. This prevents the hot path from paying the cost of complex parsing for everyday shell invocations.

How does dcg handle malformed heredocs or extraction timeouts?

The system implements graceful degradation in Tier 2. If the bounded-memory parser encounters errors or exceeds time limits, src/heredoc.rs allows the command but emits a warning. This safety mechanism ensures that resource exhaustion or syntax edge cases never cause denial-of-service for legitimate commands.

Which source files implement the three-tier heredoc scanning logic?

The core implementation resides in src/heredoc.rs, which defines the HEREDOC_TRIGGERS regex set, contains_active_heredoc_operator function, and extraction logic. The src/evaluator.rs file integrates these checks into the broader command evaluation pipeline, while benchmarks in benches/heredoc_perf.rs and tests in tests/heredoc_pack_gap.rs verify tier behavior.

Why does Tier 3 use AST pattern matching instead of additional regex checks?

AST pattern matching with ast-grep-core analyzes the actual structure of extracted code rather than surface text, eliminating false positives that regex alone cannot distinguish. This structural analysis can reliably identify destructive operations like rm -rf / while ignoring similar strings inside comments or strings, providing the precision needed for security-critical blocking decisions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →