# How dcg's Three-Tier Heredoc Scanning Architecture Works

> Discover how dcg's three-tier heredoc scanning architecture achieves sub-millisecond latency for destructive command detection through zero-allocation regex, bounded-memory parsing, and AST pattern matching.

- Repository: [Jeff Emanuel/destructive_command_guard](https://github.com/Dicklesworthstone/destructive_command_guard)
- Tags: internals
- Published: 2026-07-16

---

**The three-tier heredoc scanning architecture uses a compiled `RegexSet` for zero-allocation trigger detection, a bounded-memory parser for content extraction, and AST pattern matching for deep script analysis to identify destructive commands embedded in heredocs with sub-millisecond latency.**

The **Dicklesworthstone/destructive_command_guard** (dcg) repository protects AI agents from destructive commands that embed malicious scripts in *heredocs* or *inline* interpreter calls. To maintain extreme performance on the hot path while catching complex payloads, the detection logic in [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs) splits analysis into three successive tiers, each optimized for specific latency and accuracy requirements.

## Tier 1: Trigger Detection (Zero-Allocation Fast Path)

The first tier acts as a high-speed filter that decides whether a command *might* contain a heredoc or inline script without performing any heap allocations.

In [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs), the `HEREDOC_TRIGGERS` constant holds a compiled **`RegexSet`** containing 17 broad patterns that match indicators like `<<<`, interpreter names (`python`, `perl`, `powershell`), and execution flags (`-c`, `-e`). The set evaluates the input string in a single pass, using **`memchr`** to short-circuit immediately when the `<` character is absent.

This tier achieves **< 100 µs** latency when a pattern matches and **< 10 µs** when no match occurs, covering > 99% of everyday shell invocations with virtually no overhead.

## Tier 2: Content Extraction (Bounded-Memory Parsing)

When Tier 1 signals a potential match, the second tier extracts the exact script text for deeper inspection. The function **`contains_active_heredoc_operator`** recursively scans the command to identify active heredoc operators (`<<`) and inline-script arguments (e.g., `python -c '...'`).

A **bounded-memory parser** walks the command structure, correctly handling:
- Quoted strings (single and double)
- Command substitutions (`$()`)
- Backtick substitutions
- Escaped newlines

If extraction fails due to malformed input or timeouts, the architecture **gracefully degrades**—allowing the command but emitting a warning to ensure the hot path never crashes. This stage completes in **< 1 ms**.

## Tier 3: AST Pattern Matching (Deep Language Analysis)

The final tier performs deep, language-aware analysis of the extracted script to determine if it matches destructive patterns. While currently future-ready in the codebase, this stage plans to use **ast-grep-core** to parse the script into an abstract syntax tree.

The AST matcher runs language-specific patterns to detect dangerous sequences like `rm -rf /` or `git reset --hard`. A match produces a **BLOCK** decision with structured denial JSON; otherwise, the command is allowed. This deep inspection targets **< 5 ms** latency.

## Integration Flow: How the Tiers Interact

The three tiers operate sequentially, with each stage gated by the previous:

1. **Fast-path (Tier 1).** The command streams through `HEREDOC_TRIGGERS`. If no patterns match, dcg instantly **allows** the command.
2. **Potential-match (Tier 1 ⇒ Tier 2).** When a trigger fires, `contains_active_heredoc_operator` extracts the script body. If extraction succeeds, the pipeline moves to Tier 3; otherwise, it **allows** with a warning.
3. **Deep inspection (Tier 2 ⇒ Tier 3).** The extracted snippet passes to the AST matcher. Destructive matches trigger denial; non-matches result in **allow**.

```rust
// Tier 1: quick check (used internally by dcg)
if HEREDOC_TRIGGERS.is_match(command) {
    // Tier 2 extraction
    if let Some(script) = extract_script(command) {
        // Tier 3 AST matching (future-ready)
        if ast_matcher::is_destructive(&script) {
            deny_with_json(...);
        } else {
            allow();
        }
    } else {
        // extraction failed → allow with a warning
        warn!("heredoc extraction error, allowing command");
        allow();
    }
} else {
    // No trigger → fast-path allow
    allow();
}

```

## Performance Guarantees and Safety Properties

The three-tier design provides specific guarantees enforced by the implementation in [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs) and validated by [`benches/heredoc_perf.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/benches/heredoc_perf.rs):

- **Zero allocations on the non-match path.** The entire pipeline stays allocation-free for the overwhelming majority of commands that pass Tier 1.
- **Zero false-negatives for Tier 1.** Every possible heredoc or inline script triggers Tier 2, ensuring no destructive script slips through because the initial regex missed it.
- **Graceful fallback.** Malformed input or extraction timeouts never cause hard failures; the system allows the command but logs the event for visibility, as documented in [`docs/heredoc-error-messages.md`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/docs/heredoc-error-messages.md).

## Summary

- **Tier 1** uses a compiled `RegexSet` with 17 patterns and `memchr` optimization to filter commands in < 10 µs without allocations.
- **Tier 2** employs `contains_active_heredoc_operator` and a bounded-memory parser to extract script content while handling complex shell syntax in < 1 ms.
- **Tier 3** (future-ready) will use `ast-grep-core` for AST-based pattern matching to detect specific destructive commands in < 5 ms.
- The architecture guarantees zero allocations on the fast path and zero false negatives for heredoc detection, with graceful degradation for malformed input.

## Frequently Asked Questions

### Why does dcg use a three-tier architecture instead of a single regex?

A single regex would require either excessive memory allocations for complex parsing or miss sophisticated evasion techniques. The three-tier heredoc scanning architecture separates concerns: Tier 1 provides speed, Tier 2 handles structural extraction, and Tier 3 enables deep semantic analysis. According to the benchmarks in [`benches/heredoc_perf.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/benches/heredoc_perf.rs), this design maintains < 100 µs latency for 99% of commands while allowing complex analysis when needed.

### How does dcg handle malformed heredocs or extraction timeouts?

The system implements graceful degradation in [`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs). If the bounded-memory parser encounters unbalanced quotes, excessive recursion, or timeouts during extraction, it does not block the command. Instead, it emits a warning via the logging system (documented in [`docs/heredoc-error-messages.md`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/docs/heredoc-error-messages.md)) and allows the command to proceed, preventing denial-of-service via crafted input.

### What is the latency impact of the heredoc scanner on normal commands?

For commands that do not contain heredocs or inline scripts, the latency impact is **< 10 µs** due to the `memchr`-optimized `RegexSet` in Tier 1. This zero-allocation fast path ensures that standard shell invocations experience negligible overhead. The heavier Tier 2 and Tier 3 analysis only executes when the initial trigger patterns match, as verified by [`tests/heredoc_pack_gap.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/tests/heredoc_pack_gap.rs).

### Which source files implement the three-tier heredoc scanning architecture?

The core logic resides in **[`src/heredoc.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/heredoc.rs)**, which defines `HEREDOC_TRIGGERS` and `contains_active_heredoc_operator`. The **[`src/evaluator.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/src/evaluator.rs)** file integrates these checks into the overall pattern-matching pipeline. Performance benchmarks are located in **[`benches/heredoc_perf.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/benches/heredoc_perf.rs)**, unit tests for edge cases are in **[`tests/heredoc_pack_gap.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/tests/heredoc_pack_gap.rs)**, and the fuzzing harness for robustness testing is in **[`fuzz/fuzz_targets/heredoc_fuzz.rs`](https://github.com/Dicklesworthstone/destructive_command_guard/blob/main/fuzz/fuzz_targets/heredoc_fuzz.rs)**.