How Heredoc and Inline Script Extraction with AST Matching Works in dcg
Destructive Command Guard (dcg) uses a three-tier pipeline—trigger detection, bounded content extraction, and AST pattern matching—to analyze heredocs and inline scripts for dangerous commands while maintaining sub-millisecond latency for safe commands.
Destructive Command Guard (dcg) is a Rust-based security hook that prevents agents from executing dangerous shell commands by analyzing embedded scripts. Understanding how heredoc and inline script extraction with AST matching works reveals why the tool can process thousands of benign commands in microseconds while still catching sophisticated destructive payloads hidden behind variable aliases and indirect calls. The implementation splits analysis into distinct tiers so that cheap, allocation-free checks run first, while heavyweight structural analysis only executes when a potential risk is detected.
The Three-Tier Architecture
The pipeline is deliberately separated into three tightly-coupled stages defined in src/heredoc.rs and src/ast_matcher.rs. This design ensures that the fast path rejects safe commands quickly, while deeper inspection occurs only when triggers are present.
Tier 1: Trigger Detection
The first tier acts as a high-speed filter. In src/heredoc.rs, the check_triggers function uses a pre-compiled RegexSet containing 17 patterns to spot heredoc operators (<<, <<<, <<-, <<~) and interpreter flags like python -c, bash -c, PowerShell -Command, and cmd /c. A hand-written zero-allocation parser, contains_active_heredoc_operator, searches for the << operator without allocating memory. This stage typically completes in < 10 µs for non-matching commands and < 100 µs when a trigger is found.
Tier 2: Content Extraction
When triggers are detected, the extract_content function parses the actual script text while guarding against resource exhaustion. The extractor handles here-strings, inline-script flags, and classic heredocs using compiled regexes like HERESTRING_SINGLE_QUOTE, INLINE_SCRIPT_DOUBLE_QUOTE, and CMD_INLINE_SCRIPT. It respects ExtractionLimits (max 1 MiB, 10,000 lines, 10 heredocs, 50 ms timeout). If content is binary, oversized, or times out, the system fail-opens (allows the command) but logs the event via SkipReason. Typical latency remains ≤ 1 ms for normal bodies.
Tier 3: AST Pattern Matching
The final tier performs language-aware structural analysis using AstMatcher in src/ast_matcher.rs. It pre-compiles destructive patterns for supported languages—Python, JavaScript, TypeScript, Ruby, Bash, Go, and PHP—using ast-grep-core (AstGrep). The matcher walks the parsed AST (root.find_all(&compiled.pattern)) to identify destructive API calls that regex would miss, such as indirect executions through variable aliases. A hard timeout of 20 ms (5 s in tests) guarantees the hook never exceeds its deadline. For unsupported languages like Perl, a regex fallback (scan_executing_sink_fallback) checks for known destructive function calls.
Language Detection and Safety Mechanisms
Before AST analysis runs, ScriptLanguage::detect infers the script language from three signals in descending priority: command-prefix (e.g., python3 -c), shebang (e.g., #!/usr/bin/env ruby), and content heuristics scanning the first few lines. The selected language drives both extraction delimiter selection and the AST matcher configuration.
Resource Limits and Fallbacks
The system implements multiple hardening measures. check_binary_content aborts extraction if null bytes or high non-printable ratios are detected. Both Tier 2 (record_timeout_if_needed) and Tier 3 (run_ast_match_with_timeout) monitor a shared deadline; on expiry, the command is allowed but emits a warning. Even when AST parsing cannot run, the exec-sink fallback catches dangerous patterns like execSync, os.system, or Ruby system when they contain destructive payloads.
Pipeline Implementation
The src/evaluator.rs file orchestrates these tiers into a cohesive flow:
use destructive_command_guard::heredoc::{check_triggers, extract_content, ExtractionLimits};
use destructive_command_guard::ast_matcher::AstMatcher;
/// High-level entry point used by the hook.
pub fn analyse_command(command: &str) {
// ----- Tier 1 ---------------------------------------------------------
if check_triggers(command) == TriggerResult::NoTrigger {
// Fast-path: nothing to inspect → command allowed.
return;
}
// ----- Tier 2 ---------------------------------------------------------
let limits = ExtractionLimits::default(); // 1 MiB, 10 heredocs, 50 ms
let extraction = extract_content(command, &limits);
let scripts = match extraction {
ExtractionResult::Extracted(vec) => vec,
ExtractionResult::Skipped(_)
| ExtractionResult::NoContent
| ExtractionResult::Failed(_) => return, // fail-open
_ => return,
};
// ----- Tier 3 ---------------------------------------------------------
let matcher = AstMatcher::default(); // loads static destructive patterns
for extracted in scripts {
if let Some(match_) = matcher.has_blocking_match(
&extracted.content,
extracted.language,
) {
// A destructive AST pattern matched → deny the command.
// The matcher returns a `PatternMatch` that contains
// rule_id, severity, line number, and a short preview.
deny(match_);
return;
}
}
// No AST matches → safe to allow.
}
Why AST Matching Matters
Pure regexes cannot reliably detect indirect executions such as:
const cp = require('child_process');
cp.execSync('rm -rf /'); // alias → hard to spot with a naïve pattern
AST parsing recognizes the callee (execSync) regardless of the variable name and examines the actual argument value, including list arguments like subprocess.run(["sh","-c","rm -rf /"]). This eliminates false negatives while keeping false positives low because patterns are limited to a curated destructive set.
Practical Examples
Trigger detection identifies commands requiring deeper inspection:
use destructive_command_guard::heredoc::{check_triggers, TriggerResult};
assert_eq!(check_triggers("git status"), TriggerResult::NoTrigger);
assert_eq!(check_triggers("cat <<EOF"), TriggerResult::Triggered);
assert_eq!(check_triggers("python -c 'import os'"), TriggerResult::Triggered);
Extracting heredoc content with safety limits:
use destructive_command_guard::heredoc::{extract_content, ExtractionLimits};
let cmd = r#"cat <<'EOS'
#!/usr/bin/env python3
import os; os.system("rm -rf /")
EOS
"#;
let result = extract_content(cmd, &ExtractionLimits::default());
match result {
ExtractionResult::Extracted(contents) => {
for e in contents {
println!("Language: {:?}, Body length: {}", e.language, e.content.len());
}
}
_ => println!("No extractable content"),
}
Running structural analysis against extracted scripts:
use destructive_command_guard::ast_matcher::AstMatcher;
use destructive_command_guard::heredoc::ScriptLanguage;
let matcher = AstMatcher::default();
let script = r#"import os; os.system("rm -rf /")"#;
if let Some(m) = matcher.has_blocking_match(script, ScriptLanguage::Python) {
println!("Blocked by rule {}: {}", m.rule_id, m.reason);
}
Summary
- Three-tier design separates fast trigger detection (
src/heredoc.rslines 46-62) from expensive AST analysis to maintain < 10 µs latency for safe commands. - Bounded extraction enforces hard limits (1 MiB, 10k lines, 50 ms) and fail-open behavior to prevent denial-of-service via malicious heredocs.
- AST matching via
AstMatcherand ast-grep-core catches indirect destructive calls that regex misses, with a 20 ms hard timeout ensuring predictable performance. - Language detection uses command-prefix, shebang, and content heuristics, falling back to regex scans for unsupported grammars like Perl.
Frequently Asked Questions
What is the latency impact of dcg's heredoc analysis?
According to the dcg source code, benign commands that trigger no heredoc or inline script patterns complete in < 10 µs. Commands requiring extraction typically finish in ≤ 1 ms, while AST matching adds approximately 5 ms with a hard cap at 20 ms. This ensures the hook adds negligible overhead to interactive shells.
How does dcg handle binary content in heredocs?
The check_binary_content function in src/heredoc.rs scans for null bytes or high ratios of non-printable characters. If binary content is detected, extraction aborts immediately and the system logs a SkipReason, failing open to allow the command while preventing resource exhaustion from binary parsing.
What happens when a language isn't supported by the AST matcher?
For languages lacking an ast-grep grammar (such as Perl), dcg invokes scan_executing_sink_fallback in src/ast_matcher.rs. This lightweight regex scans for known destructive function signatures like execSync, os.system, or Ruby system and blocks commands containing these patterns when they reference dangerous payloads.
Why does dcg use a fail-open strategy for timeouts?
Both extraction and AST matching enforce strict deadlines (50 ms and 20 ms respectively). When these timeout, dcg allows the command to proceed but logs the event via SkipReason. This prevents the security hook from blocking legitimate operations due to edge-case complexity or resource contention, prioritizing availability while maintaining audit trails.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →