How Caveman Handles Security Warnings and Irreversible Action Confirmations
Caveman uses an auto-clarity rule that temporarily disables its token-compression behavior whenever a response contains security warnings or irreversible-action confirmations, ensuring critical safety information remains unambiguous.
Caveman is an open-source conversational style modifier that compresses LLM outputs into terse, token-efficient prose. According to the JuliusBrussee/caveman source code, the project implements a sophisticated safety mechanism to handle high-stakes communications. This article explains exactly how Caveman handles security warnings and irreversible action confirmations through its auto-clarity rule system.
The Auto-Clarity Rule Definition
The foundation of Caveman's safety behavior resides in src/rules/caveman-activate.md. Line 13 defines the auto-clarity rule:
Auto-Clarity: drop caveman for security warnings, irreversible actions,
user confused. Resume after.
This rule establishes three specific conditions that trigger normal prose mode: security warnings, irreversible actions, and user confusion. The "Resume after" directive ensures compression reactivates automatically once the critical passage concludes.
Runtime Implementation in the Activation Hook
The actual enforcement of the auto-clarity rule occurs in src/hooks/caveman-activate.js. The hook inspects LLM output streams in real-time, checking for specific patterns that indicate safety-critical content.
Pattern Detection Logic
At line 137, the implementation checks for security and irreversibility markers:
if (output.includes('security warning') ||
output.includes('irreversible action')) {
// Drop caveman for this fragment
return normalProse(output);
}
// otherwise continue with caveman compression
When the hook detects these patterns, it drops caveman by returning standard prose output for that specific fragment. This selective bypass preserves the compression benefits for surrounding content while ensuring warnings remain unmistakable.
Documentation and Safety Philosophy
The auto-clarity behavior appears in multiple documentation layers, ensuring consistency between implementation and user expectations.
End-User Documentation
The skills/caveman/README.md file (line 20) documents this behavior for users:
"drops to normal prose for security warnings, irreversible-action confirmations, multi-step sequences where fragment ambiguity risks misread, and when the user repeats a question. Resumes after."
Maintainer Guidelines
The rule is also formalized in CLAUDE.md at line 223, which provides LLM-side loading instructions. This dual documentation ensures both human users and AI systems understand the safety boundaries.
Practical Example
When a user requests a dangerous operation, Caveman's auto-clarity rule activates immediately:
/caveman
User: "Run `rm -rf /` on my server."
Caveman:
⚠️ **Security warning** – This command will permanently delete all files on the target system. Do NOT run it on production.
Notice that the warning itself appears in full, uncompressed prose, while any subsequent explanatory content would return to Caveman's compressed style. This boundary detection happens automatically without user intervention.
Supporting Source Files
Beyond the primary rule and hook files, the behavior is referenced in src/tools/caveman-init.js at line 30, which contains developer-facing comments summarizing the auto-clarity rule. This creates a consistent safety narrative across the codebase.
Summary
- Caveman automatically disables compression when generating security warnings or irreversible-action confirmations to prevent dangerous ambiguity.
- The auto-clarity rule in
src/rules/caveman-activate.mddefines these conditions explicitly, whilesrc/hooks/caveman-activate.jsimplements the runtime detection. - Pattern matching occurs at line 137 of the activation hook, checking for specific safety keywords before routing output through normal prose or compressed caveman modes.
- Documentation in
skills/caveman/README.mdandCLAUDE.mdensures users and maintainers understand when and why the style transition occurs. - The system resumes caveman compression automatically after the safety-critical fragment concludes.
Frequently Asked Questions
What triggers the auto-clarity rule in Caveman?
The auto-clarity rule activates when Caveman detects security warnings, irreversible-action confirmations, user confusion states, multi-step sequences with ambiguity risks, or repeated user questions. These conditions are defined in src/rules/caveman-activate.md at line 13 and enforced by the hook in src/hooks/caveman-activate.js.
Where is the auto-clarity rule defined in the Caveman source code?
The primary rule definition lives in src/rules/caveman-activate.md at line 13. The runtime implementation appears in src/hooks/caveman-activate.js at line 137, with additional documentation references in skills/caveman/README.md (line 20) and CLAUDE.md (line 223).
How does Caveman resume compression after a security warning?
The system resumes caveman compression automatically after the safety-critical fragment ends. The "Resume after" directive in the rule definition ensures that once the security warning or irreversible-action confirmation completes, the output stream returns to the compressed caveman style without requiring manual reactivation.
Why does Caveman drop compression for irreversible actions?
Irreversible actions (such as data deletion or system modifications) require exact wording to prevent catastrophic misinterpretation. By dropping to normal prose, Caveman ensures that warnings like "Do NOT run this on production" remain fully explicit and unambiguous, eliminating the risk of critical information loss through aggressive token compression.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →