What Compression Savings Each Detector Type Targets and Keeps in Caveman
Caveman's detection engine classifies payloads in engine/detect.go into specific types such as JSON, Log, and Terminal, then applies dedicated compressors that eliminate format-specific redundancy—like whitespace in JSON or ANSI escape sequences in terminal output—while retaining the semantic tokens required for accurate LLM inference.
The open-source Caveman compression engine optimizes token usage for large language model interactions by tailoring compression strategies to content types. Understanding what compression savings each detector type targets and what it keeps allows developers to predict which data survives the reduction process. At the core of this system is the Detect function in engine/detect.go, which analyzes incoming payloads and assigns a type constant that determines exactly which tokens are discarded and which are preserved.
JSON and Structured Data Detectors
JSON (TypeJSON)
The JSON detector targets structural verbosity. According to engine/detect.go#L13-L19, the compressor drops whitespace, reorders keys deterministically, and collapses repeated structures to shave thousands of tokens from large JSON blobs. It retains all structural keys, values, and array lengths—the semantic shape necessary for downstream parsing.
TOON (TypeTOON)
Implemented in engine/compressors/toon_encode.go, the TOON detector performs a "tabular-JSON" re-encoding. It stores only the column schema and minimal per-row deltas, achieving significant savings by avoiding repeated key names. The compressor keeps column names and compact row representations (numeric indices and compressed values), allowing full JSON reconstruction on decompression.
Tabular (TypeTabular)
For comma-separated and table-like data, the Tabular detector (identified via compressors.LooksTabular) encodes data as a compact matrix rather than verbose CSV/TSV text. It preserves column headers and numeric/boolean cell values while stripping formatting whitespace and redundant delimiters.
Code and Diff Detectors
Code (TypeCode)
The Code detector, defined around engine/detect.go#L31-L34, strips harmless comments, blank lines, and language-specific boilerplate. It retains language-significant tokens—including keywords, identifiers, literals, and essential syntactic markers like {} or def …:—ensuring the source remains parseable by downstream tools.
Diff (TypeDiff)
As specified in engine/detect.go#L30-L33, the Diff compressor collapses unchanged context lines and compresses repeated diff markers. It preserves the "@@ … @@" hunks, file-header lines (diff --git, ---, +++), and the actual added/removed lines that convey the semantic change.
Log and Terminal Output
Log (TypeLog)
The Log detector (engine/detect.go#L28-L34) eliminates noisy repetitions of log levels and timestamps while preserving the unique message content. Notably, it retains log level and timestamp tokens when they are required for faithful reconstruction, alongside distinct message strings carrying novel information.
Terminal (TypeTerminal)
For terminal output, the compressor (engine/detect.go#L84-L92) collapses progress-bar redraws and eliminates duplicate ANSI-escape sequences. It keeps raw ANSI escape codes necessary to reconstruct terminal state and preserves unique command-output lines that carry actual data.
Web and Configuration Content
HTML (TypeHTML)
Defined in engine/detect.go#L63-L70, the HTML detector removes insignificant whitespace and collapses repetitive tag attributes. The compressor maintains tag hierarchy, rendering-critical attributes, and any embedded scripts required for semantic interpretation.
Accessibility (TypeA11y)
While handled implicitly by the HTML compressor pipeline, the accessibility detector strips non-essential metadata but preserves required ARIA roles and labels. It retains ARIA attributes, structural headings, and alt-text that conveys meaningful information to assistive technologies.
Config (TypeConfig)
The Config detector (implemented via compressors.LooksConfig) targets configuration files such as YAML and INI. It drops comments and whitespace while keeping all configuration keys and their assigned values, ensuring the functional behavior remains intact.
Generic and Search Data
Search Result (TypeSearchResult)
For search results, the compressor (engine/detect.go#L31-L33) shortens long path listings and repetitive URL fragments. It preserves unique path or URL components that differentiate results, plus any highlighted snippets required for relevance scoring.
Text (TypeText)
The Text detector applies conservative compression, removing only obvious whitespace redundancy. As noted in engine/detect.go#L81-L82, it keeps plain-text characters intact while collapsing excessive whitespace and line-breaks.
How the Detection Engine Routes Payloads
The classification and routing logic resides in engine/detect.go. Once engine.Detect(payload) returns a type constant, the engine instantiates the matching compressor from the compressors package:
// Detect the content type.
engine := caveman.NewEngine()
typ := engine.Detect(payload) // typ is one of the constants above.
// Route to the matching compressor.
var comp compressor.Compressor
switch typ {
case engine.TypeJSON:
comp = compressors.NewJSON()
case engine.TypeLog:
comp = compressors.NewLog()
case engine.TypeCode:
comp = compressors.NewCode()
case engine.TypeDiff:
comp = compressors.NewDiff()
case engine.TypeTOON:
comp = compressors.NewTOON() // the tabular-JSON encoder.
// … other cases …
}
// Perform the compression.
compressed, stats := comp.Compress(payload)
// `stats` reports token savings for this detector type.
fmt.Printf("%s saved %d tokens (≈%0.1f%%)\n",
typ, stats.TokensSaved, stats.SavingsPct)
The stats object defined in engine/result.go exposes TokensSaved and SavingsPct, allowing precise verification that each detector type meets its specific savings target.
Summary
- Type-specific targeting: Caveman's
Detectroutine inengine/detect.goassigns one of twelve detector types, each optimized for distinct redundancy patterns found in JSON, logs, code, diffs, and terminal output. - Semantic preservation: Every compressor retains the minimal token set required for downstream utility—structural keys for JSON, syntactic markers for code, ANSI codes for terminals—while discarding formatting noise.
- Measurable savings: The
engine/result.gostatistics interface reports exact token reductions per detector type, enabling verification of compression efficacy. - Specialized encodings: Unique strategies like the TOON tabular-JSON encoder (
engine/compressors/toon_encode.go) and the Tabular matrix format achieve dramatic savings by re-encoding data structures rather than merely removing characters.
Frequently Asked Questions
How does Caveman determine which detector type to apply?
Caveman's engine/detect.go implements deterministic content analysis that inspects payload structure, prefix patterns, and formatting characteristics to assign a type constant (e.g., TypeJSON, TypeLog). This classification then triggers the corresponding compressor in the engine/compressors/ package via a type switch, ensuring each payload is processed by the algorithm optimized for its specific format.
What differentiates the TOON detector from the standard Tabular detector?
While both handle structured data, TOON (TypeTOON) specifically targets JSON-like tabular structures by re-encoding them as column schemas with per-row deltas, dramatically reducing token count for JSON arrays. The Tabular detector (TypeTabular) focuses on raw CSV/TSV formats, encoding them as compact matrices without converting to JSON intermediary, making it ideal for legacy delimiter-separated files.
Does Caveman preserve comments when compressing source code?
No. According to the Code detector implementation around engine/detect.go#L31-L34, the compressor deliberately strips comments, blank lines, and boilerplate to maximize token savings. It retains only language-significant tokens—keywords, identifiers, literals, and structural syntax markers—necessary for the source to remain parseable by compilers and interpreters.
How can I verify the compression savings for a specific detector type?
After compression, the comp.Compress(payload) method returns a stats object (defined in engine/result.go) containing TokensSaved and SavingsPct fields. By printing or logging these values as shown in the routing example, you can verify exactly how many tokens each detector type removed from your specific payload.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →