# What Compression Savings Each Detector Type Targets and Keeps in Caveman

> Discover how Caveman's detectors optimize compression for JSON, Logs, and Terminal output. Eliminate redundancy, retain semantic tokens, and boost LLM inference efficiency.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: deep-dive
- Published: 2026-09-04

---

**Caveman's detection engine classifies payloads in [`engine/detect.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/detect.go) into specific types such as JSON, Log, and Terminal, then applies dedicated compressors that eliminate format-specific redundancy—like whitespace in JSON or ANSI escape sequences in terminal output—while retaining the semantic tokens required for accurate LLM inference.**

The open-source Caveman compression engine optimizes token usage for large language model interactions by tailoring compression strategies to content types. Understanding what compression savings each detector type targets and what it keeps allows developers to predict which data survives the reduction process. At the core of this system is the `Detect` function in [`engine/detect.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/detect.go), which analyzes incoming payloads and assigns a type constant that determines exactly which tokens are discarded and which are preserved.

## JSON and Structured Data Detectors

### JSON (TypeJSON)

The JSON detector targets structural verbosity. According to `engine/detect.go#L13-L19`, the compressor drops whitespace, reorders keys deterministically, and collapses repeated structures to shave thousands of tokens from large JSON blobs. It retains all structural keys, values, and array lengths—the semantic shape necessary for downstream parsing.

### TOON (TypeTOON)

Implemented in [`engine/compressors/toon_encode.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/compressors/toon_encode.go), the TOON detector performs a "tabular-JSON" re-encoding. It stores only the column schema and minimal per-row deltas, achieving significant savings by avoiding repeated key names. The compressor keeps column names and compact row representations (numeric indices and compressed values), allowing full JSON reconstruction on decompression.

### Tabular (TypeTabular)

For comma-separated and table-like data, the Tabular detector (identified via `compressors.LooksTabular`) encodes data as a compact matrix rather than verbose CSV/TSV text. It preserves column headers and numeric/boolean cell values while stripping formatting whitespace and redundant delimiters.

## Code and Diff Detectors

### Code (TypeCode)

The Code detector, defined around `engine/detect.go#L31-L34`, strips harmless comments, blank lines, and language-specific boilerplate. It retains language-significant tokens—including keywords, identifiers, literals, and essential syntactic markers like `{}` or `def …:`—ensuring the source remains parseable by downstream tools.

### Diff (TypeDiff)

As specified in `engine/detect.go#L30-L33`, the Diff compressor collapses unchanged context lines and compresses repeated diff markers. It preserves the "@@ … @@" hunks, file-header lines (`diff --git`, `---`, `+++`), and the actual added/removed lines that convey the semantic change.

## Log and Terminal Output

### Log (TypeLog)

The Log detector (`engine/detect.go#L28-L34`) eliminates noisy repetitions of log levels and timestamps while preserving the unique message content. Notably, it retains log level and timestamp tokens when they are required for faithful reconstruction, alongside distinct message strings carrying novel information.

### Terminal (TypeTerminal)

For terminal output, the compressor (`engine/detect.go#L84-L92`) collapses progress-bar redraws and eliminates duplicate ANSI-escape sequences. It keeps raw ANSI escape codes necessary to reconstruct terminal state and preserves unique command-output lines that carry actual data.

## Web and Configuration Content

### HTML (TypeHTML)

Defined in `engine/detect.go#L63-L70`, the HTML detector removes insignificant whitespace and collapses repetitive tag attributes. The compressor maintains tag hierarchy, rendering-critical attributes, and any embedded scripts required for semantic interpretation.

### Accessibility (TypeA11y)

While handled implicitly by the HTML compressor pipeline, the accessibility detector strips non-essential metadata but preserves required ARIA roles and labels. It retains ARIA attributes, structural headings, and alt-text that conveys meaningful information to assistive technologies.

### Config (TypeConfig)

The Config detector (implemented via `compressors.LooksConfig`) targets configuration files such as YAML and INI. It drops comments and whitespace while keeping all configuration keys and their assigned values, ensuring the functional behavior remains intact.

## Generic and Search Data

### Search Result (TypeSearchResult)

For search results, the compressor (`engine/detect.go#L31-L33`) shortens long path listings and repetitive URL fragments. It preserves unique path or URL components that differentiate results, plus any highlighted snippets required for relevance scoring.

### Text (TypeText)

The Text detector applies conservative compression, removing only obvious whitespace redundancy. As noted in `engine/detect.go#L81-L82`, it keeps plain-text characters intact while collapsing excessive whitespace and line-breaks.

## How the Detection Engine Routes Payloads

The classification and routing logic resides in [`engine/detect.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/detect.go). Once `engine.Detect(payload)` returns a type constant, the engine instantiates the matching compressor from the `compressors` package:

```go
// Detect the content type.
engine := caveman.NewEngine()
typ := engine.Detect(payload) // typ is one of the constants above.

// Route to the matching compressor.
var comp compressor.Compressor
switch typ {
case engine.TypeJSON:
    comp = compressors.NewJSON()
case engine.TypeLog:
    comp = compressors.NewLog()
case engine.TypeCode:
    comp = compressors.NewCode()
case engine.TypeDiff:
    comp = compressors.NewDiff()
case engine.TypeTOON:
    comp = compressors.NewTOON() // the tabular-JSON encoder.
 // … other cases …
}

// Perform the compression.
compressed, stats := comp.Compress(payload)

// `stats` reports token savings for this detector type.
fmt.Printf("%s saved %d tokens (≈%0.1f%%)\n",
    typ, stats.TokensSaved, stats.SavingsPct)

```

The `stats` object defined in [`engine/result.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/result.go) exposes `TokensSaved` and `SavingsPct`, allowing precise verification that each detector type meets its specific savings target.

## Summary

- **Type-specific targeting**: Caveman's `Detect` routine in [`engine/detect.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/detect.go) assigns one of twelve detector types, each optimized for distinct redundancy patterns found in JSON, logs, code, diffs, and terminal output.
- **Semantic preservation**: Every compressor retains the minimal token set required for downstream utility—structural keys for JSON, syntactic markers for code, ANSI codes for terminals—while discarding formatting noise.
- **Measurable savings**: The [`engine/result.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/result.go) statistics interface reports exact token reductions per detector type, enabling verification of compression efficacy.
- **Specialized encodings**: Unique strategies like the TOON tabular-JSON encoder ([`engine/compressors/toon_encode.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/compressors/toon_encode.go)) and the Tabular matrix format achieve dramatic savings by re-encoding data structures rather than merely removing characters.

## Frequently Asked Questions

### How does Caveman determine which detector type to apply?

Caveman's [`engine/detect.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/detect.go) implements deterministic content analysis that inspects payload structure, prefix patterns, and formatting characteristics to assign a type constant (e.g., `TypeJSON`, `TypeLog`). This classification then triggers the corresponding compressor in the `engine/compressors/` package via a type switch, ensuring each payload is processed by the algorithm optimized for its specific format.

### What differentiates the TOON detector from the standard Tabular detector?

While both handle structured data, **TOON** (`TypeTOON`) specifically targets JSON-like tabular structures by re-encoding them as column schemas with per-row deltas, dramatically reducing token count for JSON arrays. The **Tabular** detector (`TypeTabular`) focuses on raw CSV/TSV formats, encoding them as compact matrices without converting to JSON intermediary, making it ideal for legacy delimiter-separated files.

### Does Caveman preserve comments when compressing source code?

No. According to the Code detector implementation around `engine/detect.go#L31-L34`, the compressor deliberately strips comments, blank lines, and boilerplate to maximize token savings. It retains only language-significant tokens—keywords, identifiers, literals, and structural syntax markers—necessary for the source to remain parseable by compilers and interpreters.

### How can I verify the compression savings for a specific detector type?

After compression, the `comp.Compress(payload)` method returns a `stats` object (defined in [`engine/result.go`](https://github.com/JuliusBrussee/caveman/blob/main/engine/result.go)) containing `TokensSaved` and `SavingsPct` fields. By printing or logging these values as shown in the routing example, you can verify exactly how many tokens each detector type removed from your specific payload.