# Caveman Compression Engine Input Type Detection: JSON, Binary, and Text Classification

> Discover how the Caveman compression engine intelligently detects input types like JSON, binary, and text using a three-step heuristic for efficient data compression. Learn more.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: internals
- Published: 2026-09-06

---

**The Caveman compression engine detects input types through a sequential three-step heuristic implemented in the TypeScript SDK's `compress` method, classifying payloads as JSON, Base-64 binary, or plain text before routing to the native engine.**

The JuliusBrussee/caveman repository implements client-side input classification to optimize compression strategies for different data formats. Located in [`packages/sdk/typescript/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts) (lines 1650‑1675), the input type detection logic executes entirely within the SDK prior to transmission via `POST /sdk/v1/compress`. This design allows the stateless `caveman‑engine` process to operate efficiently without additional content sniffing, trusting the SDK's classification to either compress textual content or pass binary data through unchanged.

## How Caveman Detects Input Types

The detection algorithm proceeds through a strict priority order: JSON parsing is attempted first, followed by Base-64 validation, with plain text serving as the default fallback. This hierarchy ensures that structured data receives appropriate field-level compression while binary payloads are identified and excluded from processing.

### Step 1: JSON Structure Validation

The SDK begins classification by attempting to parse the raw string using `JSON.parse`. If parsing succeeds and returns an object or array (verified via `typeof parsed === "object"`), the engine immediately marks the input as **JSON**. This classification enables the "toon" compression mode to operate on specific object fields rather than treating the payload as an opaque string.

### Step 2: Base-64 Binary Pattern Matching

When JSON parsing throws an exception, the SDK checks for Base-64 encoded binary data using the regex `/^[A-Za-z0-9+/=]+$/` combined with a modulus validation (`payload.length % 4 === 0`). If the pattern matches, the SDK attempts `Buffer.from(payload, "base64")` conversion and scans the resulting buffer for non-printable byte values. Successful validation classifies the input as **binary**, triggering an immediate pass-through with zero token savings reported.

### Step 3: Plain Text Fallback Classification

If the payload fails both structured data and binary detection heuristics, the SDK defaults to **plain text** classification. The engine then selects either the "elision" or "toon" compression method based on the session configuration and processes the textual content for token reduction.

## Source Code Implementation in src/index.ts

The detection logic resides in the `compress` implementation within the TypeScript SDK. The following excerpt from [`packages/sdk/typescript/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts) illustrates the sequential classification strategy:

```typescript
let inputKind: "json" | "binary" | "text";
try {
  const parsed = JSON.parse(payload);
  if (typeof parsed === "object") inputKind = "json";
} catch {
  // not JSON → try binary detection
  if (/^[A-Za-z0-9+/=]+$/.test(payload) && payload.length % 4 === 0) {
    try {
      const buf = Buffer.from(payload, "base64");
      if (buf.some(b => b < 32 && b > 126)) inputKind = "binary";
    } catch {}
  }
}
if (!inputKind) inputKind = "text";

```

After classification, the SDK forwards the payload to the native engine. The engine itself never re‑interprets the data; it trusts the SDK's classification and either compresses the content or returns the original payload unchanged for binary inputs.

## Engine Behavior After Classification

The classification determines the specific compression pathway:

- **JSON payloads**: Handed to the engine as structured data, allowing field‑specific compression algorithms to reduce token count by abbreviating keys and values.
- **Plain text**: Processed through the default text compressor using elision or toon methods based on configuration.
- **Binary data**: Passes through the engine untouched. Because compression only works on textual content, the engine reports **0 tokens saved** and returns the original Base-64 string.

## Practical Implementation Examples

The following examples demonstrate how the SDK handles different input types in practice, utilizing the same detection logic found in [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts) when processing CLI inputs.

### Compressing Structured JSON Data

```typescript
import { Cave } from "@caveman/sdk";

const cave = new Cave({ baseURL: "https://caveman.example.com" });
const json = JSON.stringify({ prompt: "Write a haiku", temperature: 0.7 });

const result = await cave.compress(json);
console.log(result.compressedPayload);   // compressed JSON (or original if pass-through)

```

The SDK detects the string as JSON, enabling the engine to apply "toon" compression if configured for the session.

### Processing Plain Text Content

```typescript
const txt = "Explain the theory of relativity in simple terms.";
const result = await cave.compress(txt);
console.log(result.compressedPayload);   // usually a shortened version

```

Since the payload fails both JSON and Base-64 checks, the SDK tags it as plain text and the engine runs the default text compressor.

### Handling Base-64 Binary Blobs

```typescript
const binaryB64 = "iVBORw0KGgoAAAANSUhEUgAA...";   // PNG image data
const result = await cave.compress(binaryB64);
console.log(result.compressedPayload === binaryB64); // true – no compression applied

```

The SDK recognizes the Base-64 pattern, validates it as binary using the buffer inspection logic, and instructs the engine to skip compression. Test fixtures in [`engine/evals/fixtures/sample.ts`](https://github.com/JuliusBrussee/caveman/blob/main/engine/evals/fixtures/sample.ts) provide additional sample payloads for each input type used in validation.

## Summary

- **Three-step hierarchy**: The SDK attempts JSON parsing first, then Base-64/binary validation, defaulting to plain text if both fail.
- **Client-side processing**: All detection occurs in [`packages/sdk/typescript/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts) before transmission to the native engine, keeping the server-side process stateless.
- **Binary pass-through**: Base-64 detected binary data bypasses compression entirely, returning zero token savings since the algorithms target textual patterns.
- **Repository integration**: The CLI entry point at [`packages/cli/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/cli/src/index.ts) utilizes the same SDK methods, ensuring consistent classification across all interfaces.

## Frequently Asked Questions

### How does the Caveman SDK distinguish between JSON and plain text?

The SDK attempts `JSON.parse()` on the raw payload and verifies that the result is an object or array using `typeof` checks. If parsing succeeds and returns a structural type, the input is classified as JSON; otherwise, it proceeds to binary detection or falls back to text classification.

### Why does the Caveman engine skip compression on binary data?

Compression algorithms in the Caveman engine are designed to reduce tokens in textual content by identifying patterns, abbreviations, and structural redundancies. Binary data contains no such compressible textual patterns, so the engine performs a pass-through, reporting zero token savings to avoid wasting computational resources.

### Where is the input type detection logic located in the codebase?

The detection logic is implemented in [`packages/sdk/typescript/src/index.ts`](https://github.com/JuliusBrussee/caveman/blob/main/packages/sdk/typescript/src/index.ts) within the `compress` method (approximately lines 1650‑1675). This client-side TypeScript code executes before the SDK sends data to the native engine via the `POST /sdk/v1/compress` endpoint.

### Can the detection heuristic misclassify input types?

While the sequential checks (JSON parse → Base-64 regex → text fallback) are robust for standard inputs, edge cases exist where valid Base-64 text containing specific byte patterns might be classified as binary, or highly structured text might attempt JSON parsing. However, the engine handles these gracefully through appropriate compression or pass-through behavior without data loss.