Caveman Compression Engine Input Type Detection: JSON, Binary, and Text Classification

The Caveman compression engine detects input types through a sequential three-step heuristic implemented in the TypeScript SDK's compress method, classifying payloads as JSON, Base-64 binary, or plain text before routing to the native engine.

The JuliusBrussee/caveman repository implements client-side input classification to optimize compression strategies for different data formats. Located in packages/sdk/typescript/src/index.ts (lines 1650‑1675), the input type detection logic executes entirely within the SDK prior to transmission via POST /sdk/v1/compress. This design allows the stateless caveman‑engine process to operate efficiently without additional content sniffing, trusting the SDK's classification to either compress textual content or pass binary data through unchanged.

How Caveman Detects Input Types

The detection algorithm proceeds through a strict priority order: JSON parsing is attempted first, followed by Base-64 validation, with plain text serving as the default fallback. This hierarchy ensures that structured data receives appropriate field-level compression while binary payloads are identified and excluded from processing.

Step 1: JSON Structure Validation

The SDK begins classification by attempting to parse the raw string using JSON.parse. If parsing succeeds and returns an object or array (verified via typeof parsed === "object"), the engine immediately marks the input as JSON. This classification enables the "toon" compression mode to operate on specific object fields rather than treating the payload as an opaque string.

Step 2: Base-64 Binary Pattern Matching

When JSON parsing throws an exception, the SDK checks for Base-64 encoded binary data using the regex /^[A-Za-z0-9+/=]+$/ combined with a modulus validation (payload.length % 4 === 0). If the pattern matches, the SDK attempts Buffer.from(payload, "base64") conversion and scans the resulting buffer for non-printable byte values. Successful validation classifies the input as binary, triggering an immediate pass-through with zero token savings reported.

Step 3: Plain Text Fallback Classification

If the payload fails both structured data and binary detection heuristics, the SDK defaults to plain text classification. The engine then selects either the "elision" or "toon" compression method based on the session configuration and processes the textual content for token reduction.

Source Code Implementation in src/index.ts

The detection logic resides in the compress implementation within the TypeScript SDK. The following excerpt from packages/sdk/typescript/src/index.ts illustrates the sequential classification strategy:

let inputKind: "json" | "binary" | "text";
try {
  const parsed = JSON.parse(payload);
  if (typeof parsed === "object") inputKind = "json";
} catch {
  // not JSON → try binary detection
  if (/^[A-Za-z0-9+/=]+$/.test(payload) && payload.length % 4 === 0) {
    try {
      const buf = Buffer.from(payload, "base64");
      if (buf.some(b => b < 32 && b > 126)) inputKind = "binary";
    } catch {}
  }
}
if (!inputKind) inputKind = "text";

After classification, the SDK forwards the payload to the native engine. The engine itself never re‑interprets the data; it trusts the SDK's classification and either compresses the content or returns the original payload unchanged for binary inputs.

Engine Behavior After Classification

The classification determines the specific compression pathway:

  • JSON payloads: Handed to the engine as structured data, allowing field‑specific compression algorithms to reduce token count by abbreviating keys and values.
  • Plain text: Processed through the default text compressor using elision or toon methods based on configuration.
  • Binary data: Passes through the engine untouched. Because compression only works on textual content, the engine reports 0 tokens saved and returns the original Base-64 string.

Practical Implementation Examples

The following examples demonstrate how the SDK handles different input types in practice, utilizing the same detection logic found in packages/cli/src/index.ts when processing CLI inputs.

Compressing Structured JSON Data

import { Cave } from "@caveman/sdk";

const cave = new Cave({ baseURL: "https://caveman.example.com" });
const json = JSON.stringify({ prompt: "Write a haiku", temperature: 0.7 });

const result = await cave.compress(json);
console.log(result.compressedPayload);   // compressed JSON (or original if pass-through)

The SDK detects the string as JSON, enabling the engine to apply "toon" compression if configured for the session.

Processing Plain Text Content

const txt = "Explain the theory of relativity in simple terms.";
const result = await cave.compress(txt);
console.log(result.compressedPayload);   // usually a shortened version

Since the payload fails both JSON and Base-64 checks, the SDK tags it as plain text and the engine runs the default text compressor.

Handling Base-64 Binary Blobs

const binaryB64 = "iVBORw0KGgoAAAANSUhEUgAA...";   // PNG image data
const result = await cave.compress(binaryB64);
console.log(result.compressedPayload === binaryB64); // true – no compression applied

The SDK recognizes the Base-64 pattern, validates it as binary using the buffer inspection logic, and instructs the engine to skip compression. Test fixtures in engine/evals/fixtures/sample.ts provide additional sample payloads for each input type used in validation.

Summary

  • Three-step hierarchy: The SDK attempts JSON parsing first, then Base-64/binary validation, defaulting to plain text if both fail.
  • Client-side processing: All detection occurs in packages/sdk/typescript/src/index.ts before transmission to the native engine, keeping the server-side process stateless.
  • Binary pass-through: Base-64 detected binary data bypasses compression entirely, returning zero token savings since the algorithms target textual patterns.
  • Repository integration: The CLI entry point at packages/cli/src/index.ts utilizes the same SDK methods, ensuring consistent classification across all interfaces.

Frequently Asked Questions

How does the Caveman SDK distinguish between JSON and plain text?

The SDK attempts JSON.parse() on the raw payload and verifies that the result is an object or array using typeof checks. If parsing succeeds and returns a structural type, the input is classified as JSON; otherwise, it proceeds to binary detection or falls back to text classification.

Why does the Caveman engine skip compression on binary data?

Compression algorithms in the Caveman engine are designed to reduce tokens in textual content by identifying patterns, abbreviations, and structural redundancies. Binary data contains no such compressible textual patterns, so the engine performs a pass-through, reporting zero token savings to avoid wasting computational resources.

Where is the input type detection logic located in the codebase?

The detection logic is implemented in packages/sdk/typescript/src/index.ts within the compress method (approximately lines 1650‑1675). This client-side TypeScript code executes before the SDK sends data to the native engine via the POST /sdk/v1/compress endpoint.

Can the detection heuristic misclassify input types?

While the sequential checks (JSON parse → Base-64 regex → text fallback) are robust for standard inputs, edge cases exist where valid Base-64 text containing specific byte patterns might be classified as binary, or highly structured text might attempt JSON parsing. However, the engine handles these gracefully through appropriate compression or pass-through behavior without data loss.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →