How the FreeLLMAPI Prompt Compression Pipeline Deduplicates, Filters, and Compacts Prompts

The FreeLLMAPI prompt compression pipeline uses configurable modes to sequentially run specialized engines—including dedup for block-level deduplication and jsoncompact for JSON compaction—while a fidelity gate ensures lossless transformations.

The tashfeenahmed/freellmapi repository implements a sophisticated prompt compression pipeline that reduces token count without sacrificing semantic meaning. Located in server/src/services/compression/pipeline.ts, this system processes incoming chat requests through a series of compression engines that target specific redundancy patterns. The pipeline supports four distinct modes—off, lossless, standard, and aggressive—allowing developers to balance compression ratio against content preservation.

How the Prompt Compression Pipeline Works

The pipeline orchestrates transformation by first stabilizing critical context, then iterating through registered engines based on the active mode.

Mode Selection and Prefix Freezing

When compressRequest() receives a message array, it immediately identifies stable system prefixes via stablePrefixIndexes. These early system messages are frozen and excluded from modification, preventing the loss of crucial instructions or persona definitions. The pipeline then selects a compression mode based on user headers or auto-trigger rules, consulting the MODE_ENGINES map to determine which engines to activate.

Engine Orchestration

The MODE_ENGINES configuration in pipeline.ts defines which engines run for each mode:

  • lossless: dedup, lite, jsoncompact
  • standard: dedup, lite, read-lifecycle, toolfilter, jsoncompact
  • aggressive: All standard engines plus relevance, aging, and hard-budget

Each engine implements a uniform interface that accepts messages, configuration options, and frozen index sets, returning transformed content along with operation statistics.

Deduplication with the dedup Engine

The dedup engine, implemented in server/src/services/compression/engines/dedup.ts, eliminates repetitive text blocks across conversation history while preserving the first occurrence.

Block Detection and Hashing

The engine splits non-system messages on double-line breaks using the regex /(\n{2,})/. A block qualifies for deduplication only if it exceeds 80 characters (minBlockChars) and spans at least 3 lines (minBlockLines). Qualified blocks are hashed using SHA-256 (hash()) to create unique fingerprints for comparison.

Replacement Strategy

When the engine encounters a hash match, it retains the original block in its position and replaces subsequent duplicates with a concise placeholder: [repeated from earlier in this conversation: …]. The placeholder includes a shortened preview generated via preview() to maintain context. The engine returns the transformed message array and a blocksReplaced count. Because dedup specifies lossless: true and respects frozenMessageIndexes, stable system prefixes remain untouched.

JSON Compaction with the jsoncompact Engine

Located in server/src/services/compression/engines/jsoncompact.ts, this engine targets tool-generated JSON payloads to reduce verbose array structures.

Detecting Uniform Arrays

The engine scans messages with role tool using jsonArrayCandidates() to identify syntactically balanced [ … ] spans. Each candidate undergoes JSON.parse validation with budget constraints to prevent pathological inputs. The uniformKeys() function verifies that an array contains objects sharing identical key sets and meets the minRows threshold of 8 rows.

Table Encoding

Qualifying arrays are transformed via encodeJsonTable() into compact markers: [[json-table:v1 …]]. The marker stores column names once, followed by row data as simple arrays, dramatically reducing token count. The engine replaces the original JSON only when the encoded representation is shorter, inserting multiple non-overlapping tables while preserving message order. The operation returns tables count and maintains lossless: true fidelity.

Fidelity Verification and Statistics

After each engine execution, checkFidelity() (from fidelity-gate.ts) validates that compression preserved protected tokens, numeric literals, JSON keys, and diff hunks. If an engine violates fidelity constraints, its changes are discarded and recorded in discardedByGate. The pipeline accumulates metrics into a CompressionRequestStats object, tracking input/output characters, saved tokens, and per-engine metadata like blocksReplaced or tables counts.

Configuring the Compression Pipeline

You can invoke compression programmatically using the pipeline's exported functions:

// Compress a request using default lossless settings
import { compressRequest } from './services/compression/pipeline.js';

const messages = [
  { role: 'system', content: 'You are a helpful assistant.' },
  { role: 'user', content: 'Explain recursion.\n\nRecursion is a process...' },
  { role: 'assistant', content: 'Recursion is a process...' }
];

const result = compressRequest(messages);
console.log(result.stats.enginesApplied); // ['dedup', 'jsoncompact']
console.log(result.stats.discardedByGate); // [] if fidelity passed

To debug specific engines or customize behavior:

// Enable only deduplication for debugging
import { setCompressionConfig } from './services/compression/config.js';
import { compressRequest } from './services/compression/pipeline.js';

setCompressionConfig({
  mode: 'lossless',
  engines: { 
    dedup: { enabled: true, minBlockChars: 50 }, 
    jsoncompact: { enabled: false } 
  },
});

const result = compressRequest(messages);
console.log(result.stats.enginesApplied); // ['dedup']

To decode compacted JSON tables back to standard JSON:

import { decodeJsonTables } from './services/compression/engines/jsoncompact.js';

const compacted = `[[json-table:v1 columns=["id","value"] rows=3]]
["1","a"]
["2","b"]
["3","c"]
[[/json-table]]`;

const restored = decodeJsonTables(compacted);
// Returns: [{"id":"1","value":"a"}, {"id":"2","value":"b"}, {"id":"3","value":"c"}]

Summary

  • The prompt compression pipeline in pipeline.ts orchestrates mode selection, prefix freezing, and sequential engine execution.
  • The dedup engine eliminates repetitive text blocks using SHA-256 hashing and replacement markers, preserving the first occurrence of each block.
  • The jsoncompact engine converts large uniform JSON arrays into compact table markers, reducing token count while maintaining data integrity.
  • A fidelity gate ensures all transformations remain lossless and protect critical content like JSON keys and numeric literals.
  • Four compression modes—lossless, standard, aggressive, and off—provide flexible trade-offs between compression ratio and processing intensity.

Frequently Asked Questions

What is the minimum block size for deduplication in the FreeLLMAPI pipeline?

The dedup engine requires blocks to exceed 80 characters (minBlockChars) and span at least 3 lines (minBlockLines) before considering them for deduplication. These thresholds prevent the replacement of short, potentially significant text fragments.

How does the pipeline ensure system prompts are not modified?

The pipeline calculates stablePrefixIndexes to identify system messages at the conversation start, freezing them in the frozenMessageIndexes set. Both the dedup and jsoncompact engines check this set before modification, ensuring critical instructions remain intact.

Can I use the JSON compaction engine on non-tool messages?

No, the jsoncompact engine specifically targets messages with role tool. It ignores user and assistant messages, focusing exclusively on tool-generated JSON payloads that typically contain large, structured data arrays suitable for tabular encoding.

What happens if a compression engine corrupts the prompt?

The checkFidelity function in fidelity-gate.ts validates each engine's output against protected token patterns. If corruption is detected, the engine's changes are discarded, the engine name is added to discardedByGate, and the pipeline continues with the unmodified message state from the previous step.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →