# How the FreeLLMAPI Prompt Compression Pipeline Deduplicates, Filters, and Compacts Prompts

> Discover how the FreeLLMAPI prompt compression pipeline uses configurable modes like dedup and jsoncompact to efficiently deduplicate, filter, and compact prompts while ensuring lossless transformations.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: internals
- Published: 2026-08-31

---

**The FreeLLMAPI prompt compression pipeline uses configurable modes to sequentially run specialized engines—including `dedup` for block-level deduplication and `jsoncompact` for JSON compaction—while a fidelity gate ensures lossless transformations.**

The `tashfeenahmed/freellmapi` repository implements a sophisticated prompt compression pipeline that reduces token count without sacrificing semantic meaning. Located in [`server/src/services/compression/pipeline.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/pipeline.ts), this system processes incoming chat requests through a series of **compression engines** that target specific redundancy patterns. The pipeline supports four distinct modes—`off`, `lossless`, `standard`, and `aggressive`—allowing developers to balance compression ratio against content preservation.

## How the Prompt Compression Pipeline Works

The pipeline orchestrates transformation by first stabilizing critical context, then iterating through registered engines based on the active mode.

### Mode Selection and Prefix Freezing

When `compressRequest()` receives a message array, it immediately identifies **stable system prefixes** via `stablePrefixIndexes`. These early system messages are frozen and excluded from modification, preventing the loss of crucial instructions or persona definitions. The pipeline then selects a compression mode based on user headers or auto-trigger rules, consulting the `MODE_ENGINES` map to determine which engines to activate.

### Engine Orchestration

The `MODE_ENGINES` configuration in [`pipeline.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/pipeline.ts) defines which engines run for each mode:

- **`lossless`**: `dedup`, `lite`, `jsoncompact`
- **`standard`**: `dedup`, `lite`, `read-lifecycle`, `toolfilter`, `jsoncompact`
- **`aggressive`**: All standard engines plus `relevance`, `aging`, and `hard-budget`

Each engine implements a uniform interface that accepts messages, configuration options, and frozen index sets, returning transformed content along with operation statistics.

## Deduplication with the `dedup` Engine

The `dedup` engine, implemented in [`server/src/services/compression/engines/dedup.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/engines/dedup.ts), eliminates repetitive text blocks across conversation history while preserving the first occurrence.

### Block Detection and Hashing

The engine splits non-system messages on double-line breaks using the regex `/(\n{2,})/`. A block qualifies for deduplication only if it exceeds **80 characters** (`minBlockChars`) and spans at least **3 lines** (`minBlockLines`). Qualified blocks are hashed using **SHA-256** (`hash()`) to create unique fingerprints for comparison.

### Replacement Strategy

When the engine encounters a hash match, it retains the original block in its position and replaces subsequent duplicates with a concise placeholder: `[repeated from earlier in this conversation: …]`. The placeholder includes a shortened preview generated via `preview()` to maintain context. The engine returns the transformed message array and a `blocksReplaced` count. Because `dedup` specifies `lossless: true` and respects `frozenMessageIndexes`, stable system prefixes remain untouched.

## JSON Compaction with the `jsoncompact` Engine

Located in [`server/src/services/compression/engines/jsoncompact.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/engines/jsoncompact.ts), this engine targets tool-generated JSON payloads to reduce verbose array structures.

### Detecting Uniform Arrays

The engine scans messages with role `tool` using `jsonArrayCandidates()` to identify syntactically balanced `[` … `]` spans. Each candidate undergoes `JSON.parse` validation with budget constraints to prevent pathological inputs. The `uniformKeys()` function verifies that an array contains objects sharing identical key sets and meets the `minRows` threshold of **8 rows**.

### Table Encoding

Qualifying arrays are transformed via `encodeJsonTable()` into compact markers: `[[json-table:v1 …]]`. The marker stores column names once, followed by row data as simple arrays, dramatically reducing token count. The engine replaces the original JSON only when the encoded representation is shorter, inserting multiple non-overlapping tables while preserving message order. The operation returns `tables` count and maintains `lossless: true` fidelity.

## Fidelity Verification and Statistics

After each engine execution, `checkFidelity()` (from [`fidelity-gate.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fidelity-gate.ts)) validates that compression preserved protected tokens, numeric literals, JSON keys, and diff hunks. If an engine violates fidelity constraints, its changes are discarded and recorded in `discardedByGate`. The pipeline accumulates metrics into a `CompressionRequestStats` object, tracking input/output characters, saved tokens, and per-engine metadata like `blocksReplaced` or `tables` counts.

## Configuring the Compression Pipeline

You can invoke compression programmatically using the pipeline's exported functions:

```typescript
// Compress a request using default lossless settings
import { compressRequest } from './services/compression/pipeline.js';

const messages = [
  { role: 'system', content: 'You are a helpful assistant.' },
  { role: 'user', content: 'Explain recursion.\n\nRecursion is a process...' },
  { role: 'assistant', content: 'Recursion is a process...' }
];

const result = compressRequest(messages);
console.log(result.stats.enginesApplied); // ['dedup', 'jsoncompact']
console.log(result.stats.discardedByGate); // [] if fidelity passed

```

To debug specific engines or customize behavior:

```typescript
// Enable only deduplication for debugging
import { setCompressionConfig } from './services/compression/config.js';
import { compressRequest } from './services/compression/pipeline.js';

setCompressionConfig({
  mode: 'lossless',
  engines: { 
    dedup: { enabled: true, minBlockChars: 50 }, 
    jsoncompact: { enabled: false } 
  },
});

const result = compressRequest(messages);
console.log(result.stats.enginesApplied); // ['dedup']

```

To decode compacted JSON tables back to standard JSON:

```typescript
import { decodeJsonTables } from './services/compression/engines/jsoncompact.js';

const compacted = `[[json-table:v1 columns=["id","value"] rows=3]]
["1","a"]
["2","b"]
["3","c"]
[[/json-table]]`;

const restored = decodeJsonTables(compacted);
// Returns: [{"id":"1","value":"a"}, {"id":"2","value":"b"}, {"id":"3","value":"c"}]

```

## Summary

- The **prompt compression pipeline** in [`pipeline.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/pipeline.ts) orchestrates mode selection, prefix freezing, and sequential engine execution.
- The **`dedup` engine** eliminates repetitive text blocks using SHA-256 hashing and replacement markers, preserving the first occurrence of each block.
- The **`jsoncompact` engine** converts large uniform JSON arrays into compact table markers, reducing token count while maintaining data integrity.
- A **fidelity gate** ensures all transformations remain lossless and protect critical content like JSON keys and numeric literals.
- Four compression modes—`lossless`, `standard`, `aggressive`, and `off`—provide flexible trade-offs between compression ratio and processing intensity.

## Frequently Asked Questions

### What is the minimum block size for deduplication in the FreeLLMAPI pipeline?

The `dedup` engine requires blocks to exceed **80 characters** (`minBlockChars`) and span at least **3 lines** (`minBlockLines`) before considering them for deduplication. These thresholds prevent the replacement of short, potentially significant text fragments.

### How does the pipeline ensure system prompts are not modified?

The pipeline calculates `stablePrefixIndexes` to identify system messages at the conversation start, freezing them in the `frozenMessageIndexes` set. Both the `dedup` and `jsoncompact` engines check this set before modification, ensuring critical instructions remain intact.

### Can I use the JSON compaction engine on non-tool messages?

No, the `jsoncompact` engine specifically targets messages with role `tool`. It ignores user and assistant messages, focusing exclusively on tool-generated JSON payloads that typically contain large, structured data arrays suitable for tabular encoding.

### What happens if a compression engine corrupts the prompt?

The `checkFidelity` function in [`fidelity-gate.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fidelity-gate.ts) validates each engine's output against protected token patterns. If corruption is detected, the engine's changes are discarded, the engine name is added to `discardedByGate`, and the pipeline continues with the unmodified message state from the previous step.