How FreeLLMAPI's Prompt Compression Works: A Deep Dive into the Pipeline

FreeLLMAPI reduces incoming prompts before they reach LLM providers, cutting token costs while preserving answer quality through a configurable compression pipeline with built-in fidelity checks.

FreeLLMAPI's prompt compression system shrinks chat messages to save on API costs and reduce latency. This article examines the exact mechanism implemented in the tashfeenahmed/freellmapi repository, walking through the compression pipeline from request entry to quality validation.

What Triggers Prompt Compression

Compression activates in two scenarios. First, when a request includes the X-FreeLLM-Compress header with a specific engine name. Second, when calling the dedicated preview endpoint at /api/compression/preview to inspect compression results before sending to a provider.

The Compression Pipeline Architecture

The core flow spans seven distinct stages across multiple source files in server/src/services/compression/.

1. Pipeline Entry Point

The compressRequest function in [pipeline.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/pipeline.ts) receives raw chat messages and a configuration object. This function orchestrates the entire transformation process.

2. Configuration Loading

Settings are read from [config.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/config.ts), including:

  • Enabled compression mode
  • Per-engine parameters
  • Fidelity-gate thresholds

3. Engine Selection and Execution

Based on the compression mode, the pipeline routes messages to the appropriate engine in server/src/services/compression/engines/:

Engine File Behavior
jsoncompact [engines/jsoncompact.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/engines/jsoncompact.ts) Removes redundant JSON structures, whitespace, and comments
aging [engines/aging.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/engines/aging.ts) Collapses old conversation turns into [compressed older turn] placeholders

Each engine performs message-level transformation by rewriting the prompt according to its specific strategy.

4. Fidelity Gate Validation

Before accepting compressed output, checkFidelity in [fidelity-gate.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/fidelity-gate.ts) evaluates similarity between original and compressed prompts using a cheap heuristic. If similarity falls below the configured threshold, the request reverts to uncompressed to prevent quality degradation.

5. Metrics Collection

getCompressionStats in [stats.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/stats.ts) tracks:

  • Number of compressed requests
  • Original vs. compressed character counts
  • Per-engine savings

These metrics expose through the /api/compression/stats endpoint defined in [routes/compression.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/compression.ts).

6. Response Header Construction

After compression, formatCompressionHeader builds the X-FreeLLM-Compress response header containing:

  • Compression fingerprint
  • Original and estimated token counts
  • Fidelity-gate actions taken

7. Provider Forwarding

The (possibly compressed) messages proceed to the LLM provider with full transparency to the client.

Code Examples: Using FreeLLMAPI Prompt Compression

Preview Compression Before Sending

import fetch from 'node-fetch';

const messages = [
  { role: 'system', content: 'You are a helpful assistant.' },
  { role: 'user', content: 'Explain quantum entanglement in simple terms.' },
];

const preview = await fetch('https://api.freellmapi.com/api/compression/preview', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({ messages }),
});

const { compressed, stats } = await preview.json();

console.log('Compressed messages:', compressed);
console.log('Saved chars:', stats.originalChars - stats.compressedChars);

Request Compression via Header

const response = await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'X-FreeLLM-Compress': 'jsoncompact',
  },
  body: JSON.stringify({ messages }),
});

console.log('Compression header:', response.headers.get('X-FreeLLM-Compress'));
const data = await response.json();

Key Files in the Compression System

Purpose Path
Pipeline orchestration server/src/services/compression/pipeline.ts
Configuration and fingerprinting server/src/services/compression/config.ts
JSON-compact engine server/src/services/compression/engines/jsoncompact.ts
Aging engine server/src/services/compression/engines/aging.ts
Quality validation server/src/services/compression/fidelity-gate.ts
Statistics tracking server/src/services/compression/stats.ts
HTTP routes server/src/routes/compression.ts

Summary

  • FreeLLMAPI prompt compression activates via header or preview endpoint, routing through compressRequest in pipeline.ts
  • Multiple compression engines (jsoncompact, aging) target different prompt patterns for size reduction
  • The fidelity gate guarantees quality by rejecting transformations that alter semantic meaning beyond threshold
  • Full observability via headers and stats endpoints lets clients track savings per request

Frequently Asked Questions

What compression engines does FreeLLMAPI support?

FreeLLMAPI currently implements two engines as of the main branch: jsoncompact for structural JSON optimization and aging for conversational context summarization. Additional engines can be added under server/src/services/compression/engines/.

How does the fidelity gate prevent quality loss?

The checkFidelity function in fidelity-gate.ts runs a fast similarity heuristic comparing original and compressed prompts. When similarity drops below the threshold set in config.ts, compression is discarded and the original prompt proceeds unchanged.

Can I see compression results before sending to the LLM?

Yes. The /api/compression/preview endpoint returns compressed messages and statistics without forwarding to any provider. This allows validation of size savings and content preservation before production use.

What information appears in the X-FreeLLM-Compress header?

The header contains a compression fingerprint, original character/token counts, compressed counts, and any fidelity-gate decisions. Parse this header to log savings or trigger fallback behavior when compression is rejected.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →