How FreeLLMAPI's Prompt Compression Works: A Deep Dive into the Pipeline
FreeLLMAPI reduces incoming prompts before they reach LLM providers, cutting token costs while preserving answer quality through a configurable compression pipeline with built-in fidelity checks.
FreeLLMAPI's prompt compression system shrinks chat messages to save on API costs and reduce latency. This article examines the exact mechanism implemented in the tashfeenahmed/freellmapi repository, walking through the compression pipeline from request entry to quality validation.
What Triggers Prompt Compression
Compression activates in two scenarios. First, when a request includes the X-FreeLLM-Compress header with a specific engine name. Second, when calling the dedicated preview endpoint at /api/compression/preview to inspect compression results before sending to a provider.
The Compression Pipeline Architecture
The core flow spans seven distinct stages across multiple source files in server/src/services/compression/.
1. Pipeline Entry Point
The compressRequest function in [pipeline.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/pipeline.ts) receives raw chat messages and a configuration object. This function orchestrates the entire transformation process.
2. Configuration Loading
Settings are read from [config.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/config.ts), including:
- Enabled compression mode
- Per-engine parameters
- Fidelity-gate thresholds
3. Engine Selection and Execution
Based on the compression mode, the pipeline routes messages to the appropriate engine in server/src/services/compression/engines/:
| Engine | File | Behavior |
|---|---|---|
| jsoncompact | [engines/jsoncompact.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/engines/jsoncompact.ts) |
Removes redundant JSON structures, whitespace, and comments |
| aging | [engines/aging.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/engines/aging.ts) |
Collapses old conversation turns into [compressed older turn] placeholders |
Each engine performs message-level transformation by rewriting the prompt according to its specific strategy.
4. Fidelity Gate Validation
Before accepting compressed output, checkFidelity in [fidelity-gate.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/fidelity-gate.ts) evaluates similarity between original and compressed prompts using a cheap heuristic. If similarity falls below the configured threshold, the request reverts to uncompressed to prevent quality degradation.
5. Metrics Collection
getCompressionStats in [stats.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/compression/stats.ts) tracks:
- Number of compressed requests
- Original vs. compressed character counts
- Per-engine savings
These metrics expose through the /api/compression/stats endpoint defined in [routes/compression.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/compression.ts).
6. Response Header Construction
After compression, formatCompressionHeader builds the X-FreeLLM-Compress response header containing:
- Compression fingerprint
- Original and estimated token counts
- Fidelity-gate actions taken
7. Provider Forwarding
The (possibly compressed) messages proceed to the LLM provider with full transparency to the client.
Code Examples: Using FreeLLMAPI Prompt Compression
Preview Compression Before Sending
import fetch from 'node-fetch';
const messages = [
{ role: 'system', content: 'You are a helpful assistant.' },
{ role: 'user', content: 'Explain quantum entanglement in simple terms.' },
];
const preview = await fetch('https://api.freellmapi.com/api/compression/preview', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ messages }),
});
const { compressed, stats } = await preview.json();
console.log('Compressed messages:', compressed);
console.log('Saved chars:', stats.originalChars - stats.compressedChars);
Request Compression via Header
const response = await fetch('https://api.freellmapi.com/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'X-FreeLLM-Compress': 'jsoncompact',
},
body: JSON.stringify({ messages }),
});
console.log('Compression header:', response.headers.get('X-FreeLLM-Compress'));
const data = await response.json();
Key Files in the Compression System
| Purpose | Path |
|---|---|
| Pipeline orchestration | server/src/services/compression/pipeline.ts |
| Configuration and fingerprinting | server/src/services/compression/config.ts |
| JSON-compact engine | server/src/services/compression/engines/jsoncompact.ts |
| Aging engine | server/src/services/compression/engines/aging.ts |
| Quality validation | server/src/services/compression/fidelity-gate.ts |
| Statistics tracking | server/src/services/compression/stats.ts |
| HTTP routes | server/src/routes/compression.ts |
Summary
- FreeLLMAPI prompt compression activates via header or preview endpoint, routing through
compressRequestinpipeline.ts - Multiple compression engines (
jsoncompact,aging) target different prompt patterns for size reduction - The fidelity gate guarantees quality by rejecting transformations that alter semantic meaning beyond threshold
- Full observability via headers and stats endpoints lets clients track savings per request
Frequently Asked Questions
What compression engines does FreeLLMAPI support?
FreeLLMAPI currently implements two engines as of the main branch: jsoncompact for structural JSON optimization and aging for conversational context summarization. Additional engines can be added under server/src/services/compression/engines/.
How does the fidelity gate prevent quality loss?
The checkFidelity function in fidelity-gate.ts runs a fast similarity heuristic comparing original and compressed prompts. When similarity drops below the threshold set in config.ts, compression is discarded and the original prompt proceeds unchanged.
Can I see compression results before sending to the LLM?
Yes. The /api/compression/preview endpoint returns compressed messages and statistics without forwarding to any provider. This allows validation of size savings and content preservation before production use.
What information appears in the X-FreeLLM-Compress header?
The header contains a compression fingerprint, original character/token counts, compressed counts, and any fidelity-gate decisions. Parse this header to log savings or trigger fallback behavior when compression is rejected.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →