FreeLLMAPI Prompt Compression Modes: A Complete Guide to `off`, `lossless`, `standard`, and `aggressive`
FreeLLMAPI supports four distinct prompt compression modes—off, lossless, standard, and aggressive—that control how incoming chat prompts are reduced before being sent to a language model.
The server codebase implements these modes in a dedicated compression service, allowing developers to optimize token usage based on their fidelity requirements. This article explores how each mode works, when to use them, and how to configure them via the API.
What Are Prompt Compression Modes?
Prompt compression modes determine how aggressively FreeLLMAPI reduces the token count of incoming prompts. In server/src/services/compression/types.ts, the COMPRESSION_MODES constant defines the four available options, and the CompressionMode type is derived from this array. The pipeline in server/src/services/compression/pipeline.ts then selects appropriate compression engines based on your chosen mode.
The Four Compression Modes
off — No Compression
The off mode forwards prompts unchanged. Use this for debugging or when your provider handles compression internally.
Characteristics:
- Zero token reduction
- Full semantic preservation
- Useful for inspecting raw requests
lossless — Conservative Reductions
lossless applies transformations that never alter semantic content, such as removing redundant whitespace and deduplicating identical messages.
Characteristics:
- Minor token savings
- Guaranteed semantic equivalence
- Ideal when every token matters but some inefficiency exists
standard — Balanced Optimization
standard combines lossy and lossless transformations to reduce token count while retaining overall meaning. This is the recommended default for most production workloads.
Characteristics:
- Good token savings with minimal fidelity loss
- Balanced engine selection
- General-purpose compression for most models
aggressive — Maximum Reduction
aggressive enables the most intensive optimizations, including aggressive JSON compacting and tool-result filtering. Use this when strict token limits apply.
Characteristics:
- Maximum token reduction
- Potentially lossy transformations
- Best for models that tolerate approximation
How to Configure Compression Modes
Update Settings via API
Configure the active mode through the compression settings endpoint:
// Example using fetch to update the compression configuration
await fetch('https://your-free-llmapi-instance.com/api/settings/compression', {
method: 'PUT',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
mode: 'standard', // Choose one of: off | lossless | standard | aggressive
engines: { // Engine-specific options (optional)
jsoncompact: { enabled: true },
},
trustProjectFilters: true,
prefixFreeze: false,
}),
});
The mode field accepts any CompressionMode value. Optional engines configuration allows fine-tuning specific compression engines per mode.
Preview Compression Results
Test how a prompt compresses under different modes before applying them:
const response = await fetch('https://your-free-llmapi-instance.com/api/compression/preview', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
header: 'User',
tools: [{ name: 'search', description: 'Web search', parameters: {} }],
messages: [
{ role: 'user', content: 'Explain quantum entanglement in simple terms.' },
],
previewMode: 'aggressive', // Try any of the four modes
}),
});
const { messages, stats } = await response.json();
console.log('Compressed prompt:', messages);
console.log('Compression stats:', stats);
The preview endpoint returns both the compressed messages and statistics about the reduction achieved.
Inspect Current Configuration
Retrieve the active compression settings:
const { mode, engines } = await (await fetch('/api/settings/compression')).json();
console.log(`Current mode: ${mode}`);
console.log('Engine options:', engines);
This queries server/src/routes/compression.ts, which interfaces with server/src/services/compression/config.ts for configuration management.
Key Implementation Files
| File | Purpose |
|---|---|
server/src/services/compression/types.ts |
Defines COMPRESSION_MODES and CompressionMode type |
server/src/services/compression/config.ts |
Parses, validates, and stores compression configuration |
server/src/routes/compression.ts |
API routes for settings and preview endpoints |
server/src/services/compression/pipeline.ts |
Core pipeline applying engines based on selected mode |
server/docs/compression.md |
Extended documentation on mode behavior |
Summary
- Four modes available:
off,lossless,standard, andaggressive, defined inserver/src/services/compression/types.ts - Progressive trade-offs: From zero compression (
off) to maximum reduction (aggressive) - API-configurable: Set via
PUT /api/settings/compressionwith optional engine-specific tuning - Preview before applying: Use
POST /api/compression/previewto test compression results - Pipeline-driven:
server/src/services/compression/pipeline.tsorchestrates engine selection based on mode
Frequently Asked Questions
What is the default compression mode in FreeLLMAPI?
The default mode depends on your initial configuration. Check the current setting via GET /api/settings/compression against server/src/services/compression/config.ts. Most deployments start with standard as the balanced default.
Can I use different compression modes for different API endpoints?
Yes. Mode selection is request-scoped through the preview endpoint, or you can implement middleware that dynamically swaps configurations via the settings API before forwarding requests.
Does aggressive compression affect function calling or tool results?
Potentially. The aggressive mode may apply tool-result filtering and aggressive JSON compacting. Test thoroughly with server/src/services/compression/pipeline.ts debug logging if your application relies heavily on structured tool outputs.
How do I disable compression entirely for debugging?
Set mode: 'off' via the settings endpoint. This bypasses all compression engines in server/src/services/compression/pipeline.ts and forwards prompts unchanged.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →