How to Control Prompt Compression on a Per‑Request Basis in FreeLLMAPI

Use the X-FreeLLM-Compress request header to override compression settings for individual API calls without changing global configuration.

FreeLLMAPI processes prompt compression before caching, routing, and provider selection occur. While operators set a baseline mode via dashboard or environment variables, clients can fine-tune compression behavior per-request using a dedicated HTTP header. This guide explains the complete mechanism as implemented in the tashfeenahmed/freellmapi repository.

How the X-FreeLLM-Compress Header Works

The compression pipeline in server/src/services/compression/pipeline.ts reads the X-FreeLLM-Compress header and merges it with the global configuration. The compressRequest function normalizes the header value and resolves the effective compression mode using these rules:

Header value Behavior
off Disables compression entirely for this request, regardless of global settings
on Uses the currently configured global mode (e.g., lossless, standard, or aggressive)
lossless / standard / aggressive Forces that specific engine only if it does not exceed the global mode; otherwise capped to global setting

The pipeline cannot be used to escalate compression beyond what the operator allows. If a client requests aggressive but the global mode is standard, the pipeline silently downgrades to standard. This preserves the "master‑switch" guarantee that clients cannot bypass operator-configured limits.

Sending Per‑Request Compression Headers

Disable Compression for a Single Request

Force zero compression when you need raw prompt fidelity:

await fetch('https://api.free.llm/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'X-FreeLLM-Compress': 'off',
  },
  body: JSON.stringify({
    model: 'gpt-4',
    messages: [{ role: 'user', content: 'Detailed technical specification...' }]
  })
});

Use the Globally Configured Mode

Explicitly opt into the server's current setting without hardcoding a value:

await fetch('https://api.free.llm/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'X-FreeLLM-Compress': 'on',
  },
  body: JSON.stringify({
    model: 'gpt-4o',
    messages: [{ role: 'user', content: 'Summarize this article...' }]
  })
});

Request a Specific Engine (Downgrade-Only)

Attempt aggressive compression, which succeeds only if permitted globally:

await fetch('https://api.free.llm/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'X-FreeLLM-Compress': 'aggressive',
  },
  body: JSON.stringify({
    model: 'claude-3-opus',
    messages: [{ role: 'user', content: 'Long document for analysis...' }]
  })
});

If the global mode is standard, this request receives standard compression without error.

Inspecting Compression Results

The server returns the actual compression mode applied and estimated savings in the response header X-FreeLLM-Compress:

const response = await fetch(/* ... */);
const compressionInfo = response.headers.get('X-FreeLLM-Compress');

// Example: "standard; saved~=1840"
console.log('Compression used:', compressionInfo);

This feedback loop lets clients verify whether their requested mode was honored or downgraded.

Implementation Details in the Source Code

According to the FreeLLMAPI source code, the per-request control flow spans several key files:

The behavior is documented in docs/compression.md (lines 28–38), which confirms the header value semantics and downgrade-only enforcement.

Global vs. Per‑Request Priority

The resolution hierarchy ensures operator control:

  1. Client sends X-FreeLLM-Compress header
  2. Pipeline normalizes the value
  3. Ceiling check: Requested mode is capped to global mode if higher
  4. Resolved mode executes through compression engines
  5. Result metadata returned in response header

This design balances client flexibility with operator governance—clients can reduce compression intensity or accept the default, but never exceed configured limits.

Summary

  • X-FreeLLM-Compress: off guarantees zero compression for sensitive prompts
  • X-FreeLLM-Compress: on defers to the globally configured mode
  • Specific engine values (lossless, standard, aggressive) work only at or below the global setting
  • Response header X-FreeLLM-Compress reveals the actual mode applied and character savings
  • The compressRequest function in server/src/services/compression/pipeline.ts implements all header parsing and mode resolution

Frequently Asked Questions

What happens if I send an invalid X-FreeLLM-Compress value?

The pipeline treats unrecognized values as no override, falling back to the global configuration. No error is returned; the request proceeds with default compression.

Can I use per-request compression with all FreeLLMAPI endpoints?

The header is honored by any route invoking the compression pipeline, including chat completions via server/src/routes/responses.ts. Administrative endpoints in server/src/routes/compression.ts bypass client headers and use direct configuration access.

Why can't I escalate above the global mode?

This restriction prevents clients from circumventing cost controls, latency guarantees, or content policies set by operators. The "master-switch" design in server/src/services/compression/pipeline.ts enforces this ceiling during mode resolution.

How do I check my current global compression setting?

Query the admin endpoint or inspect the FREELLMAPI_COMPRESSION environment variable. The dashboard and server/src/routes/compression.ts routes expose current configuration without requiring code changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →