What Are the Prompt Compression Modes in OmniRoute? A Complete Guide
OmniRoute supports seven distinct prompt compression modes—off, lite, standard, aggressive, ultra, rtk, and stacked—that let you trade latency for token economy using rule-based engines, SLM rewriting, or composable pipelines.
The diegosouzapw/OmniRoute repository implements a flexible compression framework that allows developers to select from predefined strategies or build custom pipelines. These prompt compression modes are defined in the core configuration layer and can be applied globally or per-request to optimize token usage.
The Seven Prompt Compression Modes Explained
In src/lib/db/compression.ts, the COMPRESSION_MODES constant enumerates the available strategies as a TypeScript set containing seven distinct options:
- off — Disables compression entirely; prompts are transmitted unchanged to the upstream model. This mode bypasses all transformation engines.
- lite — Runs a lightweight engine that performs fast, cheap transformations such as whitespace collapsing and system-prompt deduplication. Ideal for low-latency scenarios.
- standard — Activates the default "caveman" engine, a rule-based semantic condenser that applies balanced reductions including system-prompt deduplication and role compression.
- aggressive — Applies a heavier compressor that enforces summarization thresholds and stronger tool-output reductions. Configuration is controlled by the
aggressiveconfig block. - ultra — Leverages a Small-Language-Model (SLM) to rewrite the prompt, aiming for maximal token savings while preserving semantic meaning. Managed via the
ultraconfig block. - rtk — Executes a rule-based "RTK" engine that filters and deduplicates tool results according to configurable business rules.
- stacked — Enables a composable pipeline that chains multiple engines (such as lite, caveman, aggressive, ultra, and rtk) in a user-defined order defined by the
stackedPipelinesetting.
How Modes Map to Compression Engines
According to the OmniRoute source code, the framework separates the user-facing mode from the concrete engine implementation using two key mappings in src/lib/db/compression.ts.
The SINGLE_MODE_ENGINE map resolves single-mode selections to their underlying engine IDs:
const SINGLE_MODE_ENGINE: Partial<Record<CompressionMode, string>> = {
lite: "lite",
standard: "caveman",
aggressive: "aggressive",
ultra: "ultra",
rtk: "rtk",
};
When using stacked mode, the pipeline references engines listed in STACKED_PIPELINE_ENGINE_IDS, which includes:
const STACKED_PIPELINE_ENGINE_IDS = new Set([
"lite",
"caveman",
"aggressive",
"ultra",
"rtk",
"headroom",
"session-dedup",
"ccr",
"llmlingua",
]);
Configuring Prompt Compression Modes
You can interact with these modes programmatically using the settings API exposed in src/lib/db/compression.ts.
Reading the Current Mode
import { getCompressionSettings } from "@/lib/db/compression";
async function showMode() {
const cfg = await getCompressionSettings();
console.log("Current mode:", cfg.defaultMode); // e.g. "standard"
}
showMode();
Switching to Ultra Compression
import { updateCompressionSettings } from "@/lib/db/compression";
async function enableUltra() {
await updateCompressionSettings({
defaultMode: "ultra",
ultra: { enabled: true }
});
}
enableUltra();
Building a Stacked Pipeline
Configure a custom sequence running lite → caveman → rtk:
import { updateCompressionSettings } from "@/lib/db/compression";
async function setStackedPipeline() {
await updateCompressionSettings({
defaultMode: "stacked",
stackedPipeline: [
{ engine: "lite" },
{ engine: "caveman", intensity: "full" },
{ engine: "rtk", intensity: "standard" },
],
});
}
setStackedPipeline();
Verifying Pipeline Configuration
import { getCompressionSettings } from "@/lib/db/compression";
async function logPipeline() {
const cfg = await getCompressionSettings();
console.log(cfg.stackedPipeline);
}
logPipeline();
Summary
- OmniRoute provides seven prompt compression modes: off, lite, standard, aggressive, ultra, rtk, and stacked.
- The
COMPRESSION_MODESconstant insrc/lib/db/compression.tsdefines the valid mode strings. - Single modes map to engines via
SINGLE_MODE_ENGINE, where standard maps to thecavemanengine. - Stacked mode allows composing engines from the
STACKED_PIPELINE_ENGINE_IDSset, including specialized options likeheadroom,session-dedup, andllmlingua. - Use
getCompressionSettings()andupdateCompressionSettings()to read and modify configurations programmatically.
Frequently Asked Questions
What is the difference between standard and ultra prompt compression modes?
Standard mode uses the caveman engine, a rule-based semantic condenser that applies balanced reductions like role compression and deduplication. Ultra mode instead invokes a Small-Language-Model (SLM) to rewrite the entire prompt, achieving higher token savings at the cost of increased latency and compute.
Can I combine multiple compression engines in OmniRoute?
Yes. The stacked mode allows you to compose a custom pipeline using the stackedPipeline configuration array. You can chain any engines listed in STACKED_PIPELINE_ENGINE_IDS, such as lite, caveman, and rtk, applying them sequentially in your specified order.
Where are compression settings stored in OmniRoute?
Compression configurations are persisted through the database layer defined in src/lib/db/compression.ts and src/lib/db/compressionCombos.ts. The getCompressionSettings() and updateCompressionSettings() functions handle retrieval and persistence, typically backed by the application's configured database.
How do I disable prompt compression entirely?
Set the defaultMode to off using updateCompressionSettings({ defaultMode: "off" }). This mode bypasses all transformation engines and sends prompts to the upstream model exactly as constructed, consuming the maximum token allowance but eliminating compression latency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →