How TokenJuice Compression Reduces Tool Output Tokens by Up to 80% in OpenHuman

TokenJuice compression reduces agent tool output tokens by up to 80 % using a model‑driven Python side‑car that rewrites large payloads into compact semantic representations, caching the originals for on‑demand retrieval via the tokenjuice_retrieve tool.

OpenHuman’s TokenJuice (often shortened to tokenjuice) is a lightweight compressor that intercepts tool outputs before they reach the LLM. When an agent’s tool produces a large textual payload, the core system hands the raw output to the tokenjuice service, which runs a small language model to shrink the content while preserving meaning. This design pattern directly lowers token usage costs and prevents context window overflow during agent turns.

The Six‑Step Compression Pipeline

The TokenJuice compression flow operates as a coordinated pipeline across multiple Rust modules. Each step is designed to minimize latency while maximizing compression ratios.

Step 1: Configuration Loading The core reads the [tokenjuice] configuration block from the workspace config, defined in src/openhuman/config/schema/tokenjuice.rs. Key parameters include ml_compression_enabled, ml_target_ratio, and ml_max_input_chars.

Step 2: Host Invocation When a tool response exceeds size thresholds, the tokenjuice host in src/openhuman/modules/tokenjuice_host.rs calls crate::openhuman::inference::tokenjuice::ml::compress(&text, &options). This entry point bridges the synchronous tool execution flow with the asynchronous compression service.

Step 3: Python Side‑car Execution The ml::compress function spawns a short‑lived Python process coordinated by src/openhuman/runtime/python_server/kompress.rs. This module loads the configured model (ml_model_id) and executes the compress routine, rewriting the input to meet the ml_target_ratio (default 0.25) while respecting ml_max_input_chars.

Step 4: Caching and Token Generation The compressed result is stored in the TokenJuice cache (src/openhuman/inference/tokenjuice/cache.rs). A short hash token is generated and returned to the core, which replaces the original payload with a {tokenjuice} marker of minimal token count.

Step 5: Retrieval Tool If the model later requires the full original text, it invokes the built‑in tokenjuice_retrieve tool implemented in src/openhuman/inference/tokenjuice/tools.rs. The tool reads the cached entry using the hash token and streams the uncompressed content back to the model.

Step 6: Savings Tracking Every compression event updates tokenjuice_savings.json via src/openhuman/inference/tokenjuice/savings.rs. Empirical measurements in the test suite (tests/raw_coverage/..._raw_coverage_e2e.rs) demonstrate reductions from roughly 6 k tokens to 1.2 k tokens, achieving the advertised up‑to‑80 % token reduction per turn.

Configuration and Tuning

Fine‑tuning TokenJuice compression behavior starts with the configuration schema in src/openhuman/config/schema/tokenjuice.rs. The system exposes several levers to balance compression aggressiveness against information fidelity.

  • ml_compression_enabled: Boolean gate that disables the entire pipeline when set to false, causing the core to bypass the compressor and send raw outputs unchanged.
  • ml_target_ratio: Controls the desired tokens‑to‑character ratio (default 0.25). Lowering this value increases compression but risks semantic loss.
  • ml_max_input_chars: Sets the character threshold that triggers compression; content below this limit passes through uncompressed.
  • ml_model_id: Specifies the distilled model (e.g., a distilled GPT‑2 variant) used by the Python side‑car for rewriting.

Core Implementation Architecture

The TokenJuice subsystem spans seven primary source files that implement distinct responsibilities:

src/openhuman/modules/tokenjuice_host.rs Acts as the core entry point. This module determines when compression is required based on payload size and delegates to the inference layer.

src/openhuman/runtime/python_server/kompress.rs Orchestrates the Python side‑car process lifecycle. It handles model loading, process spawning, and inter‑process communication for the compression task.

src/openhuman/inference/tokenjuice/ml.rs Provides the Rust wrapper around the Python process. This file defines the compress API signature and manages the serialization of text and options across the language boundary.

src/openhuman/inference/tokenjuice/cache.rs Implements the storage layer for compressed payloads. It generates deterministic hash tokens and manages cache eviction policies to prevent memory bloat.

src/openhuman/inference/tokenjuice/tools.rs Contains the tokenjuice_retrieve tool definition. This module allows agents to recover full original text on demand by querying the cache with a hash token.

src/openhuman/inference/tokenjuice/savings.rs Persists token‑saving statistics to disk. This data powers dashboards and telemetry showing cumulative compression benefits across agent sessions.

Enabling and Using TokenJuice Compression

Enable Automatic Compression

To enable automatic compression for all qualifying tool outputs in an orchestrator turn, configure the builder with AgentTokenjuiceCompression::Auto:

use openhuman_core::openhuman::inference::tokenjuice::AgentTokenjuiceCompression;

let orchestrator = OrchestratorBuilder::new()
    .tokenjuice_compression(AgentTokenjuiceCompression::Auto)
    .build()
    .await?;

The orchestrator now invokes the TokenJuice host for every tool output that exceeds ml_max_input_chars.

Retrieve Original Payloads Manually

Models or downstream tools can recover the full uncompressed text using the tokenjuice_retrieve tool:

use openhuman_core::openhuman::inference::tokenjuice::TokenjuiceRetrieveTool;

let retrieve = TokenjuiceRetrieveTool::new();
let result = retrieve
    .run(json!({ "token": "a1b2c3d4" }))
    .await?;
println!("Full original text: {}", result.output);

The hash token (a1b2c3d4 in this example) corresponds to the compressed entry stored in the cache.

Monitor Token Savings via RPC

Frontend applications can query compression statistics through the RPC layer:

const stats = await coreRpcClient.call("openhuman.tokenjuice_savings_stats", {});
console.log(`Saved ${stats.saved_tokens} tokens (${stats.savings_percent}% reduction)`);

This integration allows real‑time display of session‑level token efficiency metrics.

Summary

Frequently Asked Questions

How does TokenJuice compression achieve an 80 % reduction in tool output tokens?

TokenJuice employs a model‑driven rewrite strategy rather than naive truncation. When a tool returns a large payload (e.g., 6 k tokens), the Python side‑car loads a distilled language model and generates a semantically equivalent paraphrase that retains critical information while minimizing token count. Empirical tests in the OpenHuman repository measured outputs compressed to approximately 1.2 k tokens, yielding an 80 % reduction before the LLM processes the turn.

What happens to the original tool output after compression?

The original text is never discarded. It is stored in the TokenJuice cache (src/openhuman/inference/tokenjuice/cache.rs), indexed by a short hash token. The LLM receives only a {tokenjuice} marker and the compact representation. If the model requires the full text later in the conversation, it can invoke the tokenjuice_retrieve tool to stream the uncached original from disk.

Can TokenJuice compression be disabled for specific agent turns?

Yes. The system respects the ml_compression_enabled boolean in the [tokenjuice] configuration block. When set to false, the core bypasses the compressor entirely, sending raw tool outputs directly to the model. Additionally, content shorter than ml_max_input_chars automatically skips compression, ensuring small payloads incur no overhead.

Which machine learning model powers the TokenJuice side‑car?

The compression service uses a configurable distilled language model specified by the ml_model_id parameter. According to the source analysis, the default configuration typically references a distilled GPT‑2 variant or similar lightweight model capable of running efficiently within a short‑lived Python process spawned by src/openhuman/runtime/python_server/kompress.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →