What Is TokenJuice Compression and How OpenHuman Reduces Token Costs

TokenJuice compression is OpenHuman's built-in, model-driven layer that shrinks large text payloads into compact placeholder tokens before sending them to an LLM, directly lowering API costs by reducing token count.

TokenJuice compression powers the cost-optimization strategy in the tinyhumansai/openhuman repository. This system intercepts lengthy outputs from agent tools—such as search results or document excerpts—and replaces them with tiny surrogate tokens before they reach the language model, preserving the full data in a local cache for later retrieval.

How TokenJuice Compression Works

Configuration Schema

TokenJuice behavior is controlled through the tokenjuice configuration block defined in src/openhuman/inference/tokenjuice/schemas.rs. The schema exposes fields such as ml_compression_enabled, ml_model_id, ml_device, ml_target_ratio, and ml_max_input_chars, which determine whether compression runs, which model performs the compression, and the target compression ratio.

Python-Side Compression Engine

When ml_compression_enabled is true, the Rust core spawns a short-lived Python sidecar via the registry in src/openhuman/runtime/python_server/registry.rs. This server loads the specified ML model (e.g., a transformer encoder) and exposes a compress(text, options) function that receives raw text and returns a compact placeholder token representing the compressed payload.

Cache Management and Retrieval

Compressed mappings are stored in workspace_dir/state/tokenjuice_savings.json. The built-in tokenjuice_retrieve tool, implemented in src/openhuman/inference/tokenjuice/tools.rs, looks up placeholder tokens in this cache and returns the original uncompressed text whenever the LLM or a downstream tool requires the full content.

Token Savings Accounting

OpenHuman tracks cumulative savings in src/openhuman/inference/tokenjuice/savings.rs. The system persists statistics to tokenjuice_savings.json and exposes JSON-RPC endpoints—tokenjuice_savings_stats and tokenjuice_savings_reset—allowing developers to query or reset counters programmatically.

Automatic Integration

Agent tools that generate large textual payloads automatically invoke the compressor when the feature is enabled. The glue code in src/openhuman/modules/tokenjuice_host.rs orchestrates the handoff between the Rust runtime and the Python compression server, ensuring large inputs are transparently replaced with minimal tokens before reaching the LLM context window.

Implementation Examples

Enable TokenJuice compression in your core configuration file:

[openhuman]

# Enable the compression layer

tokenjuice.ml_compression_enabled = true
tokenjuice.ml_model_id = "tinyhuggingface/gpt-2-compressor"
tokenjuice.ml_device = "cpu"
tokenjuice.ml_target_ratio = 0.2
tokenjuice.ml_max_input_chars = 12000

Compress large tool output programmatically:

// The core automatically compresses large tool output
let compressed_token = crate::openhuman::inference::tokenjuice::ml::compress(
    &large_text,
    &config.tokenjuice,
).await?;

Retrieve original content using the built-in tool:

use openhuman_core::openhuman::inference::tokenjuice::TokenjuiceRetrieveTool;

let tool = TokenjuiceRetrieveTool::new();
let original = tool.run(json!({ "token": compressed_token })).await?;

Query savings statistics via JSON-RPC:

let stats = client.call("openhuman.tokenjuice_savings_stats", json!({})).await?;
println!("Saved {} tokens so far!", stats["saved_tokens"]);

Performance Benefits of TokenJuice Compression

  • Cost Reduction: By sending placeholder tokens instead of full text, OpenHuman reduces the per-request token count billed by providers like OpenAI or Anthropic.
  • Latency Improvements: Smaller payloads decrease network transfer time and speed up LLM inference cycles.
  • Data Integrity: The original text remains recoverable on-demand through the tokenjuice_retrieve tool, ensuring no information loss during compression.

Summary

Frequently Asked Questions

Does TokenJuice compression degrade the quality of information sent to the LLM?

No, the LLM receives a compact placeholder token that represents the full content; when the model needs to reference specific details, the framework automatically retrieves the original text via the tokenjuice_retrieve tool. This ensures the LLM operates on the complete information without inflating the context window.

Which machine learning models can be used for the compression engine?

Any transformer or encoder model compatible with the Python sidecar can be specified via the ml_model_id field in the configuration. The system loads the designated model from Hugging Face or local paths through the registry in src/openhuman/runtime/python_server/registry.rs.

How do I retrieve the original text after it has been compressed?

Use the tokenjuice_retrieve tool implemented in src/openhuman/inference/tokenjuice/tools.rs, which accepts the placeholder token and returns the full text from the workspace cache stored at workspace_dir/state/tokenjuice_savings.json.

Is TokenJuice compression enabled by default in OpenHuman?

No, compression is opt-in. You must explicitly set ml_compression_enabled = true in the tokenjuice configuration block; until enabled, the system passes raw text directly to the LLM without modification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →