How TokenJuice Compression Middleware Works in the OpenHuman Agent Turn Path

TokenJuice compresses prompt and model-response payloads before they reach the LLM by running an ML-based neural compressor in a Python side-car, then stores the original text in an in-memory cache so the agent can retrieve it later via the tokenjuice_retrieve tool.

TokenJuice is the built-in compression layer inside the OpenHuman agent pipeline that reduces token usage and latency on every turn. According to the tinyhumansai/openhuman source code, it intercepts the TurnRequest after the orchestrator assembles it and before the LLM call, optionally shrinking the payload with a small neural model. This article traces the exact file paths, function calls, and configuration flags that govern how TokenJuice compression middleware functions within the agent turn path.

Three Stages of TokenJuice in the Agent Turn Path

1. Turn Preparation and Orchestrator Check

In src/openhuman/tools/orchestrator_tools.rs, the orchestrator builds a TurnRequest and consults the OrchestratorToolSettings field tokenjuice_compression. When this flag is set to Auto—the default—the request is routed toward the compressor before any LLM call is made.

2. ML Compression via the Host Bridge

The request text reaches src/openhuman/modules/tokenjuice_host.rs, where tokenjuice_host::install forwards it to crate::openhuman::inference::tokenjuice::ml::compress. This function delegates to a Python side-car launched by src/openhuman/runtime/python_server/server.rs. The side-car runs a small neural model—identified by ml_model_id such as gpt2_tokenjuice—and returns a compact string plus a cache token.

3. Retrieval with the tokenjuice_retrieve Tool

After the LLM returns a result, the agent can invoke the special tool implemented in src/openhuman/inference/tokenjuice/tools.rs. TokenjuiceRetrieveTool::new() looks up the original text in the in-memory TokenJuice cache using the stored token and returns the uncompressed content to the agent.

Configuring the TokenJuice Compression Middleware

The TokenJuice behavior is driven by the [tokenjuice] table inside .openhuman/config.toml. Its schema is defined in src/openhuman/config/schema/tokenjuice.rs, and changes are applied through live reload via install_from_config in src/openhuman/inference/tokenjuice/schemas.rs.

The fields that control compression are:

  • ml_compression_enabled — Enables the ML compressor. When set to false, every turn bypasses compression.
  • ml_model_id — Identifier of the on-disk model the Python side-car loads, for example gpt2_tokenjuice.
  • ml_target_ratio — Desired size-reduction ratio, such as 0.25 for 4× compression.
  • ml_max_input_chars — Maximum number of characters the compressor will accept in a single call.
  • ml_sidecar_idle_timeout_secs — How long the side-car process can stay idle before it shuts down.

Step-by-Step Agent Turn Path Flow

The full agent turn path moves through seven concrete steps, from request creation to text retrieval. Each step is implemented by specific functions and modules in the tinyhumansai/openhuman codebase.

  1. Orchestrator creates the turn. It assembles a TurnRequest containing the full prompt text.
  2. Middleware gate check. The orchestrator reads tokenjuice_compression from OrchestratorToolSettings. If the value is Auto and the config permits, the request enters tokenjuice_host::install.
  3. Neural compression. The host calls openhuman::inference::tokenjuice::ml::compress, which communicates with the Python side-car. The compressor produces a compact representation and a cache token.
  4. Original text cached. The uncompressed text and its token are stored in the in-memory TokenJuice cache located under src/openhuman/inference/tokenjuice/cache.
  5. Compressed LLM request. The smaller payload is sent to the LLM, cutting token consumption and latency.
  6. LLM response handling. If the LLM response carries a token field indicating compressed data, the agent prepares to expand it.
  7. Retrieval. The agent calls tokenjuice_retrieve, which uses TokenjuiceRetrieveTool::new() to fetch the original text from the cache and passes the uncompressed data back to subsequent tools or the UI.

Compression savings statistics are persisted to workspace_dir/state/tokenjuice_savings.json and exposed through the RPC endpoints tokenjuice_savings_stats and tokenjuice_savings_reset defined in src/openhuman/inference/tokenjuice/savings.rs.

When TokenJuice Compression Is Skipped

If tokenjuice_compression is set to Never or ml_compression_enabled is false, the turn bypasses the compressor and travels to the LLM unchanged. This fallback is useful when debugging large payloads or when the Python side-car cannot start due to missing dependencies.

Practical Code Examples

The following snippets show how to enable, configure, and interact with TokenJuice compression middleware in the OpenHuman agent turn path. These examples mirror the actual types and file paths found in the tinyhumansai/openhuman repository.

Enabling TokenJuice in config.toml


# .openhuman/config.toml

[tokenjuice]
ml_compression_enabled = true
ml_model_id = "gpt2_tokenjuice"
ml_target_ratio = 0.25
ml_max_input_chars = 100_000
ml_sidecar_idle_timeout_secs = 300

Configuring the Orchestrator Flag

use openhuman_core::openhuman::inference::tokenjuice::AgentTokenjuiceCompression;

let settings = OrchestratorToolSettings {
    tokenjuice_compression: AgentTokenjuiceCompression::Auto,
    ..Default::default()
};

Retrieving Original Text Manually

let retrieve_tool = crate::openhuman::inference::tokenjuice::TokenjuiceRetrieveTool::new();
let original = retrieve_tool
    .run(json!({ "token": token }))
    .await?
    .into_string();

Querying Savings Statistics via RPC

const stats = await coreRpcClient.call(
    "openhuman.tokenjuice_savings_stats", {}
);
console.log("Saved characters:", stats.saved_chars);

Key Source Files

These files implement the complete pipeline from configuration to compression to retrieval. Referencing them directly makes debugging and extending the middleware straightforward.

Summary

  • TokenJuice compression middleware intercepts TurnRequest objects in the orchestrator before they reach the LLM.
  • Compression is handled by ml::compress inside tokenjuice_host.rs, executed in a Python side-car.
  • The tokenjuice_compression flag in OrchestratorToolSettings controls whether compression runs on each turn.
  • Original texts are cached so the TokenjuiceRetrieveTool in tools.rs can restore compressed LLM responses.
  • Administrators configure behavior via .openhuman/config.toml, with schema validation in src/openhuman/config/schema/tokenjuice.rs.

Frequently Asked Questions

What triggers TokenJuice compression during an agent turn?

The tokenjuice_compression field inside OrchestratorToolSettings—defined in src/openhuman/tools/orchestrator_tools.rs—must be set to Auto and ml_compression_enabled must be true in the config. When both conditions are met, the orchestrator routes the TurnRequest through tokenjuice_host::install before the LLM call.

Where does the actual neural compression run?

The neural model executes in a dedicated Python side-car process launched by src/openhuman/runtime/python_server/server.rs. The Rust host bridge in src/openhuman/modules/tokenjuice_host.rs forwards text to crate::openhuman::inference::tokenjuice::ml::compress, which serializes the request, waits for the side-car, and returns the compact payload plus a cache token.

How can an agent recover the original text from a compressed response?

The agent calls the tokenjuice_retrieve tool implemented in src/openhuman/inference/tokenjuice/tools.rs. TokenjuiceRetrieveTool::new() accepts a JSON payload containing the cache token, looks up the original text in the in-memory TokenJuice cache, and returns the uncompressed string to the caller.

What happens if the Python side-car is unavailable?

If the side-car cannot start or ml_compression_enabled is set to false, the middleware skips compression and the turn proceeds to the LLM with its original payload intact. You can also force this behavior by setting tokenjuice_compression: AgentTokenjuiceCompression::Never in the orchestrator settings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →