How TokenJuice Compression Middleware Functions in the OpenHuman Agent Turn Path

The TokenJuice compression middleware intercepts model outputs in the TinyAgents pipeline to apply configurable lossy compression before payloads reach the client, preserving originals in a local content cache for later retrieval.

TokenJuice is the OpenHuman subsystem responsible for shrinking large language model outputs through context-aware compression. Integrated into the TinyAgents middleware stack, this TokenJuice compression middleware evaluates every agent turn against user-defined profiles to determine whether to compress content before transmission.

TokenJuice Compression Middleware Pipeline Integration

The compression logic is encapsulated in ContextCompressionMiddleware, which the TinyAgents harness inserts into the middleware stack during session initialization. According to the source code in openhuman/agent/tinyagents/mod_part_02.rs and openhuman/agent/tinyagents/mod_part_03.rs, the middleware is positioned after standard tool-output and budget middlewares but before final response serialization.

This placement ensures that the middleware operates on the complete model-generated payload immediately after inference completes but before cost accounting and client delivery occur. The middleware receives the raw turn context and consults the active compression profile to decide whether to invoke the TokenJuice engine.

Configuration Profiles and Compression Hints

The middleware’s behavior is governed by the AgentTokenjuiceCompression enum defined in openhuman/inference/tokenjuice/types.rs. This enum defines four distinct compression strategies:

  • Off – Disables compression entirely.
  • Light – Compresses only when the upstream model provides a positive compression hint.
  • Auto – Follows the model’s compression hint (default behavior).
  • Full – Forces compression on all eligible payloads regardless of hints.

The mapping from profile to execution logic occurs in openhuman/agent/tinyagents/host/budget_gate.rs at lines 220–225:

// Mapping of profile to hint (excerpt)
match self.compression {
    AgentTokenjuiceCompression::Off => CompressionHint::None,
    AgentTokenjuiceCompression::Light => match hint { ... },
    AgentTokenjuiceCompression::Auto | AgentTokenjuiceCompression::Full => hint,
}

This logic translates the high-level profile into a concrete CompressionHint that the middleware uses to gate the expensive compression operation.

TokenJuice Compression Execution Flow

When the active profile permits compression, the ContextCompressionMiddleware delegates to the TokenJuice host module. In openhuman/modules/tokenjuice_host.rs at line 23, the middleware invokes:

// TokenJuice host installs the compressor (excerpt)
crate::openhuman::inference::tokenjuice::ml::compress(&text, &options)

The underlying implementation in openhuman/inference/tokenjuice/ml/mod.rs performs the actual lossy reduction of the text payload. If the TokenJuice ML module is unavailable or fails to load, the system gracefully degrades by returning the original text unchanged, ensuring the agent turn path remains functional even when compression services are offline.

Content Cache and Recovery Mechanism

Upon successful compression, the original uncompressed text is persisted to the workspace directory at .<workspace>/.tokenjuice/ccr (Content Cache Repository). This storage strategy enables deterministic recovery of the full payload without re-running inference.

The retrieval capability is exposed through the tokenjuice_retrieve tool implemented in openhuman/inference/tokenjuice/tools.rs. Agents or client applications can invoke this tool to fetch the original text using the hash returned during the compression step.

Practical Implementation Examples

To enable TokenJuice compression when building a harness session:

use openhuman::inference::tokenjuice::AgentTokenjuiceCompression;
use openhuman::agent::harness::session::builder::SessionBuilder;

// Build a session with auto compression (the default)
let session = SessionBuilder::new()
    .tokenjuice_compression(AgentTokenjuiceCompression::Auto)
    .build()
    .await?;

To explicitly disable compression for debugging or low-latency scenarios:

let session = SessionBuilder::new()
    .tokenjuice_compression(AgentTokenjuiceCompression::Off)
    .build()
    .await?;

To recover a compressed payload later in the conversation:

use openhuman::inference::tokenjuice::retrieve;

// `hash` is the token returned by the compression step
let original_text = retrieve(hash.to_string(), None).await?;

Summary

  • The TokenJuice compression middleware (ContextCompressionMiddleware) is integrated into the TinyAgents middleware stack between model inference and client response delivery.
  • Compression behavior is controlled by the AgentTokenjuiceCompression enum, mapped to execution hints in budget_gate.rs (lines 220–225).
  • Active compression invokes ml::compress in the TokenJuice host module, with implementation details in openhuman/inference/tokenjuice/ml/mod.rs.
  • Original texts are archived to .<workspace>/.tokenjuice/ccr and recoverable via the tokenjuice_retrieve tool.
  • Configuration occurs at session build time through SessionBuilder, supporting granular control from full compression to complete disablement.

Frequently Asked Questions

Where exactly does the TokenJuice middleware sit in the agent turn lifecycle?

The middleware is positioned in the TinyAgents pipeline immediately after model output generation and budget accounting, but before the final RPC response or UI rendering phase. This ensures compression overhead does not affect inference latency while minimizing payload sizes for transmission.

What is the difference between the Auto and Light compression profiles?

Auto follows the model’s explicit compression hint without additional gating, compressing whenever the upstream system suggests it is beneficial. Light adds a conditional check that only compresses when the hint indicates a high-confidence opportunity for size reduction, preserving bandwidth for marginal cases.

How does the system handle TokenJuice ML module failures?

If the ML module referenced in openhuman/inference/tokenjuice/ml/mod.rs fails to load or crashes during compression, the ml::compress function returns the original text unchanged. This fail-safe behavior, implemented in the host module at openhuman/modules/tokenjuice_host.rs, guarantees agent availability even when compression services are degraded.

Can I retrieve the original uncompressed text after it has been compressed?

Yes. Every compressed payload stores its original form in the workspace’s .tokenjuice/ccr directory. You can recover the full text using the tokenjuice_retrieve tool defined in openhuman/inference/tokenjuice/tools.rs, passing the unique hash identifier returned during the initial compression call.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →