TokenJuice Compression in OpenHuman: How to Reduce LLM Token Costs by 30-50%

OpenHuman's TokenJuice compression system reduces LLM token costs by 30-50% using syntax-tree parsing and placeholder substitution to minimize transmission size while preserving semantic context.

The tinyhumansai/openhuman repository implements a fine-grained compression layer specifically designed to trim token usage without sacrificing essential context. Located under src/openhuman/inference/tokenjuice, this subsystem parses prompts into concrete syntax trees and replaces repetitive fragments with compact placeholders that reconstruct on the model side.

How TokenJuice Compression Works

The TokenJuice compression pipeline operates through six deterministic stages that transform verbose prompts into compact representations before LLM transmission.

The Compression Pipeline

  1. Prompt Construction – When assembling a conversation turn, the core invokes openhuman::inference::tokenjuice::compress via the re-export in src/openhuman/inference/tokenjuice/mod.rs.

  2. AST Generation – The utils.rs module employs tree-sitter parsers to generate a concrete syntax tree of the prompt, enabling structural analysis rather than raw token counting.

  3. Placeholder Insertion – Repetitive or deterministic sections (such as repeated document excerpts or fixed code blocks) are replaced by short placeholders like «DOC_1», with original text stored in a PlaceholderMap.

  4. Budget Enforcement – The algorithm iteratively removes the least-impactful placeholders until the token count falls below the configured target budget.

  5. Transmission – Both the compressed prompt and its placeholder map serialize into the JSON-RPC request payload.

  6. Model-Side Expansion – A tiny-agent shim within the Tokenizer runtime expands placeholders back into original text before the model processes the full prompt.

Core Implementation Files

The TokenJuice compression logic resides in four key Rust source files under the src/openhuman/inference/tokenjuice directory.

compress.rs

The src/openhuman/inference/tokenjuice/compress.rs file implements the public compress function, which accepts the original prompt string and a target token budget, returning a CompressedPrompt struct containing the compact representation and reconstruction mapping.

types.rs

src/openhuman/inference/tokenjuice/types.rs declares the core data structures: CompressedPrompt holds the compressed payload, while PlaceholderMap maintains the lookup table required for model-side reconstruction.

utils.rs

The src/openhuman/inference/tokenjuice/utils.rs module provides helper routines for token counting, tree-sitter based parsing, and deterministic placeholder generation.

mod.rs

src/openhuman/inference/tokenjuice/mod.rs publicly re-exports the compression API, enabling easy consumption by the rest of the OpenHuman core.

Implementing TokenJuice Compression in Rust

Developers integrate TokenJuice compression by calling the compress function with their raw prompt and desired token budget.

Basic Compression Example

The following example compresses a long UI-generated prompt before LLM transmission:

use openhuman::inference::tokenjuice::{compress, CompressedPrompt};

/// Example: compress a long UI‑generated prompt before sending it to the LLM.
fn main() -> Result<(), Box<dyn std::error::Error>> {
    // The raw prompt that includes a long document excerpt and a code block.
    let raw_prompt = r#"
        Summarize the following article:

        -----BEGIN ARTICLE-----
        Lorem ipsum dolor sit amet, consectetur adipiscing elit…
        (many paragraphs)
        -----END ARTICLE-----

        Also, explain the Rust snippet:

        ```rust
        fn main() {
            println!("Hello, world!");
        }
        ```

    "#;

    // Target a token budget of 800 tokens (the model’s limit is 4096).
    let target_budget = 800;

    // Perform compression.
    let CompressedPrompt { prompt, placeholder_map } = compress(raw_prompt, target_budget)?;

    // `prompt` is now a compact string containing placeholders like «DOC_1».
    // `placeholder_map` holds the original fragments for the model‑side expansion.
    println!("Compressed prompt ({} tokens):\n{}", prompt.len(), prompt);
    println!("Placeholder map: {:#?}", placeholder_map);

    // The `CompressedPrompt` can be passed directly to the Core RPC:
    // core_rpc_client.send_prompt(prompt, placeholder_map);
    Ok(())
}

Tool Integration Example

For custom tools requiring strict token limits, wrap the compression step directly in the handler:

// Inside a custom tool that needs to keep token usage low:
use openhuman::inference::tokenjuice::compress;

fn tool_handler(input: &str) -> Result<String, anyhow::Error> {
    // Assume the tool must stay under 300 tokens.
    let compressed = compress(input, 300)?;
    // Send the compressed payload to the model.
    let response = model_api::call(compressed.prompt, compressed.placeholder_map)?;
    Ok(response)
}

Architectural Benefits of TokenJuice

TokenJuice compression offers several distinct advantages for production LLM workflows:

  • Language-Agnostic Processing – Because the system operates on tree-sitter parse trees rather than language-specific tokenizers, it compresses any textual prompt regardless of programming language or natural language.

  • Deterministic Mapping – The placeholder generation algorithm produces consistent results; identical source text always yields the same placeholder identifier, facilitating caching and debugging.

  • Safety Validation – The compress function in compress.rs validates that the placeholder map does not exceed the model's maximum context length, failing gracefully with a clear error when the target budget cannot be met.

  • Extensible API – Developers can implement new placeholder strategies (such as semantic summarization) by extending utils.rs without modifying the core compression loop in compress.rs.

Summary

  • TokenJuice compression resides in src/openhuman/inference/tokenjuice within the OpenHuman repository.
  • The compress function in compress.rs returns a CompressedPrompt containing compact placeholders and a reconstruction map.
  • Tree-sitter parsing in utils.rs enables structural analysis of prompts for optimal compression.
  • The system reduces token usage by 30-50% for multi-document workflows while maintaining deterministic, reversible mappings.
  • Safety checks prevent context length violations, and the modular design supports strategy extensions without core modifications.

Frequently Asked Questions

What is TokenJuice compression?

TokenJuice compression is a subsystem in the OpenHuman framework designed to minimize LLM token transmission costs. It works by parsing prompts into concrete syntax trees and substituting repetitive fragments with compact placeholders like «DOC_1». These placeholders expand back into the original text on the model side, typically achieving 30-50% size reduction for multi-document workflows.

How does the compress function handle budget constraints?

The compress function in src/openhuman/inference/tokenjuice/compress.rs implements an iterative algorithm that removes the least-impactful placeholders until the token count falls below the specified target. If the target budget cannot be achieved even with maximum compression, the function fails gracefully with a clear error rather than sending an oversized prompt. This safety check ensures the placeholder map never exceeds the model's maximum context length.

Is TokenJuice compression language-specific?

No, TokenJuice compression is fully language-agnostic because it uses tree-sitter parsers to analyze prompt structure at the syntactic level. Unlike language-specific tokenizers, tree-sitter processes any textual content equally effectively, whether source code, natural language, or mixed documents. This architectural choice enables consistent compression across diverse prompt types without requiring language-specific plugins.

Can TokenJuice compression be used with any LLM?

Yes, TokenJuice compression works with any LLM connected through the OpenHuman inference layer because the placeholder expansion occurs in a tiny-agent shim on the model side. The compressed prompt and its PlaceholderMap transmit via standard JSON-RPC, and the reconstruction happens within the Tokenizer runtime before the model processes the full text. This design decouples the compression format from specific LLM implementations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →