How to Configure SmartCrusher for JSON Array and Object Compression in Headroom

Configure SmartCrusher through three preservation mechanisms—preserve_change_points, preserve_fields, and preserve_keys—to protect critical data while compressing large JSON payloads in the Headroom framework.

SmartCrusher is the default JSON compressor in the chopratejas/headroom repository, designed to shrink massive tool-output arrays while preserving the specific items large language models actually need. By tuning these configuration options in crates/headroom-core/src/transforms/smart_crusher/config.rs, you can ensure headers, footers, user-specific fields, and key ordering survive aggressive compression.

Understanding SmartCrusher's Preservation Mechanisms

SmartCrusher applies three additive rules to determine which JSON array elements and object keys survive compression. These rules are defined in config.rs and orchestrated through planning.rs and crusher.rs.

Change-Point Preservation with preserve_change_points

The preserve_change_points option ensures the first and last N items of any array remain untouched, protecting header and footer metadata such as timestamps, request IDs, and error messages.

In crates/headroom-core/src/transforms/smart_crusher/planning.rs, the planner automatically marks these boundary items as "anchors" (must-keep) before any scoring occurs. The comment block around line 17 documents this logic. The default configuration keeps the first and last 10 items when this option is enabled.

This mechanism is controlled by the boolean preserve_change_points field in SmartCrusherConfig, which defaults to true according to the source in config.rs.

Field-Based Preservation with preserve_fields

Use preserve_fields to retain array elements containing specific domain values. When you provide field names like ["user_id", "request_id"], SmartCrusher computes SHA-256 hashes (truncated to 8 bytes) of these names and stores them in preserve_field_hashes.

During the anchor detection phase in planning.rs, the function item_has_preserve_field_match checks if any query token matches a field value. If a match exists, the entire array element is preserved.

// config.rs defines the field
preserve_fields: Option<Vec<String>>,

By default, this option is None (disabled), requiring explicit configuration to activate.

Key-Order Preservation with preserve_keys

For JSON objects (dictionaries), preserve_keys ensures specified keys maintain their insertion order and never get dropped during serialization. This leverages serde_json::preserve_order in the underlying Rust implementation.

The TypeScript SDK exposes this as preserve_keys in HeadroomConfig (defined in sdk/typescript/src/types/config.ts), accepting an array of key names. The default is an empty list [], meaning no special key ordering is enforced unless configured.

Configuration Flow in the Source Code

The configuration struct originates in config.rs and flows through the compression pipeline:

  1. Configuration Creation: SmartCrusherConfig is instantiated with your preservation settings
  2. Planning Stage: The SmartCrusherPlanner constructor (new) receives the config and applies the three preservation rules to mark anchors
  3. Execution Stage: crates/headroom-core/src/transforms/smart_crusher/crusher.rs receives the removal list and drops only non-anchor items, guaranteeing original array order per the comment at line 109

This architecture ensures that preservation logic is applied during planning, while the crusher focuses solely on executing the removal plan.

Practical Configuration Examples

Python SDK Configuration

When using the Python SDK, pass preservation options directly to HeadroomClient or SmartCrusherConfig:

from headroom import HeadroomClient, compress, SmartCrusherConfig

# Client-based configuration

client = HeadroomClient(
    preserve_change_points=True,  # Keep first/last 10 items (default)

    preserve_fields=["user_id", "trace_id"],  # Keep items containing these fields

)

# Direct function call

config = SmartCrusherConfig(
    preserve_change_points=True,
    preserve_fields=["trace_id"],
)
payload = [{"role": "assistant", "content": "..."}]
compressed = compress(payload, config=config)

The HeadroomClient constructor forwards these parameters to the Rust core via config.rs.

TypeScript SDK Configuration

The TypeScript SDK maps options 1-to-1 to the Rust SmartCrusherConfig:

import { HeadroomClient } from "headroom-ai";

const client = new HeadroomClient({
  preserve_change_points: true,  // Default: true
  preserve_keys: ["session_id"], // Preserve key order
  preserve_fields: ["order_id", "invoice_id"], // Field-based retention
});

These types are defined in sdk/typescript/src/types/config.ts and maintain type safety with the underlying Rust implementation.

Verifying Preservation in Tests

The repository includes integration tests demonstrating preservation behavior:

// From sdk/typescript/test/integration.test.ts around line 105
it("compress() preserves message structure for small inputs", async () => {
  const messages = [{ role: "user", content: "hello" }];
  const result = await client.compress(messages);
  expect(result.compressed).toBe(true);
});

Summary

  • SmartCrusher compresses JSON arrays in Headroom through three configurable preservation mechanisms defined in crates/headroom-core/src/transforms/smart_crusher/config.rs.
  • preserve_change_points (default true) protects the first and last 10 array items via anchor logic in planning.rs.
  • preserve_fields uses SHA-256 hashing to retain elements containing specific field values, checked by item_has_preserve_field_match.
  • preserve_keys maintains insertion order for specified object keys using serde_json::preserve_order.
  • Configuration flows from the SDK through SmartCrusherPlanner to the crusher, which respects preservation flags at line 109 of crusher.rs.

Frequently Asked Questions

How does SmartCrusher decide which array items to keep?

SmartCrusher marks items as "anchors" during the planning stage in planning.rs if they satisfy any of three rules: they fall within the first or last N items (change-point preservation), contain a field matching the query context (field-based preservation), or belong to a protected key list. Only non-anchor items are eligible for removal when crusher.rs executes the compression.

What is the default behavior for change-point preservation?

By default, preserve_change_points is set to true in SmartCrusherConfig, which protects the first and last 10 items of every array. You can disable this by setting it to false in your HeadroomClient or SmartCrusherConfig initialization, though this risks losing critical header and footer metadata.

How does field-based preservation handle sensitive data?

The preserve_fields option hashes field names using SHA-256 (truncated to 8 bytes) before comparison, storing these as preserve_field_hashes. During anchor detection in planning.rs, the system checks if query tokens match these hashes rather than performing plaintext comparison, providing a balance between functionality and privacy.

Can I use SmartCrusher without the SDK wrappers?

Yes. You can instantiate SmartCrusherConfig directly and pass it to the low-level compress function in Python, or construct the configuration object manually in Rust. The compression pipeline accepts the configuration struct directly, bypassing the high-level client abstractions while maintaining full access to preserve_change_points, preserve_fields, and preserve_keys.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →